A WebUI chat turn crosses several durability boundaries:
If the server crashes between submission and the final sidecar save, recovery has to infer what happened from pending_user_message, active_stream_id, .json.bak, _index.json, and state.db. Those safeguards are useful, but they are still reconstructing intent after the fact.
The missing primitive is a small write-ahead journal for turns: record the submitted user turn durably before the worker starts, then advance the journal as the turn progresses.
state.db.sessions or WebUI JSON sidecars.Use one JSONL file per session under the existing WebUI state area:
<SESSION_DIR>/_turn_journal/<session_id>.jsonl
Each line is an immutable event. Recovery can scan by turn_id and choose the latest status.
{
"version": 1,
"event": "submitted",
"turn_id": "20260511T001122Z-abcdef",
"session_id": "abc123",
"stream_id": "stream-xyz",
"created_at": 1778458282.123,
"role": "user",
"content": "...",
"attachments": [],
"workspace": "/workspace",
"model": "openai/gpt-5",
"model_provider": "openai"
}
Later events for the same turn_id:
{"version":1,"event":"worker_started","turn_id":"...","created_at":1778458283.0}
{"version":1,"event":"assistant_started","turn_id":"...","created_at":1778458284.0}
{"version":1,"event":"completed","turn_id":"...","created_at":1778458299.0,"assistant_message_index":12}
{"version":1,"event":"interrupted","turn_id":"...","created_at":1778458301.0,"reason":"server_startup_recovery"}
submitted -> worker_started -> assistant_started -> completed
submitted -> interrupted
worker_started -> interrupted
assistant_started -> interrupted
completed is terminal. interrupted is terminal unless a later explicit repair creates a new turn. Recovery should not silently resume a provider call.
/api/chat/start or equivalent turn-submission path:
turn_id,submitted,_run_agent_streaming, append worker_started.assistant_started.completed.interrupted with a reason.The submitted event uses synchronous fsync on every write today. This is a deliberate tradeoff between latency and crash-safety guarantees:
The submitted event is the durability anchor for the entire recovery story. If the server crashes before the worker starts, the journal must reflect that the user message was received. Async writes risk losing that guarantee: a crash shortly after a non-fsync’d write could leave the journal silent while pending_user_message still exists, creating ambiguity during recovery. The current design avoids that ambiguity at the cost of one extra disk round-trip per turn submission.
Reported fsync latency varies significantly across storage backends. Approximate qualitative ranges to keep in mind:
These ranges are order-of-magnitude guidance, not benchmarks. Exact figures depend on hardware, kernel version, filesystem mount options, and concurrent load. Do not commit specific millisecond claims to documentation without measured evidence.
If evidence suggests the synchronous write is a bottleneck, measure before changing anything:
append_turn_journal_event helper to record wall-clock time for each event type (submitted, worker_started, etc.).strace -e fsync or kernel tracing (ftrace, perf) to confirm where time is spent.Making journal writes asynchronous is a valid future optimization, but it requires:
Async journal writes are not part of the initial implementation. They belong in a follow-up RFC once the synchronous baseline is proven stable and the recovery semantics are well-understood.
On startup, for each journal file:
completed: no action.submitted or worker_started and no matching user message exists in sidecar:
submitted, worker_started, or assistant_started and no completed assistant turn exists:
.json.bak and state.db recovery still run first so the sidecar is as complete as possible before journal reconciliation.audit_session_recovery() can report:
turn_journal_pending_turn — repairable if the user message is absent from sidecar.turn_journal_interrupted_turn — ok/warn depending on whether a visible marker exists.turn_journal_malformed_event — manual review.Safe repair should only materialize submitted user messages and interruption markers when the journal event content is valid JSON and the target message is absent.
Initial read-only endpoint can be folded into the existing recovery audit:
GET /api/session/recovery/audit
Later, if needed:
GET /api/session/turn-journal?session_id=<id>
The latter should be diagnostic-only and redact or omit large attachment payloads.
turn_id so browser retry and server retry do not duplicate the same user message.assistant_started but before completed.The first implementation PR should be deliberately small:
append_turn_journal_event(session_id, event)read_turn_journal(session_id)submitted before worker startDo not combine the first implementation with replay/repair. Replay is where most of the bugs in WAL systems live; ship the writer and audit first, prove the format, then add repair.