Persistence
Persistence lets a session survive process restarts and gives product surfaces a stable session ID. A resumed session can repopulate task lists, execution history, artifacts, and delivery summaries without replaying the completed run.
A3S Code keeps four related but different records:
File Session Store
Go selects the built-in file store with FileSessionStoreDir; custom
SessionStore trait implementations remain a Rust embedding capability.
The Fact Log Is The Control Source
Every coding turn appends immutable facts to a per-session log in the
workspace, under .a3s/effect-log/<thread>.jsonl. The thread is the session
ID when it is at most 128 characters of ASCII letters, digits, _, ., :,
or -; otherwise it is s- followed by the hex-encoded ID (truncated to 128
characters). Two sessions in one workspace therefore fold two separate logs. send, stream, attachment turns,
resumeRun, and exact recovery all choose the next step the same way: they
fold the log with a3s-effect (ingest_coding for a new fact,
resume_coding to continue). See core/src/fact_control.rs.
This has concrete consequences for persistence and recovery:
- A stored
model.turnfact is never sent to the model again. - A loop checkpoint or a session snapshot does not choose the next model call. Only the log does.
- A confirmation parks until a
confirmation.answeredfact arrives, and a question parks until aquestion.answeredfact arrives. No in-process channel or timer settles them. - A tool call whose result is missing from the log runs once on resume.
- A steer is another
user.messagefact. - When the tool-round cap is reached, the next completion is sent with an empty tool list.
The fact log lives in the workspace, not in SessionSnapshotV1. Keep the
workspace (or at least .a3s/effect-log/) together with the session store if
you expect a restarted process to continue a parked run.
Atomic Snapshot Generations
session.save() gathers the current persisted state into one
SessionSnapshotV1 and calls SessionStore::save_snapshot once. The envelope
contains:
schema_versionandSessionData- tool artifacts and the retained artifact URIs used by artifact GC
- trace events and run records
- verification reports
- delegated subagent task snapshots
- review findings, when the session has any
The file store writes one complete JSON envelope to a synced temporary file and
atomically replaces v1/sessions/id_<base64url session id>.json under the
store directory. Readers therefore observe the previous generation or the next
one, not a new conversation paired with old run or trace fragments. Loop
checkpoints live beside it in v1/loop_checkpoints/. The memory store
publishes the same aggregate under one lock.
Older files remain readable. A legacy <session-id>.json in the store root is
a read-only migration source. A bare SessionData file is combined with the
fragment locations for artifacts, traces, runs, verification reports, and
subagent tasks during load, then restored through the v1 in-memory shape. Once
a new aggregate is saved, the single envelope is authoritative. A document that
already looks like an aggregate but has a malformed or unsupported schema is
rejected rather than reinterpreted as fragment data.
Custom stores must implement save_snapshot explicitly. Its default returns an
error; it does not fan an aggregate out into independent writes or silently
acknowledge a no-op. The default load_snapshot only assembles fragments on a
best-effort basis, and capabilities() lets a host distinguish that behavior
from an atomic backend.
Negotiated Durability
SessionStore::capabilities() returns SessionStoreCapabilities, which says
which guarantees a backend implements. Hosts should check a flag before relying
on it. The built-in stores advertise:
The WAL stays unencrypted even when encryption at rest is enabled; it holds
digests only. The trait defaults for save_snapshot_cas (with an expected
digest), acquire_writer_lease, and watch_commits return an error, so a
custom store that does not implement them fails loudly.
Resume A Session
resumeSession restores a saved session snapshot. It is not the same as
resumeRun: use resumeSession when the user is continuing a saved
conversation, and use resumeRun(runId) to continue an interrupted run.
Resume validates the snapshot schema before restoring any history or runtime
evidence.
Resume An Interrupted Run
session.resumeRun(runId) (Node.js) / session.resume_run(run_id) (Python) /
session.ResumeRun(ctx, runID) (Go) / AgentSession::resume_run (Rust)
continues work that stopped before it settled, for example after a crash or
while a confirmation was parked. The fact log decides what happens next:
- If the
SessionStoreholds a loop checkpoint forrunId, the checkpoint must belong to this run and session, or the call fails withrefusing to resume checkpoint '<id>'. The session's fact-log thread is then deleted and reseeded with the checkpoint's messages, and the log is folded. The checkpoint supplies facts; it does not pick the next model call. - Otherwise, if the session's fact log already has facts, that log is folded as it is. A quiescent log takes no step.
- Otherwise the call fails. Without a store the error says
resume_run requires a session_store on this session; with a store it saysno loop checkpoint found for run '<id>'.
Resume folds at most 32 steps per call. When a checkpoint is used, the resumed work is recorded as a new run (the checkpoint's run is not modified), the checkpoint's token usage and tool-call count are added to the result, and the completion gate applies: a turn that mutated the workspace completes only with a Passed verification report bound to that mutation.
Headless hosts that need an exact, replay-safe recovery identity use
spawnRecoveryWithRunId(checkpointRunId, runId) (spawn_recovery_with_run_id
in Python, SpawnRecoveryWithRunID in Go). It requires a checkpoint for
checkpointRunId, admits the recovery under the host-chosen runId, and
resumes the fact log in a detached worker. Replaying the same recovery runId
returns the existing snapshot instead of executing again. The checkpoint run
itself stays immutable.
Completed runs are not resumed. Their final state is available through
runs(), runEvents(runId), artifacts, verification reports, and session
snapshots.
Loop checkpoint format
A LoopCheckpoint is written through the session's checkpoint sink after each
tool result lands on the fact log. Its messages hold that tool call and its
result; it also records the run and session IDs, the run's capability binding,
a turn counter, the tool-call count, and checkpoint_ms. SessionStore provides save_loop_checkpoint,
load_loop_checkpoint, and delete_loop_checkpoint; the file store writes
atomically. LoopCheckpoint::ensure_loadable() rejects checkpoints from a
future schema version, and ensure_owned_by rejects a checkpoint that belongs
to another run or session. When a run ends in-process (completed, failed, or
cancelled), its checkpoint is deleted.
Workflow Checkpoints
WorkflowCheckpoint journals completed orchestration steps so an interrupted
workflow does not run them again. Its fields are
schema_version, workflow_id, steps, and checkpoint_ms, with the schema
pinned by the WORKFLOW_CHECKPOINT_SCHEMA_VERSION constant. A resuming run
skips the recorded steps and re-dispatches only the rest.
SessionStore provides save_workflow_checkpoint /
load_workflow_checkpoint / delete_workflow_checkpoint (default no-ops; the
file store writes atomically). Loads from a future, incompatible schema version
are rejected by WorkflowCheckpoint::ensure_loadable().
The SDKs expose the journal through parallelResumable(specs, workflowId)
(Node.js), parallel_resumable (Python), and
ParallelResumable(ctx, specs, workflowID) (Go); each requires a session store.
See Orchestration and
Multi-Machine.
Resuming On A Different Node
Loop and workflow checkpoints are serializable and live in the shared
SessionStore. A host can therefore recover an interrupted run or workflow on
a different node from the one that started it. For a run, the recovering node
seeds its own fact log from the checkpoint and folds from there. The framework
owns the serializable contract; the host owns placement and transport.
Memory And Sessions
Session persistence stores conversation and replay evidence. Memory stores reusable task facts. Use both when you want resumable workflows that also learn from repeated tasks.
Operational Notes
Keep session stores and .a3s/effect-log/ out of public commits when they may
contain prompts, tool output, or private file paths.