Persistence

Persistence lets a session survive process restarts and gives product surfaces a stable session ID. A resumed session can repopulate task lists, execution history, artifacts, and delivery summaries without replaying the completed run.

A3S Code keeps four related but different records:

RecordWritten byRead byPurpose
Fact logEvery coding turn (send, stream, attachments)The next turn, resumeRun, and exact recoveryThe only control source. The next transition is chosen by folding this log.
SessionSnapshotV1session.save() or autoSaveagent.resumeSession(id, options)Rehydrate one versioned generation with the conversation, artifacts, traces, run records, verification reports, and subagent task snapshots.
Loop checkpointEach tool result that lands on the fact logresumeRun(runId) and spawnRecoveryWithRunIdPortable evidence of a run in progress, stored in the SessionStore so another process or node can recover it. Deleted when the run ends in-process.
Workflow checkpointparallelResumable / workflow combinatorsparallelResumable(specs, workflowId)Skip orchestration steps that already completed after a process restart.

File Session Store

Rust
Node.js
Python
Go
Rust
use a3s_code_core::{Agent, SessionOptions};
let options = SessionOptions::new()
.with_session_id("release-review")
.with_file_session_store("./.a3s/sessions")
.with_auto_save(true);
let session = agent
.session_builder("/repo")
.options(options)
.build()
.await?;
session.send("Review release readiness", None).await?;
session.save().await?;

Go selects the built-in file store with FileSessionStoreDir; custom SessionStore trait implementations remain a Rust embedding capability.

The Fact Log Is The Control Source

Every coding turn appends immutable facts to a per-session log in the workspace, under .a3s/effect-log/<thread>.jsonl. The thread is the session ID when it is at most 128 characters of ASCII letters, digits, _, ., :, or -; otherwise it is s- followed by the hex-encoded ID (truncated to 128 characters). Two sessions in one workspace therefore fold two separate logs. send, stream, attachment turns, resumeRun, and exact recovery all choose the next step the same way: they fold the log with a3s-effect (ingest_coding for a new fact, resume_coding to continue). See core/src/fact_control.rs.

This has concrete consequences for persistence and recovery:

  • A stored model.turn fact is never sent to the model again.
  • A loop checkpoint or a session snapshot does not choose the next model call. Only the log does.
  • A confirmation parks until a confirmation.answered fact arrives, and a question parks until a question.answered fact arrives. No in-process channel or timer settles them.
  • A tool call whose result is missing from the log runs once on resume.
  • A steer is another user.message fact.
  • When the tool-round cap is reached, the next completion is sent with an empty tool list.

The fact log lives in the workspace, not in SessionSnapshotV1. Keep the workspace (or at least .a3s/effect-log/) together with the session store if you expect a restarted process to continue a parked run.

Atomic Snapshot Generations

session.save() gathers the current persisted state into one SessionSnapshotV1 and calls SessionStore::save_snapshot once. The envelope contains:

  • schema_version and SessionData
  • tool artifacts and the retained artifact URIs used by artifact GC
  • trace events and run records
  • verification reports
  • delegated subagent task snapshots
  • review findings, when the session has any

The file store writes one complete JSON envelope to a synced temporary file and atomically replaces v1/sessions/id_<base64url session id>.json under the store directory. Readers therefore observe the previous generation or the next one, not a new conversation paired with old run or trace fragments. Loop checkpoints live beside it in v1/loop_checkpoints/. The memory store publishes the same aggregate under one lock.

Older files remain readable. A legacy <session-id>.json in the store root is a read-only migration source. A bare SessionData file is combined with the fragment locations for artifacts, traces, runs, verification reports, and subagent tasks during load, then restored through the v1 in-memory shape. Once a new aggregate is saved, the single envelope is authoritative. A document that already looks like an aggregate but has a malformed or unsupported schema is rejected rather than reinterpreted as fragment data.

Custom stores must implement save_snapshot explicitly. Its default returns an error; it does not fan an aggregate out into independent writes or silently acknowledge a no-op. The default load_snapshot only assembles fragments on a best-effort basis, and capabilities() lets a host distinguish that behavior from an atomic backend.

Negotiated Durability

SessionStore::capabilities() returns SessionStoreCapabilities, which says which guarantees a backend implements. Hosts should check a flag before relying on it. The built-in stores advertise:

CapabilityFile storeMemory storeMeaning
atomic_session_snapshotsyesyessave_snapshot commits one complete generation.
aggregate_casyesyessave_snapshot_cas commits only when the expected digest matches (or is omitted).
append_only_event_logyesnoDigest-only Intent/Committed WAL records around each atomic replace.
lease_fencingyesnoacquire_writer_lease publishes a durable epoch; stale holders fail closed.
encrypted_at_restwith a keynoFileSessionStore::with_encryption_key seals documents with AES-256-GCM.
watchyesyeswatch_commits delivers SessionStoreCommitEventV1 values after durable commits.
reference_aware_artifact_gcyesyesArtifact eviction keeps content reachable from retained URIs.

The WAL stays unencrypted even when encryption at rest is enabled; it holds digests only. The trait defaults for save_snapshot_cas (with an expected digest), acquire_writer_lease, and watch_commits return an error, so a custom store that does not implement them fails loudly.

Resume A Session

Rust
Node.js
Python
Go
Rust
use a3s_code_core::SessionOptions;
let resumed = agent
.resume_session_async(
"release-review",
SessionOptions::new().with_file_session_store("./.a3s/sessions"),
)
.await?;

resumeSession restores a saved session snapshot. It is not the same as resumeRun: use resumeSession when the user is continuing a saved conversation, and use resumeRun(runId) to continue an interrupted run. Resume validates the snapshot schema before restoring any history or runtime evidence.

Resume An Interrupted Run

session.resumeRun(runId) (Node.js) / session.resume_run(run_id) (Python) / session.ResumeRun(ctx, runID) (Go) / AgentSession::resume_run (Rust) continues work that stopped before it settled, for example after a crash or while a confirmation was parked. The fact log decides what happens next:

  1. If the SessionStore holds a loop checkpoint for runId, the checkpoint must belong to this run and session, or the call fails with refusing to resume checkpoint '<id>'. The session's fact-log thread is then deleted and reseeded with the checkpoint's messages, and the log is folded. The checkpoint supplies facts; it does not pick the next model call.
  2. Otherwise, if the session's fact log already has facts, that log is folded as it is. A quiescent log takes no step.
  3. Otherwise the call fails. Without a store the error says resume_run requires a session_store on this session; with a store it says no loop checkpoint found for run '<id>'.

Resume folds at most 32 steps per call. When a checkpoint is used, the resumed work is recorded as a new run (the checkpoint's run is not modified), the checkpoint's token usage and tool-call count are added to the result, and the completion gate applies: a turn that mutated the workspace completes only with a Passed verification report bound to that mutation.

Rust
Node.js
Python
Go
Rust
let result = session.resume_run("run-abc123").await?;
println!("{}", result.usage.total_tokens);

Headless hosts that need an exact, replay-safe recovery identity use spawnRecoveryWithRunId(checkpointRunId, runId) (spawn_recovery_with_run_id in Python, SpawnRecoveryWithRunID in Go). It requires a checkpoint for checkpointRunId, admits the recovery under the host-chosen runId, and resumes the fact log in a detached worker. Replaying the same recovery runId returns the existing snapshot instead of executing again. The checkpoint run itself stays immutable.

Completed runs are not resumed. Their final state is available through runs(), runEvents(runId), artifacts, verification reports, and session snapshots.

Loop checkpoint format

A LoopCheckpoint is written through the session's checkpoint sink after each tool result lands on the fact log. Its messages hold that tool call and its result; it also records the run and session IDs, the run's capability binding, a turn counter, the tool-call count, and checkpoint_ms. SessionStore provides save_loop_checkpoint, load_loop_checkpoint, and delete_loop_checkpoint; the file store writes atomically. LoopCheckpoint::ensure_loadable() rejects checkpoints from a future schema version, and ensure_owned_by rejects a checkpoint that belongs to another run or session. When a run ends in-process (completed, failed, or cancelled), its checkpoint is deleted.

Workflow Checkpoints

WorkflowCheckpoint journals completed orchestration steps so an interrupted workflow does not run them again. Its fields are schema_version, workflow_id, steps, and checkpoint_ms, with the schema pinned by the WORKFLOW_CHECKPOINT_SCHEMA_VERSION constant. A resuming run skips the recorded steps and re-dispatches only the rest.

SessionStore provides save_workflow_checkpoint / load_workflow_checkpoint / delete_workflow_checkpoint (default no-ops; the file store writes atomically). Loads from a future, incompatible schema version are rejected by WorkflowCheckpoint::ensure_loadable().

The SDKs expose the journal through parallelResumable(specs, workflowId) (Node.js), parallel_resumable (Python), and ParallelResumable(ctx, specs, workflowID) (Go); each requires a session store. See Orchestration and Multi-Machine.

Resuming On A Different Node

Loop and workflow checkpoints are serializable and live in the shared SessionStore. A host can therefore recover an interrupted run or workflow on a different node from the one that started it. For a run, the recovering node seeds its own fact log from the checkpoint and folds from there. The framework owns the serializable contract; the host owns placement and transport.

Memory And Sessions

Session persistence stores conversation and replay evidence. Memory stores reusable task facts. Use both when you want resumable workflows that also learn from repeated tasks.

Operational Notes

Keep session stores and .a3s/effect-log/ out of public commits when they may contain prompts, tool output, or private file paths.