Memory
Memory records reusable facts about previous work. It should help the harness recall patterns without flooding every prompt. A3S Code has two layers:
- Session memory — a
MemoryStoreevery session gets by default. The runtime extracts memories from completed turns and recalls related items before each turn. - Durable memory — an optional, host-bound A3S Memory repository
(
DurableMemorySession). Recall admits only nodes the host explicitly activated. This is a Rust-host integration surface.
Default Store
Every session gets a memory store. Unless you pass one, the session uses a
file-backed store at memory_dir from the loaded config, or
<workspace>/.a3s/memory when memory_dir is not set. If the file store
cannot be created, session creation fails with a session-initialization error
that names the directory.
Override Stores
Pass a typed store object to override the backend for one session:
Automatic Recall
At the start of each turn, Core recalls up to five session memories similar to
the prompt and adds them to the budgeted context assembly (see
Context). When a durable-memory binding is present, its
admitted Active nodes are added too; if both layers return the same normalized
text, the durable item is used once. Each recalled item is reported as a
memory_recalled event, followed by one memories_searched event. Recall is
skipped entirely when the session is bound to an exact cognitive package.
LLM Extraction
LLM memory extraction is enabled by default when memory is available. Every
completed turn with a non-empty prompt and response is submitted to the active
model. The model, not a keyword list or tool-type heuristic, decides whether
the turn contains anything that can change a future answer or action. It must
return an empty items array when nothing qualifies. Tool calls and tool
results are context for that judgment; they are never mechanically copied into
long-term memory.
Automatic output is deliberately narrow. The runtime accepts only semantic
memories for durable facts, preferences, and decisions, or procedural
memories for reusable workflows and failure lessons. Each item must include a
validated source, scope, and non-empty future-value reason, with
importance of at least 0.70 and confidence of at least 0.75. Stored
metadata also records the workspace, session ID, and a3s.memory.durable.v1
schema. The runtime rejects malformed output and obvious API keys, tokens,
password assignments, and private keys; these checks validate structure and
safety rather than judging semantic value.
The extraction prompt includes up to five related existing memories. The model
can use supersedes for a directly replaced or consolidated memory and
conflicts_with for a contradiction that should remain visible. The runtime
accepts only relation IDs it supplied to the model. It removes accepted
superseded items, preserves conflicts, and includes relation annotations in
future recall.
Extraction jobs are queued in FIFO order per session. Streaming runs extract in
the background after the final event, so a slow extraction does not delay the
response; non-streaming send runs extract before returning. Graceful session
close waits up to five seconds for extraction accepted before close.
Tune extraction in the config file:
Store Hygiene
Default memory stores return the canonical item for normalized exact duplicate
content, raising importance and preserving useful tags, provenance, and
relation metadata. Distinct but related wording remains separate unless the LLM
explicitly consolidates it through supersedes.
Automatic pruning removes stale, low-importance memories when configured, but it
hard-protects curated memories: pinned/protected items, frequently recalled
items, consolidated memories, and memories carrying supersedes or
conflicts_with relation metadata.
Write and Recall Explicitly
rememberSuccess and rememberFailure are explicit SDK operations for a host
that intentionally wants to store those records. The agent runtime does not
call them for each tool result.
Use recalled memory as supporting context. Verification evidence still comes from current commands and traces.
Durable Memory (Rust Host)
A Rust host can bind a session to one exact A3S Memory repository namespace
with SessionOptions::with_durable_memory. Code owns extraction, activation
gates, and context admission; the repository owns namespace isolation,
revisions, and atomic change sets; the host owns repository, embedding
provider, and vector-index construction plus tenant/principal/scope selection.
Active recall
DurableMemorySession::active_recall is the only binding mode
(DurableMemoryMode::ActiveRecall):
- Extraction writes evidence-backed Candidate nodes into the namespace. Candidates are never queried for prompt context.
- The host activates one exact candidate revision with
activate_candidate(DurableMemoryActivation). Activation requires new Manual or Verification decision evidence; LLM confidence and importance cannot authorize it. - Recall runs a bounded, Active-only lexical query
(
a3s.memory.lexical.word-cjk-bigram.v1), optionally followed by a bounded number of one-hopRelatedToreads.DurableMemoryRecallPolicy::try_newtakes the result bound and minimum lexical score;try_with_related_lookupsenables relation reads (default zero). - After final context assembly Code records an admission for each selected
node revision. A stale or inactive admission is removed before the model
call. The host calls
record_useonly when a node was actually used.
Code fills an omitted tenant_id and principal from the namespace; explicit
values must match. The live repository is not serialized. The session snapshot
stores the secret-free DurableMemoryBindingV1, and resume fails closed unless
the host injects an equivalent binding again.
Semantic recall and refresh
DurableMemorySemanticRecall::new pairs a host-supplied EmbeddingProvider
with a caller-owned A3S Memory VectorIndex; attach it with
with_semantic_recall. At query time Code embeds the query, searches only the
partition derived from the namespace and serving generation, re-reads every
candidate from the repository, and requires Active status plus the exact
node revision and content digest. Verified lexical and semantic ranks are
fused with reciprocal-rank fusion (a3s.code.memory.hybrid.rrf-k60.v1). Any
embedding, vector, or verification failure drops the semantic branch and keeps
the lexical result.
Vectors are not updated implicitly. Refresh them from a complete Active-only repository snapshot either on demand:
or on a session-owned schedule, which always requires index-revision CAS:
Hosts read maintenance health with memoryMaintenanceHealth() (Node),
memory_maintenance_health() (Python), or MemoryMaintenanceHealth (Go).
Production qualification harness
core/examples/durable_memory_prod1_host is a host qualification harness
(DM-PROD1) that exercises a Redis VectorIndex with index-revision CAS,
fenced leases, restart, failover, and drift against a real OpenAI-compatible
embedding provider. It requires the dm-prod1-host feature and is not part of
any release profile. See
DURABLE_MEMORY_PRODUCTION_QUALIFICATION.md
and the full integration manual
DURABLE_MEMORY.md.