Context
A3S Code treats context as a budgeted resource. The model should see the smallest useful context for the current decision, not every available file, skill, memory, and tool log.
Sources
Context can come from:
- user prompt and conversation history
- AGENTS.md project instructions
- skills and agent definitions
- memory stores
- file search and direct tool results
- MCP tools and context providers
- delegated task summaries
- trace events
AGENTS.md injection, skills discovery, memory APIs, direct tool results,
delegated task helpers, and traceEvents() are available from every SDK.
Custom context providers are a Rust-host extension point (see below).
Assembly
Context is resolved once at the start of each turn. ContextAssembler
deduplicates items, always keeps required items such as AGENTS.md, then ranks
the rest by score and relevance. The default balanced policy selects at most 12
items and 4,000 estimated tokens, with at most 6 items and 2,500 tokens from any
one source, so a single provider cannot crowd out the others. The selected
items are rendered into the system prompt.
Long grep output, logs, and child transcripts should be preserved outside the prompt and summarized into prompt-safe evidence. Use session.traceEvents() for compact runtime evidence.
Context providers (Rust Host)
A context provider implements the ContextProvider trait from
a3s_code_core::context. All providers registered on a session are queried
concurrently with the turn's prompt; empty results are dropped. A provider that
fails is skipped with a warning unless its failure_mode() returns
ContextProviderFailureMode::FailClosed, in which case the turn fails.
on_turn_complete is an optional hook called after each turn.
with_fs_context(root) registers the built-in FileSystemContextProvider.
The module also ships ripgrep, static, skill-catalog, and recent-workspace-file
providers. Node.js, Python, and Go do not expose custom context providers; use
AGENTS.md, skills, memory, or workspace retrieval from those SDKs.
Exact Cognitive Packages (Rust Host)
An embedding Rust host can bind a session to one exact A3S Use
cognitive-package generation. A3S Code does not install packages, resolve a
Registry entry, or select latest; the host injects both an immutable
CognitivePackageBindingV1 and a provider holding the matching Knowledge
lease.
The durable a3s.code.cognitive-package-session-binding.v1 identity includes
the package id/version, lifecycle generation, generation digest, capability
snapshot digest, exact Knowledge surface, and prompt-injection limits. Each
typed request and cited Markdown response repeats that binding and is validated
before content enters model context.
The hard limits are four documents, 6 KiB per document, and 6 KiB total. The host may choose smaller limits in the binding. Provider failure, malformed citations, request mismatch, or generation drift fails closed instead of falling back to unrelated retrieval.
Use with_cognitive_context; adding a cognitive provider through the generic
context-provider list is rejected because its binding could not be persisted
correctly. An exact cognitive package cannot be combined with any other
context provider (with_context_provider or with_fs_context), and both
session and durable memory recall are suppressed. Code-owned workspace instructions and skills remain
available.
The binding is stored in the session snapshot and emitted as
cognitive_context_bound. On resume, the host must re-inject a provider with
the same binding; a missing or different generation is rejected. This typed
boundary is currently a Rust-host integration surface rather than a Node.js,
Python, or Go session option.
Session-owned workspace retrieval
A3S Code can build a bounded retrieval index for one workspace and one session. Construction runs asynchronously, results are verified against the current source, and every vector is released when the session closes. The runtime does not install or require a vector database.
Workspace Retrieval is a host capability, not a switch the model can turn on. Omit the typed option to keep it disabled. A disabled session constructs no additional catalog, invokes no embedding provider, and exposes no semantic or hybrid search mode to the model.
Enabled sessions do not register a separate vector_db tool. Instead, the
single model-facing search schema adds mode: "semantic" and
mode: "hybrid". Both modes query the session-owned projection through the
same governed tool path as grep, glob, and BM25; disabled sessions omit them
from the schema rather than advertising calls that cannot run.
Choose the smallest useful retrieval surface
Dense semantic search necessarily needs a text-to-vector function, but that function can be an in-process CPU callback. A3S Code does not require a remote API, GPU, bundled model, or runtime model download. Model revision, license, artifact verification, caching, and credentials remain the host's responsibility.
Asynchronous vector projection lifecycle
The semantic serving projection is a session-owned, exact A3S Memory vector index, not a durable or shared vector database. Product builds use a separate session-local a3s-vec FTS collection for lexical ranking; its collection handles are short-lived and bounded. Session construction returns before corpus embedding finishes. The background indexer reads admitted text, publishes immutable per-file partitions atomically, and reconciles later source revisions without re-embedding unchanged files. Reopening a session builds new projections; closing it cancels outstanding provider work and releases all semantic and lexical state.
A session without Workspace Retrieval reports disabled. An enabled session
moves through building, ready, degraded, and closed. Queries can
use published coverage while construction continues. Hosts that need a
stronger first-query boundary can wait for readiness for up to 30 seconds;
timeout preserves the partial fallback, while cancellation or session close
interrupts the wait.
Backend and ranking boundary
a3s-vec owns lexical FTS/BM25 postings in product builds (engine id
a3s_vec_fts_v1, enabled by the default local-code feature through
a3s-vec-fts); a minimal build can
explicitly use the portable BM25 implementation. A lexical initialization or
query failure is reported as bounded lexical degradation and cannot change the
A3S Memory semantic authority. The SDKs expose typed lexical options and no
primitive backend-name selector. Temporary lexical collections are deleted on
normal close; process-crash residue remains subject to the host operating
system's temporary-directory policy.
Ranking
One immutable chunk catalog backs incremental BM25, Memory-authoritative exact
vectors, stable source anchors, and exact-literal or Code Intelligence
candidates. Hybrid mode combines independent one-based ranks with
reciprocal-rank fusion (k = 60) instead of mixing incomparable raw scores.
RRF-only is the default. The optional deterministic reranker is bounded,
model-free CPU code that reduces duplicate evidence while protecting exact
identifiers.
Text admission and chunking
Only manifest-admitted UTF-8 text and source files enter the catalog.
Generated files, oversized files, credentials, key material, .a3s control
paths, and non-text assets are excluded before chunking and embedding. PDF,
Office, image, audio, OCR, and other knowledge compilation belongs to a
separate knowledge compiler; Workspace Retrieval does not guess how to parse
those formats.
Built-in typed strategies cover line/byte chunks, fixed UTF-8 windows, and recursive separators. Trusted Rust hosts can supply a custom splitter whose ranges preserve UTF-8 boundaries, cover admitted bytes, and always make forward progress. Node.js, Python, and Go accept typed built-in strategy objects; primitive strategy names are rejected.
SDK control surface
CLI activation
The a3s CLI keeps semantic retrieval disabled unless a trusted user ACL or a
file selected explicitly with --config enables it. An automatically
discovered workspace .a3s/config.acl may only disable an inherited retrieval
route; it cannot authorize source egress or choose an embedding backend.
Remote embedding needs a separate provider route and an explicit source-egress grant:
Local CPU embedding is mutually exclusive with the remote fields and does not need a source-egress grant:
An explicit artifact_manifest is revision- and SHA-256-bound, and the CLI
loads it without downloading anything. When artifact_manifest is omitted, the
CLI uses the A3S Power-managed Xenova/all-MiniLM-L6-v2 bundle (384
dimensions, pinned revision and SHA-256 digests), installed on first use under
the A3S data root; offline mode or A3S_NO_AUTO_INSTALL disables that install
and requires the bundle to be present already. Run a3s config validate and inspect the redacted
workspaceRetrieval section from a3s config show before creating a session.
The embedding route is independent from default_model, so selecting DeepSeek
for chat and tool calls does not implicitly turn that chat endpoint into an
embedding service.
The CLI's local_cpu adapter is shipped for Linux x64/ARM64, Windows x64, and
Apple Silicon. Intel macOS 12 (x86_64) builds intentionally omit the optional
ONNX adapter. On Intel, leave retrieval model-free or use a separately
authorized remote embedding route.
Provider descriptors lock identity, model, dimension, and normalization. Runtime validation rejects partial, duplicate, unknown, dimension-mismatched, non-finite, non-normalized, or descriptor-drifted responses. Diagnostics do not copy input text, vectors, remote response bodies, credentials, or endpoint values.
Quality and safety evidence
Status snapshots report coverage, queue depth, failures, vector memory, batching, request amplification, non-text provider inputs, and post-close release. The cross-SDK real-model fixtures report task accuracy, Precision@5, Recall@5, MRR, nDCG@5, document request amplification, construction, readiness, turn and close latency, non-text provider inputs, and the post-close release rate. They are a portability gate, not a claim that one model or reranker is best for every repository.
Before rendering a result, A3S Code rereads the authoritative file and verifies the full-file digest and exact chunk byte range. Deleted, stale, unreadable, or superseded candidates are not exposed. See the operations runbook and qualification record for production thresholds, rollback rules, and reproducible evaluation.
Compaction
Automatic compaction is off by default. Enable it for long sessions:
The default threshold is 0.80. When maxContextTokens is omitted, Core uses
the selected model's declared context window when available. Before each model
request, Core accounts for the system prompt, conversation, tool calls and
results, and exposed tool schemas. At the threshold it:
- Prunes or truncates oversized older tool output, leaving a marker that tells the model to re-read the file or re-run the command.
- Summarizes the older prefix with the model, treating every transcript entry as untrusted data. The summary is capped at 8,000 tokens.
- Keeps up to the 20 most recent messages intact (half of a shorter history) and aims to land at 60% of the trigger watermark.
Compaction never loses the task goal. Before summarizing, Core pins the goal
outside the model: the most recent ## Goal section in a user message (for
example, from an earlier summary), otherwise the first user turn. If the new
summary omits or rewrites that section, Core puts the pinned ## Goal back.
Rolling compactions therefore keep the original paths and constraints even
when a later summary forgets them.
The summary participates in later compactions, so long sessions can roll
forward through repeated compression; this does not enlarge the model's
physical single-request context window. A successful context_compacted event
includes the cumulative summary so hosts that supply external history can
persist the same compact generation across turns.