Context
A3S Code treats context as a budgeted resource. The model should see the smallest useful context for the current decision, not every available file, skill, memory, and tool log.
Sources
Context can come from:
- user prompt and conversation history
- AGENTS.md project instructions
- skills and agent definitions
- memory stores
- file search and direct tool results
- MCP tools and context providers
- delegated task summaries
- trace events
AGENTS.md injection, skills discovery, memory APIs, direct tool results,
delegated task helpers, and traceEvents() are part of the documented Node SDK
surface. Validate MCP context behavior against your own live integrations
before documenting it as product behavior.
Assembly
Long grep output, logs, and child transcripts should be preserved outside the prompt and summarized into prompt-safe evidence. Use session.traceEvents() for compact runtime evidence.
Exact Cognitive Packages (Rust Host)
An embedding Rust host can bind a session to one exact A3S Use
cognitive-package generation. A3S Code does not install packages, resolve a
Registry entry, or select latest; the host injects both an immutable
CognitivePackageBindingV1 and a provider holding the matching Knowledge
lease.
The durable a3s.code.cognitive-package-session-binding.v1 identity includes
the package id/version, lifecycle generation, generation digest, capability
snapshot digest, exact Knowledge surface, and prompt-injection limits. Each
typed request and cited Markdown response repeats that binding and is validated
before content enters model context.
The hard limits are four documents, 6 KiB per document, and 6 KiB total. The host may choose smaller limits in the binding. Provider failure, malformed citations, request mismatch, or generation drift fails closed instead of falling back to unrelated retrieval.
Use with_cognitive_context; adding a cognitive provider through the generic
context-provider list is rejected because its binding could not be persisted
correctly. An exact cognitive package cannot accompany general-purpose RAG or
graph providers, and personal-memory recall is suppressed. Code-owned
workspace instructions and skills remain available.
The binding is stored in the session snapshot and emitted as
cognitive_context_bound. On resume, the host must re-inject a provider with
the same binding; a missing or different generation is rejected. This typed
boundary is currently a Rust-host integration surface rather than a Node.js,
Python, or Go session option.
Compaction
Enable automatic compaction for long sessions:
When maxContextTokens is omitted, Core uses the selected model's declared
context window when available. Before each model request, Core accounts for the
system prompt, conversation, tool calls and results, and exposed tool schemas.
At the configured threshold it bounds oversized tool output, summarizes the
older safe prefix, keeps recent messages, and continues the same task. The
summary participates in later compactions, so long sessions can roll forward
through repeated compression; this does not enlarge the model's physical
single-request context window. A successful context_compacted event includes
the cumulative summary so hosts that supply external history can persist the
same compact generation across turns.
Python exposes the same override as max_context_tokens.