Tools
A3S Code keeps a registry of available tools. toolNames() returns the current
session tool surface. That surface is assembled from the workspace capability
set plus session-level integrations, so a non-local workspace can intentionally
hide tools it cannot service.
Tool activity reaches clients through the session event stream, so a UI can render starts, output, errors, and completion without parsing terminal text.
Tool Surface
Use toolNames() / toolDefinitions() in tests or application diagnostics when
a workflow depends on a specific model-visible tool. Tool visibility is not a
security grant. Model-selected tool calls inside send, run, and stream
pass through active-skill restrictions, permission policy, confirmation, hooks,
budget, queue/timeouts, cancellation, recursive-call protection, output
sanitization, artifact limits, and workspace path checks. Direct SDK calls such
as session.tool(...) are host control-plane calls with a different explicit
policy, described below.
Use activeTools() for a different question: which tool calls are currently
running in an active operation.
One Invocation Kernel
The runtime attaches an origin to every invocation:
Model sub-runs created by a host-direct Skill, Task, or custom tool start again
with the Agent origin. A public custom tool can make nested calls only through
InvocationRuntime, which always creates the governed nested origin. This
prevents one direct call to arbitrary extension code from becoming a reusable
authority token.
All origins use the same invocation kernel. Pre-hooks can block them; budget checks run before side effects; queue/timeout, cancellation, recursive invocation protection, post-hooks, and security-provider output sanitization remain active. A nested call cannot fall back to a raw registry while a scoped invoker is installed.
The host-direct policy is not an end-user authorization system. The embedding application must authenticate and authorize a user before translating their request into a direct SDK helper. Closing the session cancels an in-flight host-direct tool through the session cancellation scope.
Bounded Tool Contracts
Governed agent, nested, and session calls validate arguments against each tool's cached JSON Schema before confirmation or side effects. Tools also publish per-invocation scheduling capabilities: read-only, idempotent, resumable, cancellation-safe, pagination support, maximum parallelism, and output kind.
read, ls, paginated search modes, Git log/list/diff, and web_fetch return explicit
continuation cursors or offsets. Shell capture is byte-bounded across both
streams, retains the beginning and end with exact accounting, and still
exposes stdout and stderr separately to a sandbox host. A command deadline
includes both output draining and child.wait(), so closing both pipes cannot
evade it. Timeout and cancellation terminate the complete Unix process group.
Large file changes replace full before/after event payloads with bounded
previews and unified diff text plus hashes, sizes, and artifact references.
Every Tool result includes metadata.a3s_tool_result_evidence with schema
a3s.code.tool-result-evidence.v1. It records original and projected byte
counts, deterministic token estimates, an exact SHA-256 repeat key, the loss
mode, and an immutable inline or artifact content reference. These values are
Harness observations rather than provider billing records, and this evidence
does not itself transform the Tool result.
Deterministic Tool-result projection
Each session pins one a3s.code.tool-result-transform-policy.v1 policy. The
default conservative policy retains the first 100 KiB and does not fold or
sample content. The context-efficient profile keeps a UTF-8-safe 64 KiB head
and 32 KiB tail, folds three or more exact repeated lines, and samples up to 32
items from an oversized top-level JSON array.
Projection is deterministic: Core samples an oversized JSON array first,
folds exact repeated lines next, and only then applies the UTF-8-safe head/tail
bound if the result is still too large. max_output_bytes may be 1-100 KiB;
non-compatibility profiles must leave 512 bytes inside that bound for the
transformation marker.
The policy is stored in the session snapshot. Resume inherits it when omitted and rejects an explicitly different policy, so replay cannot silently change what the model observed.
Every result records original_bytes, projected_bytes, original/projected
token estimates using utf8-bytes-ceil-div-4/v1, source and projection
digests, byte/token deltas, repeat_key, content_ref, and
transform_algorithm. loss_mode is one of none, bounded_preview,
head_tail, deterministic_transform, or composite. Lossless results use an
inline SHA-256 reference; lossy results retain the complete original under an
immutable a3s://tool-output/... artifact URI.
Binary-safe local downloads
The download tool writes an HTTP(S) resource into a writable local workspace.
It is not registered for S3, browser, or other non-local backends. Model-selected
calls are workspace mutations, so the normal permission policy and HITL path
apply; a direct session.tool(...) call is a privileged host decision.
Before each redirect hop, the shared safe HTTP transport validates the target
against SSRF-sensitive addresses. Direct connections reject mixed
public/private DNS answers and pin the validated addresses for that hop.
Redirects are bounded and revalidated; cross-origin hops do not inherit
credentials or If-Range. An explicitly configured proxy resolves hostnames,
while URL-level literal-address checks still run.
The downloader probes Range support, requires exact Content-Range and body
boundaries, and uses a stable validator before parallelizing independent ranges.
It retries a bounded set of transport, rate-limit, and server failures, then
falls back to one sequential response when a parallel protocol or network
attempt is unsafe. Data streams into an adjacent temporary file. Cancellation,
timeouts, size failures, and checksum failures remove that partial file; only a
fully synced and optionally verified file is atomically promoted.
Result metadata contains the workspace path, byte count, content type, strategy, connection count, Range support, overwrite status, safe source anchors, and the verified digest when requested. Signed query parameters are used for requests but removed from source anchors and diagnostics, so they are not leaked through tool metadata.
Repository-context modes
Use read.files when the relevant paths are already known and one bounded
response is cheaper than several tool turns:
The shared byte budget includes headers and the continuation text. Results
stay in request order, and one unreadable member does not discard successful
members. When metadata.batch.truncated is true, copy
metadata.batch.continuation into the next call's files value; its offsets
and remaining limits resume without repeating completed lines.
search is the single model-facing workspace search tool. Always pass mode
and query: use grep for regular-expression content search, glob for path
discovery, and bm25 for native lexical relevance ranking. These modes are
separate planes: grep never opens durable zvec FTS, and bm25 never becomes
the exact-match authority. When the host
explicitly enables Workspace Retrieval, the same schema also advertises
semantic for exact cosine ranking over the session-owned, Memory-authoritative
index and
hybrid for reciprocal-rank fusion across exact, lexical, symbol, and semantic
evidence. Disabled sessions omit those two modes. The shared path field
scopes all modes; include filters candidate files for grep, BM25, semantic,
and hybrid search.
In mode: "grep", output_mode selects the smallest useful result shape:
Built-in workspace backends perform non-content scans without constructing
discarded match text. With default local-code, manifest-backed workspaces
also build a lazy in-tree trigram candidate cache under .a3s-code/grep-trigram
so literal needles open fewer files before the exact regex scan. Non-literal
patterns and index failures fail open to today's full scan. S3 results set
metadata.search.truncated and warn when
the backend's object scan limit makes totals or paths incomplete.
In mode: "glob", the query is the glob pattern. Backend relevance or
recency order is preserved by default; set sort: "path" before cursor
pagination when stable lexical pages are required.
In mode: "bm25", the query is plain text. A bounded native Rust scorer
splits code identifiers and CJK text, ranks 80-line chunks, and returns only
top-k snippets plus source anchors. It uses workspace search to narrow
candidates and keeps at most 256 files, 512 KiB per file, and 16 MiB in memory;
no database, embedding model, or external reranker is required.
Semantic and hybrid calls use the same tool rather than a separate vector database tool:
Index construction is asynchronous and file-atomic. Results are reread and digest-verified against current source before rendering, and closing the session releases every vector. See Workspace Retrieval for activation, partial-readiness behavior, embedding routes, resource bounds, and lifecycle.
For exact-string changes, call edit with dry_run: true to return the same
before/after diff metadata without writing. Then apply the edit with
expected_replacements set to the previewed count, and optionally add
max_replacements as an independent upper bound. Dry runs advertise read-only
capability and are safe for batch parallelization.
batch accepts at most 32 calls and applies at most 16-way concurrency. It
fans out only when every child declares safe read-only, idempotent behavior;
mutating and unknown tools are serialized. A partial batch identifies failed
indices while treating the orchestration as completed, so callers retry only
failed items. Multi-item task fan-out has the same 32-task bound and settles
cancelled children before publishing terminal state.
The model-facing task schema always uses tasks with 1-32 items. One item runs
a focused child and may use background; multiple independent items run
concurrently and cannot set background: true. Each item accepts agent,
description, prompt, optional max_steps, and optional output_schema.
min_success_count is valid only with allow_partial_failure: true and must be
between 1 and the submitted item count. Provider and child runtimes own typed
retry policy; the fan-out layer never replays a branch based on error text.
Structurally gated web search
web_search reports complete, partial, or failed in metadata. Its default
path executes headless engines first, conventional HTTP/RSS engines only when
the combined structural retrieval requirements are not met, and native APIs
only when both earlier tiers remain insufficient. Browser discovery and pool creation are
therefore lazy. A3S Code v8.5.1 pins a3s-search v3.1.0 and uses the packaged
or shared-cache Moli runtime by default. Minimal Rust embeddings can disable
default features; Chrome and Lightpanda remain explicit configured backends.
Moli resolution checks an explicit executable, the package sidecar, the verified
per-user cache, and a discoverable system installation before attempting an
HTTPS download of the pinned release. The cache is protected by a cross-process
install lock and atomic receipts, so multiple a3s-code processes reuse one
installation. Set auto_download_moli = false for an offline/strict deployment.
On Linux musl, the release emits MOLI_UNAVAILABLE because upstream Moli has no
musl asset; provide a system/explicit executable or choose another backend.
The final cascade fails closed when it still does not satisfy the structural
retrieval requirements. Successful JSON output keeps the result-array contract.
Insufficient JSON output is an error with a typed
retrieval_requirements_not_met envelope, the candidate rows, observed
retrieval health, and the required structural thresholds. Search does not use
an external semantic verifier or reranker.
Session-scoped closed/open/half-open circuit state skips known quota,
permission, rate-limit, transport, repeated-empty, and timeout failures without
retrying them on every request; Retry-After is retained. Delegated research
contexts share search bulkheads, bounded browser retry budgets, and
identical-request coalescing. Request-scoped proxies reach the lazy browser
tier, and search_coalescing metadata reports leader, shared, bypassed, and
abandoned requests. An explicit engines argument runs only the requested
tiers. Empty results with engine errors are failures rather than successful
empty searches. Timeout, cancellation, invalid-argument, partial-failure, and
rate-limit errors carry structured error kinds, while tier decisions, retrieval
health, engine outcomes, attempt duration, and retry context remain available
in metadata.
web_fetch also preserves failure semantics instead of deriving retry advice
from rendered messages. Request and response-body I/O failures, HTTP 408, and
HTTP 5xx responses use the typed transport kind. HTTP 429 uses
rate_limited and carries a parsed Retry-After delay when the server supplies
one; the outer tool deadline uses timeout. Other HTTP status failures remain
ordinary status errors unless the runtime has typed evidence that retrying is
safe.
Structured Output with generate_object
The generate_object tool asks the configured LLM for a JSON value, validates
the response against a JSON Schema, and returns the validated value only on a
zero-exit result. Root objects, arrays, enums, constants, composition keywords,
and local $ref definitions are supported. The active timeout starts after
model-generation admission, is forwarded to clients that accept an active
transport budget, and remains in force across bounded schema repairs. It works
in two ways:
- Agent-driven: The LLM sees
generate_objectin its tool list and calls it autonomously when structured output is needed. - Direct call: Your application calls
session.tool('generate_object', ...)to bypass model-driven tool selection. The tool itself still calls the configured LLM.
Parameters
Modes
- tool: Forces a synthetic tool whose parameters are the schema when the provider supports forced tool calls.
- prompt: Appends schema instructions to the prompt. It is useful as a prompt-only fallback, but depends more on model compliance.
- auto: Selects forced-tool mode when available, otherwise prompt mode.
- strict: Uses provider-native strict JSON Schema when supported, otherwise safely falls back to forced-tool or prompt mode.
- json: Uses provider-native JSON-object mode when supported, otherwise safely falls back to forced-tool or prompt mode.
Every resolved mode retains the provider-facing response schema as host-only
validation metadata; that metadata is never serialized as an extra provider
request field. Rust composite clients can use
structured::is_complete_streamed_value(...) to accept only a complete JSON
value that validates against this schema, including endpoints that omit a
terminal stream event. They can also inspect
LlmClient::has_distinct_non_streaming_transport() before treating a blocking
call as an independent fallback instead of replaying the same streaming
failure mode under another method name.
Streaming
When called through session.stream(), partial objects are emitted as
tool_output_delta events. Snapshots are rate-limited to one event per 100 ms;
objects over the event budget emit byte accounting instead of duplicating the
full value in the stream:
Repair Loop
If the LLM output fails schema validation, the tool automatically retries by feeding the validation errors back to the model. This handles edge cases like missing required fields or wrong enum values without application-level retry logic.
Direct Tool Calls
SDK callers can call deterministic tools directly:
Direct calls execute inside the session workspace and should be treated as privileged host operations. They do not claim the session's single-flight conversation lease because they do not update transcript history.
For delegated child work, use SDK helpers over the same core tools:
Automatic subagent delegation also uses these core tools. autoParallel: false
disables only automatic parallel fan-out; it does not remove manual task
fan-out or session.tasks(...).
Programmatic Tool Calling
PTC is the next step beyond a single direct call. The program tool runs a sandboxed JavaScript script in an embedded QuickJS VM. The script defines async function run(ctx, inputs) and replaces repeated model-tool turns with one bounded program.
Instead of spending LLM turns on:
the model can ask program to run a script:
Run it through the SDK helper with either inline source or a workspace-relative .js or .mjs file path:
session.program(...) is equivalent to session.tool('program', { type: 'script', language: 'javascript', ... }) but uses SDK-native naming. When allowedTools / allowed_tools is omitted, the script can call every registered tool except program. Provide an allow-list when a workflow should run with a smaller capability surface.
ctx.readFile(path, options) returns the selected text without the read
Tool's line anchors or continuation footer, which makes it suitable for string
processing. Use ctx.read(path, options) when the script needs the complete
Tool result, including line-numbered output, exit code, metadata, and
continuation evidence. Both forms invoke the same governed read Tool and
consume the same script call budget.
The QuickJS VM receives no filesystem, network, subprocess, or environment permissions. The only useful capabilities are the ctx methods wired back to A3S Code tools. PTC returns a readable ToolResult.output and structured data in ToolResult.metadataJson. Keep large raw output out of the prompt; summarize findings, evidence refs, risks, and suggested next actions.
In a3s code, PTC is also used by DynamicWorkflowRuntime. Recursive
program, dynamic_workflow, and the removed parallel_task alias are kept out
of the default TUI PTC allow-list. QuickJS may call task with one item, but
direct multi-item fan-out is blocked. A dynamic workflow schedules a Flow step
named task for local fan-out; the TUI host executes it outside QuickJS. After
/login,
dynamic workflow PTC steps may call ctx.tool("runtime", ...) when a host
runtime tool is registered in the session.
Authorizing a model-selected dynamic_workflow authorizes its private QuickJS
execution engine; callers do not need a second program(*) rule for that
implementation detail. Every Tool named in allowed_tools still re-enters the
normal permission, confirmation, Hook, budget, sandbox, and cancellation path.
Dynamic workflows default structured generation to single-flight. Set
limits.maxConcurrentGenerations to 2-4 only when independent
generate_object steps should fan out; each admitted step receives a client
fork bound to its exact run and step identity. Providers that cannot fork a
session remain single-flight. The Rust helper
dynamic_workflow::recover_dynamic_workflow_step_output(...) can recover a
completed durable step only when the run id, original query, and step id all
match. It does not act as a cross-run query cache and never promotes an
incomplete step.
Tool Results
ToolResult includes name, output, exitCode, and optional metadataJson.
Direct session.tool(...), session.program(...), session.git(...),
session.writeFile(...), session.ls(...), session.editFile(...), and
session.patchFile(...) calls return ToolResult. The typed read/search/shell
helpers return simpler values: readFile, grep, and bash return strings,
while glob returns a string array. Long outputs should be summarized before
they are fed back to the model.
Verification
Use verification commands to turn "done" into evidence: