Are you an LLM? View https://a3s-lab.github.io/Code/llms.txt for optimized Markdown documentation, or https://a3s-lab.github.io/Code/llms-full.txt for full documentation bundle. This page is also available as Markdown at https://a3s-lab.github.io/Code/v7.0.1/en/guide/verification.md
The harness treats "done" as something that must be proven, not merely
claimed. When the model says a task is complete, that assertion is worth nothing
on its own. Verification turns the claim into evidence: you declare commands
that must succeed, the runtime executes them, and the result carries a report
you can inspect, gate on, or surface to a user.
Verification is session-scoped. The Rust core runs each command, records its
exit status and output, and rolls every report up into a single summary that
travels alongside the turn result.
Product UIs can present a delivery summary first, followed by the command output
and file-level changes used to support it.
A verification command is a small, named check: an id, a kind, a
human-readable description, and the command to run. Mark a check required
when a failure should be treated as a hard failure rather than a warning.
Every turn's send() result also carries read-only verification fields, so you
can gate on the outcome without issuing a separate verification call. Use these
to decide whether the turn actually accomplished what it claimed.
Beyond the per-turn fields, the session exposes the full set of reports, a
structured summary, the available presets, and a human-readable digest. The
digest is the quickest way to show a person why a turn passed or failed.
verificationPresets() returns workspace-aware check templates inferred from
files such as Cargo.toml, package.json, pyproject.toml, and go.mod.
Treat them as starting points: review the commands, timeouts, and required
flags for the project before gating releases or user-visible automation.
Turn verification answers whether one agent task produced its claimed result.
Repository qualification answers a different question: whether every public
A3S Code capability still satisfies its contract across Core, SDKs, resource
limits, and supported deployment surfaces. A green build alone cannot answer
that question.
The repository therefore separates four evidence classes:
Evidence class
What it proves
What it does not prove
Deterministic correctness
Activation, successful behavior, invalid input, permissions, cancellation, lifecycle, ordering, and cleanup against fixed oracles
Real provider, browser, or object-store availability
Deterministic resource gates
Bounds on calls, retries, records, bytes, queues, candidates, tool rounds, and retained state
Wall-clock latency on every machine
Release performance qualification
Release-build p50, p95, maximum, resource accounting, workload parameters, and machine metadata for stable local work
Remote model or public-search latency
External qualification
Compatibility with a named live model, browser, collector, or storage service under recorded conditions
Hermetic reproducibility or a universal performance claim
The capability ledger maps all 20 advertised product areas to executable
evidence and keeps any unresolved gap visible. CI rejects a capability-map
change that is not reflected in that ledger. Node.js and Python gates build and
load their native modules before exercising the public wrappers; compile-only
Rust checks are not counted as SDK runtime evidence. Go runs through its
versioned bridge with the race detector.
Performance checks distinguish work amplification from timing. Ordinary CI
gates deterministic ceilings such as provider requests, vector bytes, scratch
space, retries, and post-close retention. The dedicated release-profile
workflow uses warmups and repeated samples, emits machine-readable JSON, and
retains it as a CI artifact. Network-dependent DeepSeek and browser timings are
reported separately, because combining them with local execution would make a
regression indistinguishable from provider or network variance.
The 2026-08-18 release-profile run on four logical x86-64 Linux CPUs passed all
six reports:
Profile
Fixed workload
Observed p95
Objective
Agent convergence
Four completion, guard, and recovery cases
4/4 cases
4/4 cases
Workspace Retrieval
25,000 × 384 exact / deterministic hybrid
15.590 / 38.506 ms
≤ 30 / 100 ms
Flow / State Graph
1,000-step projection / 11,008-record replay
130.067 / 125.526 ms
each ≤ 2,000 ms
Code Intelligence
5,000 files; cold / warm workspace symbols
754.397 ms cold / 0.519 ms warm
≤ 5,000 / 250 ms
Context / memory
25,000 context inputs / 2,500 memories
136.740 / 0.123 ms
≤ 500 / 250 ms
File persistence
1,272,624-byte synchronized save / load
338.887 / 1.028 ms
≤ 1,000 / 500 ms
Resource gates also passed for request amplification, vector and rerank bytes,
RSS deltas, serialized graph bytes, process cleanup, overwrite without
accumulation, and delete cleanup. Companion hermetic CI completed a MinIO
roundtrip, the production Chrome/CDP and Google-parser path against controlled
HTTPS, and exact service/span receipt by a local OpenTelemetry Collector.
See the
performance qualification record
for p50/p95/max values, inclusion rules, machine metadata, resource counts,
workflow links, and Artifact SHA-256 digests. These are regression ceilings for
the locked profiles, not universal hardware or remote-service SLAs.
See the
capability verification and performance contract
for the current evidence ledger, gap closure, external boundaries,
qualification commands, and the completion rule. A repository-wide claim is
not complete while that ledger contains an unresolved Code-owned gap.
Without verification, an agent run ends on the model's word. With it, the run
ends on observable evidence: a build that compiled, a test suite that passed, a
linter that stayed quiet. The summary text gives you the audit trail; the
counts on the result let you fail closed in automation.