Verification
The harness treats "done" as something that must be proven, not merely claimed. When the model says a task is complete, that assertion is worth nothing on its own. Verification turns the claim into evidence: you declare commands that must succeed, the runtime executes them, and the result carries a report you can inspect, gate on, or surface to a user.
Verification is session-scoped. The Rust core runs each command, records its exit status and output, and rolls every report up into a single summary that travels alongside the turn result.
Product UIs can present a delivery summary first, followed by the command output and file-level changes used to support it.
Running Verification Commands
A verification command is a small, named check: an id, a kind, a
human-readable description, and the command to run. Mark a check required
when a failure should be treated as a hard failure rather than a warning.
The subject (here release-readiness) labels the batch so multiple
verification passes within one session stay distinct in the reports.
Reading The Post-Turn Summary
Every turn's send() result also carries read-only verification fields, so you
can gate on the outcome without issuing a separate verification call. Use these
to decide whether the turn actually accomplished what it claimed.
Inspecting Reports And Summaries
Beyond the per-turn fields, the session exposes the full set of reports, a structured summary, the available presets, and a human-readable digest. The digest is the quickest way to show a person why a turn passed or failed.
verificationPresets() returns workspace-aware check templates inferred from
files such as Cargo.toml, package.json, pyproject.toml, and go.mod.
Treat them as starting points: review the commands, timeouts, and required
flags for the project before gating releases or user-visible automation.
How A3S Code Itself Is Qualified
Turn verification answers whether one agent task produced its claimed result. Repository qualification answers a different question: whether every public A3S Code capability still satisfies its contract across Core, SDKs, resource limits, and supported deployment surfaces. A green build alone cannot answer that question.
The repository therefore separates four evidence classes:
The capability ledger maps all 27 advertised product areas to executable evidence and keeps any unresolved gap visible. CI rejects a capability-map change that is not reflected in that ledger. Node.js and Python gates build and load their native modules before exercising the public wrappers; compile-only Rust checks are not counted as SDK runtime evidence. Go runs through its versioned bridge with the race detector.
Performance checks distinguish work amplification from timing. Ordinary CI gates deterministic ceilings such as provider requests, vector bytes, scratch space, retries, and post-close retention. The dedicated release-profile workflow uses warmups and repeated samples, emits machine-readable JSON, and retains it as a CI artifact. Network-dependent DeepSeek and browser timings are reported separately, because combining them with local execution would make a regression indistinguishable from provider or network variance.
The v8.2.0 live-model matrix uses every model declared by the qualification ACL. It checks model-selected Tool calls and Hook rewrites, an evidence-gated multi-file coding task, automatic and explicit SubAgents, Skill discovery and execution, concurrent PTC reads, persisted A3S Flow replay, and public steer/interrupt behavior. Deterministic fixtures independently cover the same control paths, including stale and idempotent receipts, denial, budgets, cancellation, and cleanup.
Latest controlled qualification
The 2026-08-18 release-profile run on four logical x86-64 Linux CPUs passed all six reports:
Resource gates also passed for request amplification, vector and rerank bytes, RSS deltas, serialized graph bytes, process cleanup, overwrite without accumulation, and delete cleanup. Companion hermetic CI completed a MinIO roundtrip, the production Chrome/CDP and Google-parser path against controlled HTTPS, and exact service/span receipt by a local OpenTelemetry Collector.
See the performance qualification record for p50/p95/max values, inclusion rules, machine metadata, resource counts, workflow links, and Artifact SHA-256 digests. These are regression ceilings for the locked profiles, not universal hardware or remote-service SLAs.
See the capability verification and performance contract for the current evidence ledger, gap closure, external boundaries, qualification commands, and the completion rule. A repository-wide claim is not complete while that ledger contains an unresolved Code-owned gap.
Why This Matters
Without verification, an agent run ends on the model's word. With it, the run ends on observable evidence: a build that compiled, a test suite that passed, a linter that stayed quiet. The summary text gives you the audit trail; the counts on the result let you fail closed in automation.