Verification
The harness treats "done" as something that must be proven, not merely claimed. A turn that changed the workspace cannot complete on assistant text. It completes only when a Passed verification report is bound to the exact mutation effect digest of that turn, or when the host supplied a waiver for that digest. Neither the report binding nor the waiver can be produced by the model's prose.
Verification is session-scoped. The Rust core runs each command, records its exit status and output, and rolls every report up into a single summary that travels alongside the turn result.
Product UIs can present a delivery summary first, followed by the command output and file-level changes used to support it.
The Completion Gate
Every turn ends at the completion gate (core/src/harness_loop.rs). The gate
inspects the turn's mutation ledger, the verification reports collected during
the turn, host waivers, and any external observations bound to the session. It
decides in this order:
- An external observation that still requires a workspace change keeps the turn open. This is the only case that may continue the turn: once per turn, and only when continuation is enabled and the run is not being forced to finish. Otherwise the turn is incomplete.
- A background workspace child task whose writes were not observed makes the turn incomplete.
- Incomplete workspace observation makes the turn incomplete. A waiver or a report over a partial path list cannot close it.
- An empty ledger (no workspace mutation) completes as
Narrative. - A host waiver whose
effect_digestequals the ledger digest completes asWaived. - A report bound to the ledger digest with every required check Passed
completes as
Verified. - Anything else is incomplete.
In the last case the turn fails with this message, which names the digest:
send() and stream() surface it as an error. The mutation-gate path does not
spend a continuation turn: another model turn cannot close the gate without new
tool evidence. A successful Rust AgentResult exposes the outcome in
completion:
Mutation effect digest
The ledger records, per turn:
- successful
write,edit,patch, anddownloadcalls, with a SHA-256 content digest of the written content; changed_pathsreported by tool metadata (for example frombash), even when the command failed;- effects of nested tool calls;
- workspace paths changed outside any recorded tool, found by comparing the
workspace before and after the turn (
.a3sand.gitare excluded).
The effect digest is a SHA-256 over the recorded tool|path|content_digest
lines. A later write to the same path changes the digest, so a report bound to
an earlier state no longer matches.
What binds a Passed report
A report satisfies the gate only when all of these hold:
effect_digestequals the current ledger digest;- it contains at least one required check, and every required check is
passed; - the report status is neither
failednorneeds_review.
Report authorship is checked before a report is collected. Tool metadata may
carry verification_report plus verification_author; reports authored as
editor are rejected, while host and verifier reports are accepted. Built-in
bash produces evidence in two ways:
- A command that covers a workspace preset command (see
presets) or an active
/goalacceptance criterion from.a3s/loops/*/ACCEPTANCE.mdyields ashell:<id>report for that check. Such a report has no digest by itself. - When the same command also contains an existence check (
test -f,test -e,[ -f ... ],[ -e ... ],[[ -f ... ]], or[[ -e ... ]]) that names a mutated path whose current on-disk SHA-256 equals the digest recorded in the ledger, the runtime binds unbound required-pass reports to the ledger digest. If the command exited 0 and no report was bound, a host reportshell:mutation_path_verify:<path>is synthesized. An existence check without recorded content can bind an existing Passed report but never synthesizes one.
Reports recorded on the session with verifyCommands() or
recordVerificationReports() carry no effect digest, and they do not seed a
later turn's gate. Use them for host-side release checks and summaries; use a
completion attestor when host evidence must close a
live turn.
Host Waivers
CompletionWaiverV1 { effect_digest, reason } is a host- or user-confirmed
exception for one exact digest. Both fields must be non-blank. The model cannot
create one from prose, and Waived is reported separately from Verified.
Waivers are fixed when the session is built, so they fit replay and resume
flows in which the digest is already known.
The Python SessionOptions has no public setter for completion waivers.
Completion Attestor
A host that runs third-party work cannot predict the digest when it builds the
session. CompletionAttestor (core/src/completion_attestor.rs) closes that
gap without adding a bypass:
- It is called after the ledger digest exists and before the gate decides. It is skipped when the ledger is empty.
- It receives the digest and one
MutatedPathRecord { path, content_digest }per mutated path, so the host can re-read each file and compare it with the recorded content. - It may return a
VerificationReport. The report is accepted as a host report and appended to the turn's reports. It must still bind: a wrong digest, a missing required check, or a non-Passed required check leaves the turn incomplete. - It is Rust-only (
SessionOptions::with_completion_attestor), is not a tool, cannot be granted to the model, and has no observe-only mode that lets an unverified mutation complete.
Read-only Verifier
verifier_enabled is off by default. When it is on and a turn mutated the
workspace, the runtime runs one extra verifier turn before the gate. That turn
cannot present write, edit, patch, download, Skill, task, or
batch. Bash commands containing >, rm , mv , or tee , and mutating
git tool calls (checkout, creating a branch, stash with a message or
untracked files, worktree create or remove), are refused. Its reports are accepted as
verifier reports and still have to bind; a Failed or NeedsReview result stays a
non-pass. The verifier runs at most once per turn.
Running Verification Commands
A verification command is a small, named check: an id, a kind, a
human-readable description, and the command to run. Mark a check required
when a failure should be treated as a hard failure rather than a warning.
verifyCommands() runs each command through bash as a host-controlled,
escalated call with the given timeout and records the report on the session.
The subject (here release-readiness) labels the batch so multiple
verification passes within one session stay distinct in the reports. Reports
use schema a3s.verification_report.v1; check statuses are passed, failed,
needs_review, and skipped. Rust callers can also require a specific exit
code with VerificationCommand::with_expect_exit.
Reading The Post-Turn Summary
Every turn's send() result also carries read-only verification fields, so you
can gate on the outcome without issuing a separate verification call. Use these
to decide whether the turn actually accomplished what it claimed.
A turn that the gate rejected does not reach these fields: the call returns the completion-gate error instead.
Inspecting Reports And Summaries
Beyond the per-turn fields, the session exposes the full set of reports, a structured summary, the available presets, and a human-readable digest. The digest is the quickest way to show a person why a turn passed or failed.
The summary contains status, report_count, required_check_count,
pending_required_check_count, failed_check_count, residual_risk_count,
pending_subjects, and failed_subjects.
Verification presets
verificationPresets() returns workspace-aware check templates inferred from
project files:
A bash command that covers one of these commands produces a shell:<id>
report, as described above. Treat presets as starting points: review the
commands, timeouts, and required flags for the project before gating releases
or user-visible automation.
How A3S Code Itself Is Qualified
Turn verification answers whether one agent task produced its claimed result. Repository qualification answers a different question: whether every public A3S Code capability still satisfies its contract across Core, SDKs, resource limits, and supported deployment surfaces. A green build alone cannot answer that question.
The repository therefore separates four evidence classes:
The capability ledger (manual/CAPABILITY_VERIFICATION.md) maps every product
area advertised in the README to executable evidence and keeps any unresolved
gap visible. CI runs scripts/check_capability_verification.py and fails when
an advertised capability has no ledger entry. The SDK runtime workflow builds
and loads the Node.js and Python native modules before running their test
suites; compile-only Rust checks are not counted as SDK runtime evidence. Go
tests run through the bridge with the race detector.
Performance checks distinguish work amplification from timing. Ordinary CI gates deterministic ceilings such as provider requests, vector bytes, scratch space, retries, and post-close retention. The release-profile performance workflow writes one JSON report per profile and keeps the reports as a workflow artifact. A separate hermetic-integrations workflow exercises an S3-compatible object store fixture, a controlled Chrome CDP browser path, and receipt by a local OpenTelemetry Collector.
See the performance qualification record for recorded runs, inclusion rules, and machine metadata. Those values are regression ceilings for the locked profiles, not universal hardware or remote-service SLAs. See the capability verification and performance contract for the current evidence ledger, external boundaries, and qualification commands.
Why This Matters
Without verification, an agent run ends on the model's word. With it, the run ends on observable evidence: a build that compiled, a test suite that passed, a linter that stayed quiet. The summary text gives you the audit trail; the counts on the result let you fail closed in automation.