Verification

The harness treats "done" as something that must be proven, not merely claimed. A turn that changed the workspace cannot complete on assistant text. It completes only when a Passed verification report is bound to the exact mutation effect digest of that turn, or when the host supplied a waiver for that digest. Neither the report binding nor the waiver can be produced by the model's prose.

Verification is session-scoped. The Rust core runs each command, records its exit status and output, and rolls every report up into a single summary that travels alongside the turn result.

Product UIs can present a delivery summary first, followed by the command output and file-level changes used to support it.

The Completion Gate

Every turn ends at the completion gate (core/src/harness_loop.rs). The gate inspects the turn's mutation ledger, the verification reports collected during the turn, host waivers, and any external observations bound to the session. It decides in this order:

  1. An external observation that still requires a workspace change keeps the turn open. This is the only case that may continue the turn: once per turn, and only when continuation is enabled and the run is not being forced to finish. Otherwise the turn is incomplete.
  2. A background workspace child task whose writes were not observed makes the turn incomplete.
  3. Incomplete workspace observation makes the turn incomplete. A waiver or a report over a partial path list cannot close it.
  4. An empty ledger (no workspace mutation) completes as Narrative.
  5. A host waiver whose effect_digest equals the ledger digest completes as Waived.
  6. A report bound to the ledger digest with every required check Passed completes as Verified.
  7. Anything else is incomplete.

In the last case the turn fails with this message, which names the digest:

Text
completion gate: workspace mutation <digest> has no bound Passed verification
and no host waiver. Assistant text does not count. ...

send() and stream() surface it as an error. The mutation-gate path does not spend a continuation turn: another model turn cannot close the gate without new tool evidence. A successful Rust AgentResult exposes the outcome in completion:

CompletionTerminalMeaning
NarrativeNo workspace mutation; a final answer is enough.
Verified { effect_digest }Required checks passed and were bound to the mutation digest.
Waived { effect_digest }A host waiver covered the mutation digest. This is not a pass.
DistinctPlan steps closed different digests; no combined digest is produced.

Mutation effect digest

The ledger records, per turn:

  • successful write, edit, patch, and download calls, with a SHA-256 content digest of the written content;
  • changed_paths reported by tool metadata (for example from bash), even when the command failed;
  • effects of nested tool calls;
  • workspace paths changed outside any recorded tool, found by comparing the workspace before and after the turn (.a3s and .git are excluded).

The effect digest is a SHA-256 over the recorded tool|path|content_digest lines. A later write to the same path changes the digest, so a report bound to an earlier state no longer matches.

What binds a Passed report

A report satisfies the gate only when all of these hold:

  • effect_digest equals the current ledger digest;
  • it contains at least one required check, and every required check is passed;
  • the report status is neither failed nor needs_review.

Report authorship is checked before a report is collected. Tool metadata may carry verification_report plus verification_author; reports authored as editor are rejected, while host and verifier reports are accepted. Built-in bash produces evidence in two ways:

  • A command that covers a workspace preset command (see presets) or an active /goal acceptance criterion from .a3s/loops/*/ACCEPTANCE.md yields a shell:<id> report for that check. Such a report has no digest by itself.
  • When the same command also contains an existence check (test -f, test -e, [ -f ... ], [ -e ... ], [[ -f ... ]], or [[ -e ... ]]) that names a mutated path whose current on-disk SHA-256 equals the digest recorded in the ledger, the runtime binds unbound required-pass reports to the ledger digest. If the command exited 0 and no report was bound, a host report shell:mutation_path_verify:<path> is synthesized. An existence check without recorded content can bind an existing Passed report but never synthesizes one.

Reports recorded on the session with verifyCommands() or recordVerificationReports() carry no effect digest, and they do not seed a later turn's gate. Use them for host-side release checks and summaries; use a completion attestor when host evidence must close a live turn.

Host Waivers

CompletionWaiverV1 { effect_digest, reason } is a host- or user-confirmed exception for one exact digest. Both fields must be non-blank. The model cannot create one from prose, and Waived is reported separately from Verified. Waivers are fixed when the session is built, so they fit replay and resume flows in which the digest is already known.

Rust
use a3s_code_core::harness_loop::CompletionWaiverV1;
use a3s_code_core::SessionOptions;
let waiver = CompletionWaiverV1::new(digest, "Reviewed by the release owner")
.ok_or_else(|| a3s_code_core::CodeError::Config("waiver needs digest and reason".into()))?;
let options = SessionOptions::new().with_completion_waivers(vec![waiver]);
TypeScript
const session = agent.session('/repo', {
completionWaivers: [
{ effect_digest: digest, reason: 'Reviewed by the release owner' },
],
});
Go
session, err := agent.Session(ctx, "/repo", &code.SessionOptions{
CompletionWaivers: []code.CompletionWaiver{
{EffectDigest: digest, Reason: "Reviewed by the release owner"},
},
})

The Python SessionOptions has no public setter for completion waivers.

Completion Attestor

A host that runs third-party work cannot predict the digest when it builds the session. CompletionAttestor (core/src/completion_attestor.rs) closes that gap without adding a bypass:

  • It is called after the ledger digest exists and before the gate decides. It is skipped when the ledger is empty.
  • It receives the digest and one MutatedPathRecord { path, content_digest } per mutated path, so the host can re-read each file and compare it with the recorded content.
  • It may return a VerificationReport. The report is accepted as a host report and appended to the turn's reports. It must still bind: a wrong digest, a missing required check, or a non-Passed required check leaves the turn incomplete.
  • It is Rust-only (SessionOptions::with_completion_attestor), is not a tool, cannot be granted to the model, and has no observe-only mode that lets an unverified mutation complete.
Rust
use std::sync::Arc;
use a3s_code_core::verification::{VerificationCheck, VerificationReport, VerificationStatus};
use a3s_code_core::{CompletionAttestor, MutatedPathRecord, SessionOptions};
struct HostCiAttestor;
impl CompletionAttestor for HostCiAttestor {
fn attest(
&self,
effect_digest: &str,
paths: &[MutatedPathRecord],
) -> Option<VerificationReport> {
if !paths.iter().all(|record| host_ci_accepts(&record.path, &record.content_digest)) {
return None;
}
let check = VerificationCheck::required("host:ci", "test", "Host CI accepted the change")
.with_status(VerificationStatus::Passed);
Some(VerificationReport::new("host-ci", vec![check]).with_effect_digest(effect_digest))
}
}
fn host_ci_accepts(_path: &str, _content_digest: &str) -> bool {
// Host-specific: re-read the file, compare the digest, consult CI.
false
}
let options = SessionOptions::new().with_completion_attestor(Arc::new(HostCiAttestor));

Read-only Verifier

verifier_enabled is off by default. When it is on and a turn mutated the workspace, the runtime runs one extra verifier turn before the gate. That turn cannot present write, edit, patch, download, Skill, task, or batch. Bash commands containing >, rm , mv , or tee , and mutating git tool calls (checkout, creating a branch, stash with a message or untracked files, worktree create or remove), are refused. Its reports are accepted as verifier reports and still have to bind; a Failed or NeedsReview result stays a non-pass. The verifier runs at most once per turn.

Rust
Node.js
Python
Go
Rust
use a3s_code_core::SessionOptions;
let options = SessionOptions::new().with_verifier(true);

Running Verification Commands

A verification command is a small, named check: an id, a kind, a human-readable description, and the command to run. Mark a check required when a failure should be treated as a hard failure rather than a warning. verifyCommands() runs each command through bash as a host-controlled, escalated call with the given timeout and records the report on the session.

Rust
Node.js
Python
Go
Rust
use a3s_code_core::verification::VerificationCommand;
let commands = vec![
VerificationCommand::required(
"build",
"build",
"Project compiles",
"cargo build --all-features",
)
.with_timeout_ms(120_000),
VerificationCommand::required(
"tests",
"test",
"Unit tests pass",
"cargo test",
),
];
let report = session
.verify_commands("release-readiness", &commands)
.await?;
println!("{report:#?}");

The subject (here release-readiness) labels the batch so multiple verification passes within one session stay distinct in the reports. Reports use schema a3s.verification_report.v1; check statuses are passed, failed, needs_review, and skipped. Rust callers can also require a specific exit code with VerificationCommand::with_expect_exit.

Reading The Post-Turn Summary

Every turn's send() result also carries read-only verification fields, so you can gate on the outcome without issuing a separate verification call. Use these to decide whether the turn actually accomplished what it claimed.

Rust
Node.js
Python
Go
Rust
let result = session
.send("Apply the fix and run the checks", None)
.await?;
let summary = result.verification_summary();
println!("{:?}", result.completion);
println!("{:?}", summary.status);
println!("{}", summary.pending_required_check_count);
println!("{}", summary.failed_check_count);
println!("{}", summary.report_count);
println!("{}", result.verification_summary_text());
if summary.failed_check_count > 0 {
return Err(a3s_code_core::CodeError::Session(
"turn reported done but verification failed".to_string(),
));
}

A turn that the gate rejected does not reach these fields: the call returns the completion-gate error instead.

Inspecting Reports And Summaries

Beyond the per-turn fields, the session exposes the full set of reports, a structured summary, the available presets, and a human-readable digest. The digest is the quickest way to show a person why a turn passed or failed.

Rust
Node.js
Python
Go
Rust
let reports = session.verification_reports();
let summary = session.verification_summary();
let presets = session.verification_presets();
let text = session.verification_summary_text();
println!(
"{} reports, status {:?}, {} presets",
reports.len(),
summary.status,
presets.len()
);
println!("{text}");

The summary contains status, report_count, required_check_count, pending_required_check_count, failed_check_count, residual_risk_count, pending_subjects, and failed_subjects.

Verification presets

verificationPresets() returns workspace-aware check templates inferred from project files:

PresetDetected fromCommands
rust-defaultCargo.tomlcargo fmt -- --check, cargo check, cargo test (required); cargo clippy -- -D warnings (optional)
node-defaultpackage.json scripts test, typecheck, linttest required; typecheck and lint optional; run through the detected package manager (packageManager, lockfile, or npm)
python-defaultpyproject.toml, tests/, pytest.ini, Ruff or mypy settingspython -m pytest (required when tests are configured); python -m ruff check . and python -m mypy . (optional when configured)
go-defaultgo.modgo test ./... (required); go vet ./... (optional)

A bash command that covers one of these commands produces a shell:<id> report, as described above. Treat presets as starting points: review the commands, timeouts, and required flags for the project before gating releases or user-visible automation.

How A3S Code Itself Is Qualified

Turn verification answers whether one agent task produced its claimed result. Repository qualification answers a different question: whether every public A3S Code capability still satisfies its contract across Core, SDKs, resource limits, and supported deployment surfaces. A green build alone cannot answer that question.

The repository therefore separates four evidence classes:

Evidence classWhat it provesWhat it does not prove
Deterministic correctnessActivation, successful behavior, invalid input, permissions, cancellation, lifecycle, ordering, and cleanup against fixed oraclesReal provider, browser, or object-store availability
Deterministic resource gatesBounds on calls, retries, records, bytes, queues, candidates, tool rounds, and retained stateWall-clock latency on every machine
Release performance qualificationRelease-build p50, p95, maximum, resource accounting, workload parameters, and machine metadata for stable local workRemote model or public-search latency
External qualificationCompatibility with a named live model, browser, collector, or storage service under recorded conditionsHermetic reproducibility or a universal performance claim

The capability ledger (manual/CAPABILITY_VERIFICATION.md) maps every product area advertised in the README to executable evidence and keeps any unresolved gap visible. CI runs scripts/check_capability_verification.py and fails when an advertised capability has no ledger entry. The SDK runtime workflow builds and loads the Node.js and Python native modules before running their test suites; compile-only Rust checks are not counted as SDK runtime evidence. Go tests run through the bridge with the race detector.

Performance checks distinguish work amplification from timing. Ordinary CI gates deterministic ceilings such as provider requests, vector bytes, scratch space, retries, and post-close retention. The release-profile performance workflow writes one JSON report per profile and keeps the reports as a workflow artifact. A separate hermetic-integrations workflow exercises an S3-compatible object store fixture, a controlled Chrome CDP browser path, and receipt by a local OpenTelemetry Collector.

See the performance qualification record for recorded runs, inclusion rules, and machine metadata. Those values are regression ceilings for the locked profiles, not universal hardware or remote-service SLAs. See the capability verification and performance contract for the current evidence ledger, external boundaries, and qualification commands.

Why This Matters

Without verification, an agent run ends on the model's word. With it, the run ends on observable evidence: a build that compiled, a test suite that passed, a linter that stayed quiet. The summary text gives you the audit trail; the counts on the result let you fail closed in automation.

  • Telemetry — inspect trace events and verification reports as runtime evidence.
  • Limits — bound how much work a turn can do before verification runs.