Verification

The harness treats "done" as something that must be proven, not merely claimed. When the model says a task is complete, that assertion is worth nothing on its own. Verification turns the claim into evidence: you declare commands that must succeed, the runtime executes them, and the result carries a report you can inspect, gate on, or surface to a user.

Verification is session-scoped. The Rust core runs each command, records its exit status and output, and rolls every report up into a single summary that travels alongside the turn result.

Product UIs can present a delivery summary first, followed by the command output and file-level changes used to support it.

Running Verification Commands

A verification command is a small, named check: an id, a kind, a human-readable description, and the command to run. Mark a check required when a failure should be treated as a hard failure rather than a warning.

Rust
Node.js
Python
Go
Rust
use a3s_code_core::verification::VerificationCommand;
let commands = vec![
VerificationCommand::required(
"build",
"build",
"Project compiles",
"cargo build --all-features",
)
.with_timeout_ms(120_000),
VerificationCommand::required(
"tests",
"test",
"Unit tests pass",
"cargo test",
),
];
let report = session
.verify_commands("release-readiness", &commands)
.await?;
println!("{report:#?}");

The subject (here release-readiness) labels the batch so multiple verification passes within one session stay distinct in the reports.

Reading The Post-Turn Summary

Every turn's send() result also carries read-only verification fields, so you can gate on the outcome without issuing a separate verification call. Use these to decide whether the turn actually accomplished what it claimed.

Rust
Node.js
Python
Go
Rust
let result = session
.send("Apply the fix and run the checks", None)
.await?;
let summary = result.verification_summary();
println!("{:?}", summary.status);
println!("{}", summary.pending_required_check_count);
println!("{}", summary.failed_check_count);
println!("{}", summary.report_count);
println!("{}", result.verification_summary_text());
if summary.failed_check_count > 0 {
return Err(a3s_code_core::CodeError::Session(
"turn reported done but verification failed".to_string(),
));
}

Inspecting Reports And Summaries

Beyond the per-turn fields, the session exposes the full set of reports, a structured summary, the available presets, and a human-readable digest. The digest is the quickest way to show a person why a turn passed or failed.

Rust
Node.js
Python
Go
Rust
let reports = session.verification_reports();
let summary = session.verification_summary();
let presets = session.verification_presets();
let text = session.verification_summary_text();
println!(
"{} reports, status {:?}, {} presets",
reports.len(),
summary.status,
presets.len()
);
println!("{text}");

verificationPresets() returns workspace-aware check templates inferred from files such as Cargo.toml, package.json, pyproject.toml, and go.mod. Treat them as starting points: review the commands, timeouts, and required flags for the project before gating releases or user-visible automation.

How A3S Code Itself Is Qualified

Turn verification answers whether one agent task produced its claimed result. Repository qualification answers a different question: whether every public A3S Code capability still satisfies its contract across Core, SDKs, resource limits, and supported deployment surfaces. A green build alone cannot answer that question.

The repository therefore separates four evidence classes:

Evidence classWhat it provesWhat it does not prove
Deterministic correctnessActivation, successful behavior, invalid input, permissions, cancellation, lifecycle, ordering, and cleanup against fixed oraclesReal provider, browser, or object-store availability
Deterministic resource gatesBounds on calls, retries, records, bytes, queues, candidates, tool rounds, and retained stateWall-clock latency on every machine
Release performance qualificationRelease-build p50, p95, maximum, resource accounting, workload parameters, and machine metadata for stable local workRemote model or public-search latency
External qualificationCompatibility with a named live model, browser, collector, or storage service under recorded conditionsHermetic reproducibility or a universal performance claim

The capability ledger maps all 26 advertised product areas to executable evidence and keeps any unresolved gap visible. CI rejects a capability-map change that is not reflected in that ledger. Node.js and Python gates build and load their native modules before exercising the public wrappers; compile-only Rust checks are not counted as SDK runtime evidence. Go runs through its versioned bridge with the race detector.

Performance checks distinguish work amplification from timing. Ordinary CI gates deterministic ceilings such as provider requests, vector bytes, scratch space, retries, and post-close retention. The dedicated release-profile workflow uses warmups and repeated samples, emits machine-readable JSON, and retains it as a CI artifact. Network-dependent DeepSeek and browser timings are reported separately, because combining them with local execution would make a regression indistinguishable from provider or network variance.

Latest controlled qualification

The 2026-08-18 release-profile run on four logical x86-64 Linux CPUs passed all six reports:

ProfileFixed workloadObserved p95Objective
Agent convergenceFour completion, guard, and recovery cases4/4 cases4/4 cases
Workspace Retrieval25,000 × 384 exact / deterministic hybrid15.590 / 38.506 ms≤ 30 / 100 ms
Flow / State Graph1,000-step projection / 11,008-record replay130.067 / 125.526 mseach ≤ 2,000 ms
Code Intelligence5,000 files; cold / warm workspace symbols754.397 ms cold / 0.519 ms warm≤ 5,000 / 250 ms
Context / memory25,000 context inputs / 2,500 memories136.740 / 0.123 ms≤ 500 / 250 ms
File persistence1,272,624-byte synchronized save / load338.887 / 1.028 ms≤ 1,000 / 500 ms

Resource gates also passed for request amplification, vector and rerank bytes, RSS deltas, serialized graph bytes, process cleanup, overwrite without accumulation, and delete cleanup. Companion hermetic CI completed a MinIO roundtrip, the production Chrome/CDP and Google-parser path against controlled HTTPS, and exact service/span receipt by a local OpenTelemetry Collector.

See the performance qualification record for p50/p95/max values, inclusion rules, machine metadata, resource counts, workflow links, and Artifact SHA-256 digests. These are regression ceilings for the locked profiles, not universal hardware or remote-service SLAs.

See the capability verification and performance contract for the current evidence ledger, gap closure, external boundaries, qualification commands, and the completion rule. A repository-wide claim is not complete while that ledger contains an unresolved Code-owned gap.

Why This Matters

Without verification, an agent run ends on the model's word. With it, the run ends on observable evidence: a build that compiled, a test suite that passed, a linter that stayed quiet. The summary text gives you the audit trail; the counts on the result let you fail closed in automation.

  • Telemetry — inspect trace events and verification reports as runtime evidence.
  • Limits — bound how much work a turn can do before verification runs.