验证

Harness 把“完成”视为必须被证明的事情,而不是一句声明。修改过 workspace 的回合不能靠助手文本完成。只有当一份 Passed 验证报告绑定到该回合精确的 mutation effect digest,或者宿主为该 digest 提供了 waiver 时,回合才会完成。报告绑定和 waiver 都不能由模型的文字产生。

验证以 session 为作用域。Rust core 执行每条命令,记录退出状态与输出,并把所有报告 汇总成一份摘要,随回合结果一同返回。

产品 UI 可以先展示交付摘要,再展示支撑它的命令输出和文件级变更。

完成门禁

每个回合都止于完成门禁(core/src/harness_loop.rs)。门禁检查本回合的 mutation ledger、回合内收集的验证报告、宿主 waiver,以及绑定到 session 的外部观测,并按以下 顺序判定:

  1. 仍要求 workspace 变更的外部观测会让回合保持打开。这是唯一可以继续回合的情况: 每个回合最多一次,且仅在启用 continuation、运行未被强制结束时发生;否则回合不完整。
  2. 写入未被观测到的后台 workspace 子任务会让回合不完整。
  3. workspace 观测不完整时回合不完整。waiver 或只覆盖部分路径列表的报告都不能关闭它。
  4. ledger 为空(没有 workspace 变更)时以 Narrative 完成。
  5. effect_digest 等于 ledger digest 的宿主 waiver 以 Waived 完成。
  6. 绑定到 ledger digest 且所有 required check 都为 Passed 的报告以 Verified 完成。
  7. 其他情况均为不完整。

最后一种情况下回合失败,并返回包含 digest 的消息:

Text
completion gate: workspace mutation <digest> has no bound Passed verification
and no host waiver. Assistant text does not count. ...

send() 和 stream() 会把它作为错误返回。mutation 门禁路径不会消耗 continuation 回合:没有新的工具证据,再多一个模型回合也无法关闭门禁。成功的 Rust AgentResult 在 completion 中给出结果:

CompletionTerminal含义
Narrative没有 workspace 变更,最终回答即可。
Verified { effect_digest }required check 已通过,且绑定到 mutation digest。
Waived { effect_digest }宿主 waiver 覆盖了该 digest。这不是验证通过。
Distinct计划步骤关闭了不同的 digest;不会生成合并 digest。

Mutation effect digest

ledger 按回合记录:

  • 成功的 write、edit、patch 和 download 调用,以及写入内容的 SHA-256 content digest;
  • 工具元数据报告的 changed_paths(例如来自 bash),即使命令失败也会记录;
  • 嵌套工具调用的效果;
  • 不属于任何已记录工具、通过比较回合前后 workspace 发现的变更路径(排除 .a3s 和 .git)。

effect digest 是对已记录的 tool|path|content_digest 行计算的 SHA-256。之后再写入 同一路径会改变 digest,因此绑定到旧状态的报告不再匹配。

什么样的报告能绑定为 Passed

只有同时满足以下条件的报告才能通过门禁:

  • effect_digest 等于当前 ledger digest;
  • 至少包含一个 required check,并且每个 required check 都是 passed;
  • 报告状态既不是 failed 也不是 needs_review。

报告在收集前会检查作者身份。工具元数据可以携带 verification_report 和 verification_author;作者为 editor 的报告会被拒绝,宿主和 verifier 报告会被接受。 内置 bash 通过两种方式产生证据:

  • 覆盖某条 workspace preset 命令(见预设)或当前 /goal 验收标准 (来自 .a3s/loops/*/ACCEPTANCE.md)的命令,会为该检查生成 shell:<id> 报告。 这种报告本身不带 digest。
  • 如果同一条命令还包含存在性检查(test -f、test -e、[ -f ... ]、[ -e ... ]、 [[ -f ... ]] 或 [[ -e ... ]]),且其中指定的变更路径当前磁盘内容的 SHA-256 等于 ledger 记录的 digest,运行时会把尚未绑定的 required-pass 报告绑定到 ledger digest。如果命令以 0 退出且没有报告被绑定,会合成一份宿主报告 shell:mutation_path_verify:<path>。没有记录内容的存在性检查可以绑定已有的 Passed 报告,但永远不会合成报告。

通过 verifyCommands() 或 recordVerificationReports() 记录到 session 的报告不带 effect digest,也不会作为后续回合门禁的输入。它们适合宿主侧发布检查和摘要;若需要 用宿主证据关闭一个进行中的回合,请使用 completion attestor。

宿主 Waiver

CompletionWaiverV1 { effect_digest, reason } 是宿主或用户针对某个精确 digest 确认的例外。两个字段都不能为空。模型无法通过文字创建它,Waived 也与 Verified 分开报告。waiver 在构建 session 时固定,因此适用于 digest 已知的回放和恢复流程。

Rust
use a3s_code_core::harness_loop::CompletionWaiverV1;
use a3s_code_core::SessionOptions;
let waiver = CompletionWaiverV1::new(digest, "Reviewed by the release owner")
.ok_or_else(|| a3s_code_core::CodeError::Config("waiver needs digest and reason".into()))?;
let options = SessionOptions::new().with_completion_waivers(vec![waiver]);
TypeScript
const session = agent.session('/repo', {
completionWaivers: [
{ effect_digest: digest, reason: 'Reviewed by the release owner' },
],
});
Go
session, err := agent.Session(ctx, "/repo", &code.SessionOptions{
CompletionWaivers: []code.CompletionWaiver{
{EffectDigest: digest, Reason: "Reviewed by the release owner"},
},
})

Python SessionOptions 没有设置 completion waiver 的公开 setter。

Completion Attestor

运行第三方工作的宿主在构建 session 时无法预知 digest。CompletionAttestor (core/src/completion_attestor.rs)在不引入旁路的前提下解决这个问题:

  • 它在 ledger digest 产生之后、门禁判定之前被调用;ledger 为空时跳过。
  • 它接收 digest,以及每个变更路径对应的 MutatedPathRecord { path, content_digest }, 宿主可以重新读取文件并与记录的内容比较。
  • 它可以返回一份 VerificationReport。该报告作为宿主报告被接受并追加到本回合报告中, 但仍必须满足绑定条件:digest 错误、缺少 required check,或 required check 不是 Passed,都会让回合保持不完整。
  • 它只在 Rust 中提供(SessionOptions::with_completion_attestor),不是工具,不能 授予模型,也没有让未验证变更完成的仅观察模式。
Rust
use std::sync::Arc;
use a3s_code_core::verification::{VerificationCheck, VerificationReport, VerificationStatus};
use a3s_code_core::{CompletionAttestor, MutatedPathRecord, SessionOptions};
struct HostCiAttestor;
impl CompletionAttestor for HostCiAttestor {
fn attest(
&self,
effect_digest: &str,
paths: &[MutatedPathRecord],
) -> Option<VerificationReport> {
if !paths.iter().all(|record| host_ci_accepts(&record.path, &record.content_digest)) {
return None;
}
let check = VerificationCheck::required("host:ci", "test", "Host CI accepted the change")
.with_status(VerificationStatus::Passed);
Some(VerificationReport::new("host-ci", vec![check]).with_effect_digest(effect_digest))
}
}
fn host_ci_accepts(_path: &str, _content_digest: &str) -> bool {
// Host-specific: re-read the file, compare the digest, consult CI.
false
}
let options = SessionOptions::new().with_completion_attestor(Arc::new(HostCiAttestor));

只读 Verifier

verifier_enabled 默认关闭。开启后,如果回合修改了 workspace,运行时会在门禁之前 额外运行一个 verifier 回合。该回合不会呈现 write、edit、patch、download、 Skill、task 或 batch。包含 >、rm 、mv 或 tee 的 bash 命令,以及会修改 状态的 git 工具调用(checkout、创建分支、带消息或未跟踪文件的 stash、创建或删除 worktree)都会被拒绝。它的报告作为 verifier 报告被接受,但仍必须绑定;Failed 或 NeedsReview 结果仍不算通过。每个回合最多运行一次 verifier。

Rust
Node.js
Python
Go
Rust
use a3s_code_core::SessionOptions;
let options = SessionOptions::new().with_verifier(true);

执行验证命令

一条验证命令是一个小的、具名的检查:id、kind、可读的 description,以及要执行的 command。当失败应被视为硬性失败而非警告时,把检查标记为 required。 verifyCommands() 会以宿主控制的提权调用通过 bash 执行每条命令,应用给定超时, 并把报告记录到 session。

Rust
Node.js
Python
Go
Rust
use a3s_code_core::verification::VerificationCommand;
let commands = vec![
VerificationCommand::required(
"build",
"build",
"Project compiles",
"cargo build --all-features",
)
.with_timeout_ms(120_000),
VerificationCommand::required(
"tests",
"test",
"Unit tests pass",
"cargo test",
),
];
let report = session
.verify_commands("release-readiness", &commands)
.await?;
println!("{report:#?}");

subject(此处为 release-readiness)为这批检查打标签,使同一 session 内的多次验证 在报告中保持区分。报告使用 schema a3s.verification_report.v1;检查状态为 passed、failed、needs_review 和 skipped。Rust 调用方还可以用 VerificationCommand::with_expect_exit 要求特定退出码。

读取回合后摘要

每个回合的 send() 结果都携带只读的验证字段,因此无需单独发起验证调用即可基于结果 做判断。用这些字段确认回合是否真正完成了它声称的事情。

Rust
Node.js
Python
Go
Rust
let result = session
.send("Apply the fix and run the checks", None)
.await?;
let summary = result.verification_summary();
println!("{:?}", result.completion);
println!("{:?}", summary.status);
println!("{}", summary.pending_required_check_count);
println!("{}", summary.failed_check_count);
println!("{}", summary.report_count);
println!("{}", result.verification_summary_text());
if summary.failed_check_count > 0 {
return Err(a3s_code_core::CodeError::Session(
"turn reported done but verification failed".to_string(),
));
}

被门禁拒绝的回合不会返回这些字段:调用会直接返回完成门禁错误。

检视报告与摘要

除了每回合的字段,session 还暴露完整的报告集合、结构化摘要、可用预设,以及一段可读 的摘要文本。摘要文本是向人展示回合为何通过或失败的最快方式。

Rust
Node.js
Python
Go
Rust
let reports = session.verification_reports();
let summary = session.verification_summary();
let presets = session.verification_presets();
let text = session.verification_summary_text();
println!(
"{} reports, status {:?}, {} presets",
reports.len(),
summary.status,
presets.len()
);
println!("{text}");

摘要包含 status、report_count、required_check_count、 pending_required_check_count、failed_check_count、residual_risk_count、 pending_subjects 和 failed_subjects。

验证预设

verificationPresets() 返回根据项目文件推断出的检查模板:

预设检测依据命令
rust-defaultCargo.tomlcargo fmt -- --check、cargo check、cargo test(required);cargo clippy -- -D warnings(optional)
node-defaultpackage.json 中的 test、typecheck、lint 脚本test 为 required;typecheck 和 lint 为 optional;通过检测到的包管理器执行(packageManager、lockfile 或 npm)
python-defaultpyproject.toml、tests/、pytest.ini、Ruff 或 mypy 配置python -m pytest(配置了测试时为 required);python -m ruff check . 与 python -m mypy .(已配置时为 optional)
go-defaultgo.modgo test ./...(required);go vet ./...(optional)

覆盖其中某条命令的 bash 命令会生成 shell:<id> 报告,见上文。请把预设当作起点: 在用来拦截发布或用户可见自动化之前,应审查命令、超时和 required 标记是否适合该项目。

A3S Code 自身如何完成资格认证

回合验证回答的是一次智能体任务是否产生了它所声称的结果。仓库资格认证回答的是另一 个问题,即每项公开的 A3S Code 能力是否仍然在 Core、各语言 SDK、资源上限和受支持的 部署面上满足合同。仅有一次绿色编译无法回答这个问题。

仓库把证据分成四类:

证据类别能证明什么不能证明什么
确定性正确性用固定 Oracle 检查启用条件、成功行为、非法输入、权限、取消、生命周期、顺序和清理真实 Provider、浏览器或对象存储是否可用
确定性资源门禁限定调用、重试、记录、字节、队列、候选项、工具轮次与留存状态每台机器上的绝对耗时
Release 性能资格认证对稳定本地工作记录 Release Build 的 p50、p95、最大值、资源计量、负载参数和机器信息远程模型或公共搜索服务的延迟
外部资格认证在已记录条件下验证指定真实模型、浏览器、Collector 或存储服务的兼容性Hermetic 可复现性或普遍性能结论

能力台账(manual/CAPABILITY_VERIFICATION.md)把 README 宣称的每个产品领域逐项连接 到可执行证据,并明确显示任何尚未解决的缺口。CI 会运行 scripts/check_capability_verification.py,只要有宣称的能力缺少台账条目就会失败。SDK Runtime 工作流会先构建并加载 Node.js 和 Python Native Module,再运行它们的测试套件; 仅通过 Rust 编译检查不算 SDK Runtime 证据。Go 测试通过 Bridge 并使用 Race Detector 运行。

性能检查会区分工作放大和耗时。普通 CI 拦截 Provider Request、Vector Byte、Scratch Space、Retry 与关闭后留存等确定性上限。Release Profile 性能工作流为每个 Profile 输出一份 JSON 报告,并将这些报告保留为 Workflow Artifact。独立的 hermetic-integrations 工作流会 验证 S3 兼容对象存储 Fixture、受控 Chrome CDP 浏览器路径,以及本地 OpenTelemetry Collector 的接收。

已记录的运行、包含规则和机器信息见 性能资格认证记录。 这些数据是锁定 Profile 的回归上限,不是所有硬件或远程服务的 SLA。当前证据台账、外部边界 和资格认证命令见 能力验证与性能合同。

为何重要

没有验证,一次智能体运行止于模型的一面之词。有了验证,运行止于可观测的证据:编译通过的构建、跑通的测试套件、保持沉默的代码检查器。摘要文本提供审计轨迹;结果上的计数让你能在自动化中以失败为默认(fail closed)。

相关

  • 遥测 —— 将追踪事件和验证报告作为运行时证据进行检视。
  • 限制 —— 在验证运行之前限定一个回合能完成多少工作量。