验证
运行时将"完成"视为必须被证明 的事实,而不仅仅是被声明的结果。当模型说某个任务已完成时,这句话本身毫无价值。验证把声明转化为证据:你声明一组必须 成功的命令,运行时执行它们,结果会携带一份你可以检视、据以拦截或呈现给用户的报告。
验证是会话级的。Rust 核心运行每条命令,记录其退出状态和输出,并将每份报告汇总为一份随回合结果一同返回的摘要。
产品界面可以先展示交付摘要,再列出支持结论的命令输出和文件级改动。
运行验证命令
一条验证命令就是一个小而具名的检查:一个 id、一个 kind、一段可读的 description,以及要运行的 command。当某个失败应被视为硬失败而非警告时,将该检查标记为 required。
Rust 换行 复制 use a3s_code_core :: verification :: VerificationCommand ;
let commands = vec! [
VerificationCommand :: required (
" build ",
" build ",
" 项目可以编译 ",
" cargo build --all-features ",
)
. with_timeout_ms ( 120_000 ),
VerificationCommand :: required (
" tests ",
" test ",
" 单元测试通过 ",
" cargo test ",
),
];
let report = session
. verify_commands (" release-readiness ", & commands )
. await ? ;
println! ("{ report:#? }");
TypeScript 换行 复制 const report = await session . verifyCommands (' release-readiness ', [
{
id : ' build ',
kind : ' build ',
description : ' Project compiles ',
command : ' cargo build --all-features ',
required : true ,
timeoutMs : 120000 ,
},
{
id : ' tests ',
kind : ' test ',
description : ' Unit tests pass ',
command : ' cargo test ',
required : true ,
},
]);
console . log ( report );
Python 换行 复制 report = session . verify_commands (' release-readiness ', [
{
" id ": " build ",
" kind ": " build ",
" description ": " Project compiles ",
" command ": " cargo build --all-features ",
" required ": True ,
" timeout_ms ": 120000 ,
},
{
" id ": " tests ",
" kind ": " test ",
" description ": " Unit tests pass ",
" command ": " cargo test ",
" required ": True ,
},
])
print ( report )
Go 换行 复制 report , err := session . VerifyCommands ( ctx , " release-readiness ", [] code . VerificationCommand {
{
ID : " build ",
Kind : " build ",
Description : " 项目可以编译 ",
Command : " cargo build --all-features ",
Required : true ,
TimeoutMS : code . Ptr ( uint64 ( 120000 )),
},
{
ID : " tests ",
Kind : " test ",
Description : " 单元测试通过 ",
Command : " cargo test ",
Required : true ,
},
})
if err != nil {
return err
}
fmt . Println ( report )
subject(此处为 release-readiness)为这一批检查命名,使同一会话内的多次验证在报告中保持彼此独立。
读取回合结束后的摘要
每个回合的 send() 结果同样携带只读的验证字段,因此你无需单独发起一次验证调用即可据结果拦截。用这些字段判断该回合是否真正完成了它所声称的工作。
Rust 换行 复制 let result = session . send (" 应用修复并运行检查 ", None ) . await ? ;
let summary = result . verification_summary ();
println! ("{ :? }", summary . status );
println! ("{}", summary . pending_required_check_count );
println! ("{}", summary . failed_check_count );
println! ("{}", summary . report_count );
println! ("{}", result . verification_summary_text ());
if summary . failed_check_count > 0 {
return Err ( a3s_code_core :: CodeError :: Session (
" 当前回合声称已完成,但验证失败 " . to_string (),
));
}
TypeScript 换行 复制 const result = await session . send (' Apply the fix and run the checks ');
console . log ( result . verificationStatus );
console . log ( result . pendingVerificationCount );
console . log ( result . failedVerificationCount );
console . log ( result . verificationReportCount );
console . log ( result . verificationSummaryText );
if ( result . failedVerificationCount > 0 ) {
throw new Error (' Turn reported done but verification failed ');
}
Python 换行 复制 result = session . send (' Apply the fix and run the checks ')
print ( result . verification_status )
print ( result . pending_verification_count )
print ( result . failed_verification_count )
print ( result . verification_report_count )
print ( result . verification_summary_text )
if result . failed_verification_count > 0 :
raise RuntimeError (' Turn reported done but verification failed ')
Go 换行 复制 result , err := session . Run ( ctx , " 应用修复并运行检查 ")
if err != nil {
return err
}
summary := result . VerificationSummary
fmt . Println ( summary . Status )
fmt . Println ( summary . PendingRequiredCheckCount )
fmt . Println ( summary . FailedCheckCount )
fmt . Println ( summary . ReportCount )
fmt . Println ( result . VerificationSummaryText )
if summary . FailedCheckCount > 0 {
return errors . New (" 回合声称完成,但验证失败 ")
}
检视报告与摘要
除了逐回合的字段之外,会话还暴露完整的报告集合、一份结构化摘要、可用的预设,以及一份可读的概要。该概要是向人展示某回合为何 通过或失败的最快方式。
Rust 换行 复制 let reports = session . verification_reports ();
let summary = session . verification_summary ();
let presets = session . verification_presets ();
let text = session . verification_summary_text ();
println! (
"{} 份报告,状态 { :? } , {} 个预设 ",
reports . len (),
summary . status ,
presets . len ()
);
println! ("{ text }");
TypeScript 换行 复制 import { formatVerificationSummary } from ' @a3s-lab/code ';
const reports = session . verificationReports ();
const summary = session . verificationSummary ();
const presets = session . verificationPresets ();
// 会话辅助方法和独立格式化函数都能生成可读文本。
console . log ( session . verificationSummaryText ());
console . log ( formatVerificationSummary ( summary ));
Python 换行 复制 reports = session . verification_reports ()
summary = session . verification_summary ()
presets = session . verification_presets ()
# 会话辅助方法返回可直接打印的可读摘要。
print ( session . verification_summary_text ())
Go 换行 复制 reports , err := session . VerificationReports ( ctx )
if err != nil {
return err
}
summary , err := session . VerificationSummary ( ctx )
if err != nil {
return err
}
presets , err := session . VerificationPresets ( ctx )
if err != nil {
return err
}
text , err := session . VerificationSummaryText ( ctx )
if err != nil {
return err
}
fmt . Println ( len ( reports ), summary . Status , len ( presets ))
fmt . Println ( text )
verificationPresets() 返回根据 workspace 文件推断出的检查模板,例如
Cargo.toml、package.json、pyproject.toml 和 go.mod。请把它们当作起点:
在用来拦截发布或用户可见自动化之前,应审查命令、超时和 required 标记是否适合该项目。
A3S Code 自身如何完成资格认证
回合验证回答的是一次智能体任务是否产生了它所声称的结果。仓库资格认证回答的是另一
个问题,即每项公开的 A3S Code 能力是否仍然在 Core、各语言 SDK、资源上限和受支持的
部署面上满足合同。仅有一次绿色编译无法回答这个问题。
仓库把证据分成四类:
能力台账把 README 宣称的 27 个产品领域逐项连接到可执行证据,并明确显示任何尚未解决
的缺口。能力地图发生变化而台账没有同步时,CI 会直接失败。Node.js 和 Python 门禁会先
构建并加载 Native Module,再通过公开 Wrapper 执行测试;仅通过 Rust cargo check
不算 SDK Runtime 证据。Go 会通过带版本的 Bridge 并使用 Race Detector 运行。
性能检查会区分工作放大和耗时。普通 CI 拦截 Provider Request、Vector Byte、Scratch
Space、Retry 与关闭后留存等确定性上限。专用 Release Profile 工作流会预热并重复采样,
输出 Machine-readable JSON,并将其作为 CI Artifact 保留。依赖网络的 DeepSeek 与浏览器
耗时会单独报告,因为把这些时间混入本地执行后,就无法区分 Code Regression 与 Provider
或网络波动。
v8.3.0 的真实模型矩阵会运行资格 ACL 声明的每个模型,验证模型选择的工具与 Hook
参数改写、有证据门禁的多文件编码任务、自动与显式 SubAgents、Skill 发现和执行、
并发 PTC 读取、持久 A3S Flow 回放,以及公开 steer/interrupt。确定性 Fixture 还会
独立覆盖相同控制路径,包括过期与幂等回执、拒绝、预算、取消和清理。
最新受控资格认证
2026-08-18 的 Release Profile 在四个逻辑 CPU 的 x86-64 Linux Runner 上运行,六份报告
全部通过:
Provider Request Amplification、Vector 与 Rerank Byte、RSS Delta、Serialized Graph Byte、
Process Cleanup、无累积覆盖写入与删除清理等资源门禁也全部通过。配套 Hermetic CI 完成了
MinIO Roundtrip、受控 HTTPS 上的生产 Chrome/CDP 与 Google Parser 路径,以及本地
OpenTelemetry Collector 对指定 Service/Span 的精确接收。
性能资格认证记录
包含 p50/p95/Max、精确包含规则、机器信息、资源计数、Workflow 链接与 Artifact SHA-256
Digest。这些数据是锁定 Profile 的回归上限,不是所有硬件或远程服务的 SLA。
当前证据台账、缺口关闭情况、外部边界、资格认证命令和完成规则见
能力验证与性能合同 。
只要台账中仍有未解决的 Code-owned Gap,就不能声称仓库级目标已经完成。
为何重要
没有验证,一次智能体运行止于模型的一面之词。有了验证,运行止于可观测的证据:编译通过的构建、跑通的测试套件、保持沉默的代码检查器。摘要文本提供审计轨迹;结果上的计数让你能在自动化中以失败为默认(fail closed)。
相关
遥测 —— 将追踪事件和验证报告作为运行时证据进行检视。
限制 —— 在验证运行之前限定一个回合能完成多少工作量。