运行限制
Limit options 是 session 级控制项,用于长任务、噪声工具输出和 provider
失败场景。
会话选项
TypeScript 换行 复制 const session = agent . session (' /repo ', {
maxToolRounds : 24 ,
maxParseRetries : 3 ,
toolTimeoutMs : 120000 ,
circuitBreakerThreshold : 4 ,
autoCompact : true ,
autoCompactThreshold : 0 . 75 ,
continuationEnabled : true ,
maxContinuationTurns : 3 ,
});
Python 在 SessionOptions 上使用相同名称的 snake_case 形式(例如
opts.max_tool_rounds = 24)。Go 使用指针字段,例如
MaxToolRounds: code.Ptr(uint(24)) 和 ToolTimeoutMS: code.Ptr(uint64(120000))。
选项与默认值
回合达到 maxToolRounds 时,运行时不会注入收尾消息,而是以空工具列表发送下一次
补全请求,迫使模型用文本回答;见 core/src/fact_control.rs 中的
fact_log_tool_round_cap_uses_an_empty_tool_list 测试。修改过 workspace 的回合仍然
必须通过完成门禁 。
熔断器只重试可重试的 provider 错误;session 以流式方式输出事件时,失败的调用最多
重试一次。预算拒绝和不可重试的错误会立即失败。重复调用保护不会执行被拒绝的调用:
模型会收到一个说明重复原因的失败工具结果;如果模型再次发出同一个被拒绝的调用,
回合失败。
实用默认值
CI、发布和用户可见自动化应使用严格限制;本地探索可以放宽预算,但有副作用的任务仍要显式验证。
保留上限
session 会把 run 历史、trace events 和 subagent task 快照保存在内存里。
SessionRetentionLimits 使用保守的有限默认值,避免长时间会话无限增长。可以覆盖
任意单项 FIFO 上限;只有明确设置 unbounded: true 才会启用无限保留。
五个相互独立的上限:
默认值为 64 个 run、每个 run 2048 条 event、每个 run 8 MiB event 数据、8192 条
trace event,以及 512 个终态 subagent task。所有上限都是软上限:触发时在插入处丢弃
最旧条目,绝不返回错误。
TypeScript 换行 复制 const session = agent . session (' /repo ', {
retentionLimits : {
maxRunsRetained : 100 ,
maxEventsPerRun : 5000 ,
maxEventBytesPerRun : 16 * 1024 * 1024 ,
maxTraceEvents : 20000 ,
maxTerminalSubagentTasks : 500 ,
},
});
Python 换行 复制 opts = SessionOptions ()
opts . retention_limits = {
' max_runs_retained ': 100 ,
' max_events_per_run ': 5000 ,
' max_event_bytes_per_run ': 16 * 1024 * 1024 ,
' max_trace_events ': 20000 ,
' max_terminal_subagent_tasks ': 500 ,
}
session = agent . session (' /repo ', opts )
Go 换行 复制 session , err := agent . Session ( ctx , " /repo ", & code . SessionOptions {
RetentionLimits : & code . RetentionLimits {
MaxRunsRetained : code . Ptr ( uint ( 100 )),
MaxEventsPerRun : code . Ptr ( uint ( 5 _ 000 )),
MaxEventBytesPerRun : code . Ptr ( uint ( 16 * 1024 * 1024 )),
MaxTraceEvents : code . Ptr ( uint ( 20 _ 000 )),
MaxTerminalSubagentTasks : code . Ptr ( uint ( 500 )),
},
})
Rust 使用对应的 SessionRetentionLimits builder 方法。
预算守卫
BudgetGuard(core/src/budget.rs)是一套由宿主提供的成本与配额契约。框架本身
不强制执行预算——它只定义决策点,并咨询宿主注入的守卫。在 LLM 与工具调用处接入
三个钩子:
check_before_llm —— 在每次 LLM 调用之前。
record_after_llm —— 在每次成功的 LLM 调用之后,带上服务提供商的实际用量,
以便宿主准确累计消费。
check_before_tool —— 在每次工具调用之前。
每个 check_* 返回三种决策之一:
Allow —— 正常继续,不发出事件。
SoftLimit { resource, consumed, limit, message } —— 发出
AgentEvent::BudgetThresholdHit { kind: "soft" } 并继续执行 。会话内的钩子
可以据此采取动作,例如自动压缩或下一轮换用更便宜的模型。
Deny { resource, reason } —— 发出 BudgetThresholdHit { kind: "hard" } 并拒绝
本次调用。被拒绝的 LLM 调用以 CodeError::BudgetExhausted 使回合失败,熔断器不会
重试;被拒绝的工具调用不会执行,模型会收到一个 denied 工具结果。
会话仍然保持打开 ——调用方可以稍后重试,或在宿主重新分配预算后重试。
Node.js:session.setBudgetGuard({...})
每个回调接收单个 ctx 对象 (不是位置参数),并返回一个决策对象(或 null /
{ decision: 'allow' } 表示放行):
TypeScript 换行 复制 session . setBudgetGuard ({
checkBeforeLlm : ( ctx ) => {
// ctx.sessionId, ctx.estimatedTokens
if ( overMonthlyCap ( ctx . sessionId )) {
return {
decision : ' deny ',
resource : ' llm_tokens ',
reason : ' monthly cap ',
};
}
return { decision : ' allow ' };
},
recordAfterLlm : ( ctx ) => {
// ctx.sessionId, ctx.usage —— usage 的 key 是 camelCase:
// promptTokens, completionTokens, totalTokens, cacheReadTokens, cacheWriteTokens
addSpend ( ctx . sessionId , ctx . usage . totalTokens );
},
checkBeforeTool : ( ctx ) => {
// ctx.sessionId, ctx.toolName
return { decision : ' allow ' };
},
timeoutMs : 5000 , // 可选,默认 5000
});
Node.js 桥接层采用失败即拒绝 策略(sdk/node/src/workflow_budget.rs):check*
回调抛出异常、返回无法解析的值(包括 Promise),或未在 timeoutMs 内返回,都会被
当作拒绝 处理。预算守卫失败或卡住时,预算控制绝不会悄悄自我失效。回调必须是同步的。
recordAfterLlm 内部的失败会被忽略。
Python:会话选项 budget_guard
Python 在 budget_guard 这个 SessionOptions 字段上提供一个 BudgetGuard 形态的
对象,并受 budget_guard_timeout_ms 约束(默认 5000,必须大于零)。
session.set_budget_guard(guard, timeout_ms) 可在运行中的 session 上替换它;传入
None 可清除。未定义的方法视为放行且不执行操作。Python 回调使用位置参数 。
check_* 方法中的异常、格式错误的决策或超时都会失败关闭为拒绝
(sdk/python/src/orchestration_bridge.rs):
Python 换行 复制 class MyBudgetGuard :
def check_before_llm ( self , session_id , est_tokens ):
if over_monthly_cap ( session_id ):
return {' decision ': ' deny ', ' resource ': ' llm_tokens ', ' reason ': ' monthly cap '}
return {' decision ': ' allow '}
def record_after_llm ( self , session_id , usage ):
# usage 是一个 dict,key 为 snake_case:
# total_tokens, cache_read_tokens(以及 prompt_tokens, completion_tokens, cache_write_tokens)
add_spend ( session_id , usage [' total_tokens '])
def check_before_tool ( self , session_id , tool_name ):
return {' decision ': ' allow '}
opts = SessionOptions ()
opts . budget_guard = MyBudgetGuard ()
session = agent . session (' /repo ', opts )
Go:session.SetBudgetGuard(ctx, handlers)
Go 回调接收类型化上下文,并返回 *code.BudgetDecision。Timeout 默认为 5 秒。
未设置的处理函数视为放行;返回错误、超时或无法解析的决策都会拒绝。传入 nil
处理函数可清除守卫。
Go 换行 复制 err := session . SetBudgetGuard ( ctx , & code . BudgetGuardHandlers {
CheckBeforeLLM : func ( ctx context . Context , value code . BudgetLLMContext ) ( * code . BudgetDecision , error ) {
if overMonthlyCap ( value . SessionID ) {
return & code . BudgetDecision { Decision : " deny ", Resource : " llm_tokens ", Reason : " monthly cap "}, nil
}
return & code . BudgetDecision { Decision : " allow "}, nil
},
RecordAfterLLM : func ( ctx context . Context , value code . BudgetUsageContext ) error {
addSpend ( value . SessionID , value . Usage . TotalTokens )
return nil
},
CheckBeforeTool : func ( ctx context . Context , value code . BudgetToolContext ) ( * code . BudgetDecision , error ) {
return & code . BudgetDecision { Decision : " allow "}, nil
},
Timeout : 5 * time . Second ,
})
决策结构 {"decision": "deny", "resource": ..., "reason": ...}(以及带 resource、
consumed、limit、message 的 "soft",或 "allow")在 Node.js、Python 和 Go
SDK 中相同。