Limits

Limit options are session-level controls for long-running work, noisy tools, and provider failures.

Session Options

TypeScript
const session = agent.session('/repo', {
maxToolRounds: 24,
maxParseRetries: 3,
toolTimeoutMs: 120000,
circuitBreakerThreshold: 4,
autoCompact: true,
autoCompactThreshold: 0.75,
continuationEnabled: true,
maxContinuationTurns: 3,
});

Python uses the same names in snake_case on SessionOptions (for example opts.max_tool_rounds = 24). Go uses pointer fields such as MaxToolRounds: code.Ptr(uint(24)) and ToolTimeoutMS: code.Ptr(uint64(120000)).

Options and Defaults

Node.js optionDefaultIntent
maxToolRounds50Tool-iteration budget for one turn.
maxParseRetries2Malformed tool-call recovery budget.
toolTimeoutMsnonePer-tool timeout in milliseconds.
llmApiTimeoutMsprovider clientPer-model HTTP timeout in milliseconds.
circuitBreakerThreshold3Attempts for one model call before the turn fails on a retryable provider error.
duplicateToolCallThreshold3Identical tool calls among the last 8 before the duplicate-call guard refuses the call.
autoCompactfalseEnables automatic context compaction.
autoCompactThreshold0.80Fraction of the context window that triggers compaction (context window default 200000).
continuationEnabledtrueInjects a continuation when the model stops early.
maxContinuationTurns3Maximum continuation injections per execution.
maxExecutionTimeMsnoneAborts the execution loop after this duration.

When a turn reaches maxToolRounds, the runtime does not inject a finalization message. It sends the next completion with an empty tool list so the model must answer in text; see the fact_log_tool_round_cap_uses_an_empty_tool_list test in core/src/fact_control.rs. A turn that changed the workspace still has to pass the completion gate.

The circuit breaker retries only retryable provider errors; when the session streams events, a failed call is retried at most once. Budget denials and non-retryable errors fail immediately. The duplicate-call guard does not run a refused call: the model receives a failed tool result explaining the repeat. If the model repeats the same refused call again, the turn fails.

Practical Defaults

Use strict limits for CI, release, and user-facing automation. Use larger budgets for exploratory local coding sessions, but keep verification commands explicit and required when the task has side effects.

Retention Limits

A session keeps its run history, trace events, and subagent task snapshots in memory. SessionRetentionLimits applies conservative finite defaults so a long-lived session cannot grow those stores without bound. Override any individual FIFO cap, or explicitly set unbounded: true to opt into unlimited retention.

Five independent caps:

FieldEffect when capped
max_runs_retainedWhen a new run pushes past the cap, the oldest run and all of its events are dropped.
max_events_per_runThe oldest events in a run are FIFO-dropped. The run snapshot's event_count is not decremented — it stays the cumulative total ever recorded.
max_event_bytes_per_runOldest run events are dropped until the serialized-byte cap is met; an oversized single event is not retained.
max_trace_eventsThe oldest event in the trace sink is dropped on each new write past the cap.
max_terminal_subagent_tasksThe oldest terminal (completed / failed / cancelled) subagent task snapshot is dropped past the cap. Running tasks are never dropped.

The defaults are 64 runs, 2048 events per run, 8 MiB of event data per run, 8192 trace events, and 512 terminal subagent tasks. All caps are soft: enforcement drops the oldest entry on insert and never returns an error.

TypeScript
const session = agent.session('/repo', {
retentionLimits: {
maxRunsRetained: 100,
maxEventsPerRun: 5000,
maxEventBytesPerRun: 16 * 1024 * 1024,
maxTraceEvents: 20000,
maxTerminalSubagentTasks: 500,
},
});
Python
opts = SessionOptions()
opts.retention_limits = {
'max_runs_retained': 100,
'max_events_per_run': 5000,
'max_event_bytes_per_run': 16 * 1024 * 1024,
'max_trace_events': 20000,
'max_terminal_subagent_tasks': 500,
}
session = agent.session('/repo', opts)
Go
session, err := agent.Session(ctx, "/repo", &code.SessionOptions{
RetentionLimits: &code.RetentionLimits{
MaxRunsRetained: code.Ptr(uint(100)),
MaxEventsPerRun: code.Ptr(uint(5_000)),
MaxEventBytesPerRun: code.Ptr(uint(16 * 1024 * 1024)),
MaxTraceEvents: code.Ptr(uint(20_000)),
MaxTerminalSubagentTasks: code.Ptr(uint(500)),
},
})

Rust uses the corresponding SessionRetentionLimits builder methods.

Budget Guard

BudgetGuard (core/src/budget.rs) is a host-supplied cost / quota contract. The framework does not enforce budgets itself — it defines the decision points and consults a guard the host plugs in. Three hooks are wired at the LLM / tool call site:

  • check_before_llm — before each LLM call.
  • record_after_llm — after each successful LLM call, with the actual provider usage, so the host keeps its running spend total accurate.
  • check_before_tool — before each tool call.

Each check_* returns one of three decisions:

  • Allow — proceed normally, no event.
  • SoftLimit { resource, consumed, limit, message } — emits an AgentEvent::BudgetThresholdHit { kind: "soft" } and proceeds. In-session hooks can react (auto-compact, swap to a cheaper model next turn).
  • Deny { resource, reason } — emits BudgetThresholdHit { kind: "hard" } and refuses the call. A denied LLM call fails the turn with CodeError::BudgetExhausted and is not retried by the circuit breaker; a denied tool call is not run and the model receives a denied tool result. The session stays open — the caller can retry later or after the host re-allocates budget.

Node — session.setBudgetGuard({...})

Each callback takes a single ctx object (not positional arguments) and returns a decision dict (or null / { decision: 'allow' } to allow):

TypeScript
session.setBudgetGuard({
checkBeforeLlm: (ctx) => {
// ctx.sessionId, ctx.estimatedTokens
if (overMonthlyCap(ctx.sessionId)) {
return {
decision: 'deny',
resource: 'llm_tokens',
reason: 'monthly cap',
};
}
return { decision: 'allow' };
},
recordAfterLlm: (ctx) => {
// ctx.sessionId, ctx.usage — usage keys are camelCase:
// promptTokens, completionTokens, totalTokens, cacheReadTokens, cacheWriteTokens
addSpend(ctx.sessionId, ctx.usage.totalTokens);
},
checkBeforeTool: (ctx) => {
// ctx.sessionId, ctx.toolName
return { decision: 'allow' };
},
timeoutMs: 5000, // optional, default 5000
});

The Node bridge fails closed (sdk/node/src/workflow_budget.rs): a check* callback that throws, returns something unreadable (including a Promise), or does not return within timeoutMs is treated as a deny. A budget control never silently disables itself when the guard fails or stalls. Callbacks must be synchronous. A failure inside recordAfterLlm is ignored.

Python — budget_guard session option

Python supplies a BudgetGuard-shaped object on the budget_guard SessionOptions field, bounded by budget_guard_timeout_ms (default 5000, must be greater than zero). session.set_budget_guard(guard, timeout_ms) replaces it on a live session; pass None to clear it. Methods that aren't defined behave as Allow / no-op. Python callbacks use positional arguments. An exception, a malformed decision, or a timeout in a check_* method fails closed as a deny (sdk/python/src/orchestration_bridge.rs):

Python
class MyBudgetGuard:
def check_before_llm(self, session_id, est_tokens):
if over_monthly_cap(session_id):
return {'decision': 'deny', 'resource': 'llm_tokens', 'reason': 'monthly cap'}
return {'decision': 'allow'}
def record_after_llm(self, session_id, usage):
# usage is a dict with snake_case keys:
# total_tokens, cache_read_tokens (plus prompt_tokens, completion_tokens, cache_write_tokens)
add_spend(session_id, usage['total_tokens'])
def check_before_tool(self, session_id, tool_name):
return {'decision': 'allow'}
opts = SessionOptions()
opts.budget_guard = MyBudgetGuard()
session = agent.session('/repo', opts)

Go — session.SetBudgetGuard(ctx, handlers)

Go callbacks receive typed contexts and return a *code.BudgetDecision. Timeout defaults to 5 seconds. A nil handler allows; a returned error, a timeout, or an unreadable decision denies. Pass nil handlers to clear the guard.

Go
err := session.SetBudgetGuard(ctx, &code.BudgetGuardHandlers{
CheckBeforeLLM: func(ctx context.Context, value code.BudgetLLMContext) (*code.BudgetDecision, error) {
if overMonthlyCap(value.SessionID) {
return &code.BudgetDecision{Decision: "deny", Resource: "llm_tokens", Reason: "monthly cap"}, nil
}
return &code.BudgetDecision{Decision: "allow"}, nil
},
RecordAfterLLM: func(ctx context.Context, value code.BudgetUsageContext) error {
addSpend(value.SessionID, value.Usage.TotalTokens)
return nil
},
CheckBeforeTool: func(ctx context.Context, value code.BudgetToolContext) (*code.BudgetDecision, error) {
return &code.BudgetDecision{Decision: "allow"}, nil
},
Timeout: 5 * time.Second,
})

The decision shape {"decision": "deny", "resource": ..., "reason": ...} (and "soft" with resource, consumed, limit, message, or "allow") is the same across the Node.js, Python, and Go SDKs.