Cluster Extension Points

A cluster host platform runs long-lived agent sessions across many nodes. The framework does not ship a scheduler or a placement engine. Instead it exposes a small set of seams: it defines the decision points, emits structured events, and lets the host supply the policy. Everything below is something you wire from outside the framework — you never fork it.

This page shows the equivalent native seams in Rust, Node.js, Python, and Go. The configuration shape follows each language's conventions, while the underlying capability and lifecycle contract stay the same.

Identity labels

Every session can carry four opaque identity labels. The framework does not interpret them: it exposes them through session getters, persists them in SessionData, and restores them on resume. This is how a host attributes a session to a tenant, a principal, an agent template, and a wider correlation chain. The one exception is a durable-memory binding: when a session has one, tenant_id and principal default to the binding's namespace, and explicit values must match it exactly.

Pair identity labels with a sessionStore / session_store so the labels survive a process restart. On resume, caller-supplied options win, so you can relabel a session as you move it between nodes.

Rust
Node.js
Python
Go
Rust
use a3s_code_core::{Agent, SessionOptions};
#[tokio::main]
async fn main() -> a3s_code_core::Result<()> {
let agent = Agent::new("agent.acl").await?;
let options = SessionOptions::new()
.with_tenant_id("tenant-example")
.with_principal("principal-example")
.with_agent_template_id("agent-template-example")
.with_correlation_id("trace-example")
.with_file_session_store("./.a3s/sessions");
let session = agent
.session_builder("/path/to/project")
.options(options)
.build()
.await?;
println!("tenant: {:?}", session.tenant_id());
println!("principal: {:?}", session.principal());
println!("template: {:?}", session.agent_template_id());
println!("correlation: {:?}", session.correlation_id());
session.close().await;
agent.close().await;
Ok(())
}

Budget / cost guard

A budget guard lets the host gate every LLM call against a cost or token budget. The framework calls your guard before each LLM request and after it returns. The guard is policy you own; the framework only enforces the decision you hand back.

Rust
Node.js
Python
Go

The native guard decision shape is equivalent across all four SDKs:

Return valueEffect
None / null / { decision: 'allow' }Proceed with the LLM call.
{ decision: 'soft', resource, consumed, limit, message? }Emits BudgetThresholdHit (kind soft) and proceeds.
{ decision: 'deny', resource, reason }Emits BudgetThresholdHit (kind hard) and fails the call with Budget exhausted on '<resource>': <reason>.

A denied call raises RuntimeError in Python, rejects in Node, and returns an error in Go; the session stays open. A missing guard method is treated as the permissive default in every SDK. In every SDK, callback errors, timeouts, and malformed check* returns fail closed as deny.

Cluster event vocabulary

AgentEvent reserves three cluster-level variants with stable JSON tags, so a host and external producers can target one schema:

  • BudgetThresholdHit { resource, kind, consumed, limit, message? } (budget_threshold_hit) — the agent loop emits it on the session event stream when a budget guard returns soft (kind soft) or deny (kind hard, with the reason in message).
  • PassivationRequested { reason, deadline_ms? } (passivation_requested) — the host is asking the session to reach a safe, persistable state so it can be evicted from this node. deadline_ms, when present, is the Unix-epoch deadline in milliseconds before forced close.
  • PeerInvocation { from_session_id, from_tenant_id?, correlation_id? } (peer_invocation) — another session invoked this one. The labels let the receiver attribute the call back to its origin tenant and correlation chain.

The framework produces only BudgetThresholdHit. It never emits PassivationRequested or PeerInvocation; a host that uses them emits them itself through the transport it owns. These variants are not hook event types, so registerHook / register_hook handlers do not receive them — observe them on the event stream (see Hooks for the hook event list).

Deterministic IDs and time (replay)

A cluster that wants bit-identical replay of a run on a different node must remove the two sources of nondeterminism in a normal run: random IDs and the wall clock. The Rust core models both behind a HostEnv { id_generator, clock }. The default pairs a UUID generator with the system clock; replay tooling swaps in a SequentialIdGenerator and a FixedClock so that re-executing the same inputs produces the same IDs and timestamps, and therefore the same output, on any node.

All four SDKs expose the same deterministic configuration. The ID prefix and fixed timestamp are independent — an omitted field keeps its system-backed default — so set both for replay, and recreate the configuration from the same values at the start of each replay.

Rust
Node.js
Python
Go
Rust
use std::sync::Arc;
use a3s_code_core::{
host_env::{FixedClock, HostEnv, SequentialIdGenerator},
Agent, SessionOptions,
};
#[tokio::main]
async fn main() -> a3s_code_core::Result<()> {
let agent = Agent::new("agent.acl").await?;
let host_env = Arc::new(HostEnv::new(
Arc::new(SequentialIdGenerator::new("replay")),
Arc::new(FixedClock::new(1_700_000_000_000)),
));
let session = agent
.session_builder(".")
.options(SessionOptions::new().with_host_env(host_env))
.build()
.await?;
println!("{}", session.session_id());
session.close().await;
agent.close().await;
Ok(())
}

Loop checkpoints and run resumption

With a sessionStore / session_store configured, the agent loop persists a loop checkpoint, keyed by run id, at the end of each turn once that turn's tool calls have settled. The checkpoint is deleted when the run finishes in-process. Any node that shares the same store can recover the run. The checkpoint supplies facts, not decisions: the recovering node seeds a fresh fact-log thread from it and folds that log, which chooses the next step.

Rust
Node.js
Python
Go
Rust
use a3s_code_core::{Agent, SessionOptions};
#[tokio::main]
async fn main() -> a3s_code_core::Result<()> {
let agent = Agent::new("agent.acl").await?;
let session = agent
.session_builder(".")
.options(
SessionOptions::new()
.with_file_session_store("./.a3s/sessions")
.with_session_id("session-from-node-a"),
)
.build()
.await?;
let result = session.resume_run("run-id-from-node-a").await?;
println!("{}", result.text);
session.close().await;
agent.close().await;
Ok(())
}

The resumed work is recorded as a new run; the checkpoint's run is left intact in the store. When no checkpoint exists but the session's fact log already has facts, resume_run folds that log instead. The two errors below fire only when the fact log is also empty:

  • resume_run requires a session_store on this session: no store was configured; fall back to a fresh session.
  • no loop checkpoint found for run 'X': the run never reached its first checkpoint, or the checkpoint was deleted when the run ended; treat the run as settled or lost.

A tool call whose result is missing from the log runs once on resume. Hosts that need an exact, replay-safe recovery id use spawnRecoveryWithRunId. See Persistence for the full rules.

Retention caps for long-running sessions

A session that runs for hours or days accumulates state in four in-memory stores: run records, per-run event buffers, trace events, and terminal subagent task snapshots. Left unbounded, these grow with session age — fine for short-lived sessions, a real leak for long-lived ones.

SessionRetentionLimits caps each store. The defaults are:

FieldDefault
max_runs_retained64
max_events_per_run2,048
max_event_bytes_per_run8 MiB
max_trace_events8,192
max_terminal_subagent_tasks512

Omitted fields keep these defaults; set unbounded: true (Rust: SessionRetentionLimits::unbounded()) only when unlimited retention is intentional. Caps are soft: when a store reaches its cap, the oldest entry is dropped on insert and no error is returned. Running subagent tasks are never dropped — only terminal snapshots are evicted, oldest completion first.

Use retentionLimits in Node, opts.retention_limits in Python, or SessionOptions.RetentionLimits in Go. Rust hosts use SessionOptions::with_retention_limits(...). See Limits.


See also: Multi-machine · Persistence · Limits · Hooks