Auto Compact
A long session grows its transcript every turn. Enable autoCompact and the
runtime folds earlier context once the transcript approaches
autoCompactThreshold of the model's context window, so you do not manage
tokens by hand.
How compaction works
Three session options control it:
autoCompact/auto_compact— off by default.autoCompactThreshold/auto_compact_threshold— a fraction of the context window, clamped to 0.0–1.0, default 0.8. Lower it to compact earlier.maxContextTokens/max_context_tokens— the context window used for the calculation. It defaults to the configured model's declared context limit, falling back to 200,000 tokens.
Session turns run on the fact log. After each user message or tool result, the
runtime measures the folded transcript. When it reaches about
maxContextTokens × autoCompactThreshold × 4 characters (roughly four
characters per token) and the current turn has not compacted yet, the runtime
records a compaction.done fact that folds the earlier transcript into a single
leading context message, and the stream emits a context_compacted event. The
fold keeps the earlier text as is; it does not ask the model to summarize it.
Without autoCompact, the same fold happens only at a fixed
1,000,000-character ceiling.
The SDK history API (history() / History) returns the messages the session
currently carries.
Compaction in a Meta Harness recipe
If the session uses an explicit Meta Harness recipe,
compaction is one of its stock parts. compactAfterChars / compact_after_chars
sets that part's character threshold directly; when omitted, the threshold
derived from autoCompact above applies. toolBudget / tool_budget belongs to
the budget part: it is the number of successful tool calls allowed per user
message, after which the next tool call is denied and the turn ends.