Skip to main content
src/agent/agent-loop.ts is the core. There is one loop, one agent, one flat message history. Everything else — streaming, tools, sandbox, persistence — exists to feed it or consume what it emits.
Agent Loop

The loop, in one paragraph

For each turn we call our internal streamText({ maxSteps: 1, tools, messages }) once. The model can either emit text, or emit tool calls. Tool calls returned in a single step are executed in parallel via Promise.all. The results are appended to messages and the loop runs again. We stop when finish_reason is "stop" and there are no pending tool calls, or when maxTurns is hit.
maxSteps: 1 is deliberate — the outer loop owns turns, not the SDK. This gives us a single place to compact history, run self-heal, dedup tool calls, and emit progress.

finish_reason quirk

Some providers return tool calls with finish_reason: "stop" instead of "tool_calls". The loop must clear the local tool-call buffer on either value, or the next turn will replay stale calls.
Don’t simplify this to just "tool_calls". It will look fine in tests with one provider and silently fail with another.

Context: routing at 75 %, compaction at 90 %

Two decoupled steps, in this order, every turn. Routing (75 %). When projected tokens cross 75 % of the current model’s window and CONTEXT_AUTO_UPGRADE is on, the loop swaps to the largest-context model available and re-budgets. Nothing is discarded — the context just gets a bigger window. Compaction (90 %). If projected tokens are still past 90 % of the selected model’s window — no larger model exists, or the context exceeds even the biggest one — the loop pauses, runs ContextManager.compact(), and resumes. The compactor:
  • preserves the system prompt and the most recent user/assistant pair verbatim,
  • summarises older turns into a single rolled-up message,
  • copies .mana/memory.md from the sandbox into the summary so durable facts survive.
Why the split: compacting early thrashes — we lose recent context we could have kept simply by moving to a bigger window. Leaving it later than 90 % trips provider context-overflow rejections, which are unrecoverable mid-stream. Both thresholds are tuned. Don’t change them casually.

Self-heal

Tool errors are caught and fed back to the model as tool results. The model is asked to retry. After 3 retries on the same error key the loop gives up and surfaces the error. Without the cap, a flapping sandbox or a syntactically broken template can burn an entire session looping on the same fix. The cap is per error, not global — a session can absorb many distinct errors.
Agent Loop

Provider failover

If the default provider returns a stream-level error, the loop transparently retries on the configured fallback provider. If the user explicitly picked a model, we don’t silently swap it — they get the error back.

Empty-step retry

Some providers (Gemini, occasionally) sometimes return a step with zero text and zero tool calls — a real empty turn. Without intervention the loop would terminate thinking the model was done. We detect empty steps and retry once with the same messages. If the second attempt is also empty, we give up.

What ends a session

  • finish_reason === "stop" with no pending tool calls.
  • maxTurns reached (AGENT_MAX_TURNS_PER_REQUEST, default 15).
  • 3-retry cap hit on a single error.
  • abortSignal fired (explicit POST /vibe/abort/:projectId).
  • Provider stream error with no failover available.
Closing the SSE connection does not end the session — that’s intentional, so a flaky network doesn’t kill an in-progress build.