Skip to main content
The agent emits events, not request/response payloads. Two transports speak the same event shape, and both are first-class.
Streaming

Transports

SSE — Server-Sent Events

Both keep the connection alive with a : heartbeat\n\n SSE comment every 15 s, which beats the typical 30 s idle timeout on Cloudflare / nginx / ELB. Without it, intermediaries silently drop SSE during long agent turns. V1 wraps every chunk as data: {"type":…,"data":…} and terminates with data: [DONE]. V2 puts the same type on the SSE event: line and terminates with event: done.

WebSocket

WS /api/v1/ws — bidirectional, 1 MB max payload. The server pushes the same chunk types as SSE. Client-to-server messages are subscribe, unsubscribe, cancel (stops the run), submit_answers (resolves an ask_user_question), console_log_data, and ping.

Event types

Every event carries a type from StreamChunkType in src/types/index.ts — that union is the list. A representative sample: All events are also fan-out via EventEmitter to any other subscribers attached to the same project (Redis pub/sub when queue mode is on — see below).

Client close ≠ session abort

Closing the SSE connection does not stop the agent. This is intentional: a flaky network shouldn’t kill a 5-minute build. To explicitly stop a session:
The next agent turn checks the abort signal and exits cleanly.
Streaming

ask_user flow

Timeout is 5 minutes. After that the tool rejects and the agent self-heals (usually by making a reasonable default choice).

Queue mode (Redis-backed)

When QUEUE_ENABLED=true:
Two effects:
  1. Horizontal scaling. Many web pods can accept requests; a smaller worker pool processes them. A pod restart doesn’t kill in-flight builds.
  2. Idempotency. Re-connecting a client to an in-flight session simply re-subscribes to the channel — no duplicate work.
If queue mode events aren’t reaching the client, look at the Redis pub/sub channel name in src/queue/ and the relay generator in src/routes/vibe.ts.

Backpressure

Neither SSE writer buffers without limit. Before enqueuing a chunk the stream waits while controller.desiredSize <= 0, backing off 2 ms → 50 ms until the consumer drains. A slow client (a paused tab, a stalled proxy) therefore slows the generator rather than growing a queue until the process OOMs, and browsers absorb short stalls in their own buffer without the user noticing.