From token generation to character render: the full streaming pipeline
This question is the senior-level version of "explain how the browser renders a page." It's designed to surface whether you understand the full path data takes — from a token being sampled inside the LLM to a character lighting up on the user's screen. Most candidates know one or two stages. The strong answer connects all of them and identifies where latency is added or hidden.
Let me walk you through it stage by stage. I'll mark the latency budget at each step so you can see where time goes.
Stage 0: Token sampling inside the model
Time so far: 0–800ms (this is the TTFT — time-to-first-token).
Inside the LLM server, the user's prompt has been tokenized, run through the model's forward pass, and the model has produced a probability distribution over the next token. A sampler (greedy, top-p, top-k, temperature, etc.) picks one. That's one token.
This stage is dominated by:
- Prefill latency — the cost of running the prompt through the model once to set up the KV cache. This scales with prompt length. A 4K-token prompt is much slower to start than a 200-token prompt.
- Model size — a 70B model has higher per-token latency than a 7B model.
- Server load — at the API provider level, your request might wait in a queue.
Why this matters for the frontend: there's nothing you can do to make TTFT faster. But you can mask it with a typing indicator, an optimistic message echo, or a cancel button that signals "we're working on it." Perceived latency lives here.
Stage 1: Server emits the token
Time so far: TTFT + ~5ms.
The model server takes the sampled token, decodes it (token ID → bytes), and writes it to the response stream. If this is OpenAI/Anthropic-style SSE, the write looks like:
data: {"type":"token","value":"Hello"}\n\n
If it's a streaming JSON SDK pattern, it might write a JSON line directly. If it's a websocket, it's a frame.
The write goes through the HTTP server's buffer. Most servers default to some buffering for performance. If you've ever seen "tokens stream in bursts of 5" instead of one at a time, this is why — Nagle's algorithm, framework buffering, or proxy buffering somewhere upstream is batching writes. Disabling Nagle (TCP_NODELAY) or flushing after each write is required for true token-by-token streaming. Most production setups do this.
Stage 2: Network transit
Time so far: TTFT + ~5ms + RTT (round-trip time).
The bytes travel over TCP from the API server to the user's browser. RTT depends on geography:
- Same continent: 30–80ms
- Cross-continent (US ↔ EU): 100–150ms
- Half-the-world (US ↔ Asia): 150–300ms
This is unavoidable. CDN edge POPs help if the inference is at the edge (rare), but for hosted LLM APIs the inference is in a fixed region, so RTT is fixed too. What you can do is keep the connection open with HTTP/2 keepalive, so subsequent tokens don't pay TCP handshake or TLS handshake costs again. That's "free" because every modern HTTP client does it by default.
Stage 3: Browser network stack receives the bytes
The browser's networking layer (Chromium's net stack, in Chrome) reads the bytes off the socket. They land in a kernel buffer, get copied into the browser's process memory, and become available to the JS layer.
This is invisible to your code, but it's where bytes are bytes. No characters, no JSON, no events. Just Uint8Array chunks delivered to your ReadableStream.
Stage 4: Your fetch/EventSource consumer
1const reader = response.body.getReader();2const { value } = await reader.read(); // value is Uint8Array
Each read() resolves with the next chunk the network has delivered. The size and timing of these chunks is not aligned to your tokens. You might get:
- One token per chunk (if the server is flushing aggressively and the network is low-latency).
- Five tokens per chunk (if any layer batches).
- Half a token per chunk (if a multi-byte character was split — see the TextDecoder question).
Chunk timing is also subject to the operating system's network event loop. Linux delivers chunks in TCP segments, which the browser groups based on its read scheduling. There's typically 5–20ms of variance here.
Stage 5: Decoding and parsing
Time added: ~0.1–1ms per chunk.
Your TextDecoder turns bytes into a string. Your SSE parser (or NDJSON parser, or whatever) turns the string into structured events. Each event is a token + metadata.
This is where the bugs live. UTF-8 boundary bugs, JSON parse errors on incomplete chunks, missing the data: prefix. None of this is visible if everything works — but each is a 10–50ms regression if it doesn't, plus likely visible artifacts.
Stage 6: State update
Time added: ~0.5–5ms (depends on framework and component tree).
You call setMessage(prev => prev + token) or your equivalent. React schedules a re-render. The old text + new token becomes the new state.
In React 18+, multiple state updates that happen in the same microtask are batched. This is good — if 5 tokens arrive in one chunk and you call setMessage 5 times in a tight loop, only one re-render happens for the whole batch. In React 17 and earlier, you'd see 5 re-renders inside the React.unstable_batchedUpdates wrapper or none if your state updates were inside a setTimeout.
If you want the highest token throughput at the cost of some flicker, you can buffer tokens client-side (e.g. accumulate for 16ms — one frame — before flushing to React). This is what Vercel's AI SDK does internally.
Stage 7: React reconciliation and commit
Time added: ~1–8ms for a small chat component, ~10–30ms for a complex one.
React diffs the new state against the previous tree. For a chat message, the diff is usually trivial — one text node changed. The commit phase writes the new text to the DOM via nodeValue = '...' or similar.
If your message uses Markdown rendering (most chat UIs do), each setState triggers a full Markdown re-parse of the message text. For a 2000-character message, that's ~5–15ms per token if your Markdown library is naive. Caching the AST across renders or using an incremental Markdown parser (rare) is the optimization here. Most apps just live with the cost.
Stage 8: DOM mutation and style/layout
Time added: ~0.5–5ms.
The DOM API (Text.nodeValue = ..., appendChild, etc.) updates the tree. The browser invalidates layout if the change affects size — and a growing message always affects size (it's getting taller). So layout reflow happens.
Most chat UIs have only the message itself growing, so layout is local: the message box grows, sibling messages don't change. If you have an autoscroll behavior (scrollIntoView or manual scrollTop adjustment), that adds another paint.
Stage 9: Paint and pixel
Time added: ~1–4ms.
The browser composites the new layout into a paint layer, sends it to the GPU, and the GPU draws pixels. Modern compositors do this at 60 or 120Hz, so worst case you wait one frame (~16ms or ~8ms) for the next vsync.
The total budget
| Stage | Latency | Cumulative |
|---|---|---|
| Token sampling (TTFT) | 200–800ms | 200–800ms |
| Server write + flush | 5ms | 205–805ms |
| Network transit (RTT) | 30–300ms | 235–1100ms |
| Browser net stack | <1ms | 235–1100ms |
| TextDecoder + parse | <1ms | 236–1101ms |
| React state update + render | 1–10ms | 237–1111ms |
| DOM commit + layout + paint | 2–10ms | 239–1121ms |
Per-token incremental cost (after TTFT) is dominated by RTT (the time between tokens), then React rendering. Everything else is in the noise.
What to optimize and what to leave alone
Optimize: TTFT masking (typing indicator, immediate cancel button), Markdown re-parsing if it shows up in profiling (memoize the AST), and your decoder/parser pipeline (use { stream: true }, use a streaming JSON parser).
Don't optimize prematurely: the network stack, React batching (it's already optimal), or the paint pipeline. These are not the bottleneck for token streaming.
Mask, don't reduce: TTFT can't be reduced from the client. But you can make it feel shorter with a typing indicator that appears in <100ms, an optimistic echo of the user's prompt, and confidence-building UI ("Thinking…", "Searching the web…", "Generating…").
What interviewers probe
-
"Where would you add a token to appear faster?" — you can't. Token speed is set by the model + RTT. You can only mask perceived latency with optimistic UI and confidence indicators.
-
"What's the difference between TTFT and inter-token latency?" — TTFT is dominated by prefill (one-time cost per request). Inter-token latency is dominated by RTT and per-token model decode time. Different bottlenecks, different fixes.
-
"What about partial Markdown rendering?" — you have to handle the case where the LLM is mid-codeblock or mid-list. Either render incrementally (most libraries do this fine) or accept that the rendered output flickers a little as it stabilizes. Most chat UIs render Markdown live and it looks great because Markdown's grammar is forgiving.
The senior-level move is connecting the whole pipeline. Most candidates know "fetch returns chunks" and "React renders." Few articulate "here's the latency budget end-to-end, here's where it's spent, here's what I can and can't optimize." That's the answer that lands.
Related Frontend AI interview topics
Continue with nearby questions from the same topic cluster.