The best practices for LLM streaming UI are to render whole blocks instead of raw tokens, stop auto-scrolling the moment the reader scrolls up, and put a Stop button where the send button was. Streaming is a design decision, not a setting you switch on, and a stream that flickers, jumps and cannot be stopped can feel slower than a plain spinner.
For a founder this matters because the stream is the first thing a new user watches your product do. The model can be fast and right, and the screen can still tell a stranger that it is struggling. You never see that version, because you already know what the answer will say before it arrives.
The NDA that flickers into shape
Picture an AI drafting assistant for in-house lawyers at small companies. The screen is a chat panel on the right and a blank document area on the left. A lawyer types into the box: Draft a mutual NDA with a hardware supplier, two-year term, Delaware law. She clicks the arrow.
Text starts arriving almost at once, which is the part the team is proud of. But the panel shows the raw output first and formats it later. A line appears as ## 1. Definitions, pound signs and all, and a beat later it snaps into a heading and pushes everything below it down. The phrase **Confidential Information** sits on screen with two asterisks on each side until the closing pair arrives, then turns bold. A clause comparison comes in as a table, one cell at a time, so for a while the panel shows rows of vertical bars and dashes.
Meanwhile the panel keeps scrolling to the bottom. She tries to read clause one, the definition that decides everything else, and the text slides up out of view every time a new line lands. She scrolls back up. The panel drags her down again. The send arrow has turned grey, and there is no other button. If she can see by clause three that the term is wrong, she has no way to stop it. She waits for clause nine to finish, then types a correction.
Now run the same model with three changes. Markdown is held back until each line is complete, so a heading appears once, already a heading. The panel follows new text only while she is at the bottom, and the moment she scrolls up it stays put and shows a small Jump to latest pill. The grey arrow becomes a square Stop button. Same tokens, same speed. The second version reads as roughly twice as fast, because nothing on screen fights her.
Most advice says stream everything. That is half right.
The usual guidance is short: turn on streaming, because people hate waiting. Jakob Nielsen set out the reason decades ago. About 10 seconds is the limit for keeping a user's attention on a task, and past that people want feedback that the system is working. Streaming gives that feedback. So far, so good.
Where the advice goes wrong is in treating the token as the unit of display. A token is a unit of billing. It is not a unit of reading. People read in words, lines and finished items. When every token repaints the screen, the user is not watching an answer appear. They are watching a renderer change its mind. Each reflow resets where their eyes were.
So the claim of this post is simple, and some engineers will disagree with it. A stream that is fast but unstable is worse than a stream that is slightly slower and steady. Hold text back by a line, or by a table row, and the user loses almost nothing in speed and gains a page that stays still long enough to read. The goal is not the earliest possible pixel. It is the earliest moment the user can start reading and keep reading.
Why the founder never sees the flicker
In this example the founder demos the product most days. He types the NDA prompt himself, and while the text streams he talks. His eyes are on the person he is talking to, not on the panel. By the time he looks back, the draft is finished and formatted, and it is good.
When a new hire mentions the screen jumping around, he files a ticket titled Make streaming faster. Engineering spends a sprint cutting time to first token, and they succeed. The stream now starts sooner and jumps around sooner. The session logs still show the same pattern: new users send one draft request, scroll up and down for a while, and close the tab. He reads that as a prompt quality problem and starts rewriting the system prompt.
He knows what an NDA should say and where each clause will land, so the flicker costs him nothing. The lawyer trying to check the definition clause has never seen this draft before. She is the one who needs the page to hold still.
What a stream that fights the reader costs
The first cost is activation. In this tool, the moment a lawyer trusts a draft enough to edit it in place is a fair activation event. A panel she cannot read while it writes pushes that moment later, or past the point where she gives up. Product-led SaaS benchmarks put activation at 20 to 40 percent for most products, and a ten point improvement typically drives a 15 to 25 percent increase in free-to-paid conversion. The first streamed answer is one of the cheapest places to win those points.
The second cost is margin, and it is the one founders rarely connect to a Stop button. Without one, every wrong draft runs to the end. The user waits for all of it, then asks again, and you pay for two full answers to get one. AI product builders average around 52 percent gross margin, against 70 to 80 percent for traditional software, because inference sits in cost of goods. A Stop button is a design decision that also cuts output you would otherwise pay for and nobody reads.
The third cost is the room. A demo where the screen jitters reads as early, even when the draft is excellent. Investors rarely say the scroll lost them. They say it needs more polish, and the founder hears a note about colours.
Six rules for a stream people can read
None of these touch the model. They are rendering and layout decisions, which is why they get skipped.
- Render complete lines, not raw tokens. Buffer text until a line or sentence ends, then paint it with formatting already applied. An unclosed pair of asterisks or a lone pound sign should never reach the screen.
- Follow the reader, do not drag them. Auto-scroll only while the user is already at the bottom of the panel. The moment they scroll up, stop following and show a small Jump to latest button. When they click it, resume.
- Swap the send button for Stop. Put Stop in the exact spot the send arrow was, so the user does not have to hunt for it. When they press it, keep the partial answer on screen and offer Edit prompt and Retry right under it.
- Stream structured output as finished parts. Show the table header first, then add one complete row at a time. Show a list item only when it is whole. Never show a half-built table with pipes and dashes.
- Reserve space so nothing jumps. If the answer will contain a citation card, a chart or an image, draw its frame before the content arrives. Content that appears inside a fixed box does not shove the text the user is reading.
- Show the work before the first word. For an answer that runs steps, such as searching your clause library, list the steps as they happen: Reading your template, Checking Delaware rules, Drafting. A visible step is better than a blinking cursor with nothing behind it.
Our post on LLM output formatting covers what the model should write in the first place. These rules are about how that text arrives on screen, and that decision usually has no owner.
When a spinner beats a stream
Not every answer should stream. If the output is a single verdict, a number, or a short yes or no, streaming it word by word just makes a two-second answer feel like a slow one. Show a brief loading state, then the whole answer at once.
The same goes for output that drives the interface rather than being read. If the model returns structured data that fills a form, a table or a chart, stream the steps if the wait is long, but render the result when it is complete and valid. A half-filled form looks broken, not fast. This is also where a screen often beats a chat reply in the first place.
Finally, do not stream text you may have to take back. If a guardrail or a check runs after generation and can block the answer, a user who watched three paragraphs appear and then vanish trusts the product less than one who waited. Hold the answer until the check passes, and treat the failure path with the same care as any other AI error state.
An afternoon test for your stream
Pick the three requests new users send most often in their first session. Record your own screen while each one streams, with your hands off the keyboard and your mouth shut. Then watch the recordings at normal speed and count three things: how many times the text under your cursor moves, how many times raw markdown is visible, and whether there was any way to stop a wrong answer.
Next, try to read the first paragraph of each answer while it is still streaming. If you cannot finish it without scrolling back, the scroll rule is broken. Fix that first, because it is usually a few lines of front-end code. Then fix the renderer, then add Stop. Then record again and compare.
When we redesigned the product side of Dualite, an AI app builder, the screens that mattered most were the ones where non-technical builders watch the product work, such as the first run and a build that fails. That is where a stranger decides whether the tool is thinking or stuck.
You already know what the answer says
The flicker, the scroll and the missing Stop button all survive for one reason. The person who built the product reads every stream already knowing how it ends. A new user does not. For them the stream is the product, and right now it is showing them a renderer arguing with itself.
Studio Maydit is a product and web design studio for AI founders in the US, UK and Europe, and we spend most of our time on the screens where a new user decides to stay, including the ones that show a model at work. If you are comparing outside teams for this kind of work, our ranking of design agencies for AI chat interfaces is a fair place to start. Our fixed-scope projects run three to four weeks and end with a diagnosis of where the product is losing people. If your answers are good and your first sessions are short, book a 30-minute call with Studio Maydit.





