Reading time:
11 min read
Last updated:
LLM Output Formatting Best Practices for AI Products
LLM output formatting best practices: answer in the first sentence, set length by question type, and allow only the markdown your chat UI renders.
The best practice for LLM output formatting is to put the answer in the first sentence, set a length rule for each kind of question, and allow only the markdown your interface can actually render. Write those rules into the system prompt, because if you do not, the model picks a format for you, and it usually picks a report.
For a founder this is a product decision, not a prompt tweak. The reply is the screen in an AI product. If the thing the user asked for is buried under headings and caveats, they read the first lines, decide the product talks too much, and do not ask a second question. That second question is often where activation starts.
The churn answer where the reason is bullet nine
Picture an AI analyst for subscription businesses. A head of growth connects Stripe, opens the chat panel and types: why did churn go up in March?
The reply is 640 words long. It opens with a heading, Overview, and two sentences restating the question. Then comes Key Drivers, with fourteen bullets split across three subheadings. Then a bold line that starts Note: this analysis is based on available data and may not reflect all factors. Then a table with six columns: Plan, Seats, MRR, Churned, Churn Rate, Change. On a phone, the table scrolls sideways and the last two columns sit off screen.
The answer is in there. Bullet nine says the Starter plan went from $19 to $29 on March 4, and most of the cancelled accounts were on Starter. That one line is what the person needed. It is the ninth thing they would have to read, and it has the same weight as bullet three, which says seasonal patterns may play a role.
Nothing in this reply is wrong. Every bullet is true. The model did exactly what it was left to do, which is produce something that looks thorough. Nobody on the team decided that a why question should come back as a memo. Nobody decided anything at all, and that is the problem.
Why the model writes a memo when nobody asked for one
Left alone, a chat model tends to format like a careful analyst writing for a manager. Headings, bullets, a caveat, a summary at the end. That shape reads as effort, and effort is what the model has learned looks helpful. Without a rule, it applies the same shape to a yes or no question as it does to a request for a full plan.
Most advice on this topic says use headings and bullets so answers are easy to scan. That is advice for documents. A chat reply is not a document. It is one turn in a conversation, read in a narrow panel, often on a phone, by someone who asked one thing. Scannable is the wrong goal. The goal is that the first sentence is the answer and the rest is optional.
The usual fix is a line in the system prompt that says be concise. It rarely holds, because it tells the model what to avoid and not what to produce. Anthropic's own prompting guide recommends telling the model what to do instead of what not to do, and notes that a prompt written in heavy markdown tends to pull more markdown into the reply. If your system prompt is a wall of headings and bullets, your answers will look like it.
Why the long answer looks fine in the demo
In this example, the founder does not see the problem, because he never reads the answers the way a customer does. When he demos the analyst, he asks a question, scrolls straight past Overview, and points at bullet nine. He knows where the answer lives because he knows what the answer is. On a call, the long reply looks impressive. Alone, it looks like homework.
So the evidence he collects points the other way. Demos go well. Self-serve users ask one question and leave. When someone on the team suggests the replies are too long, he adds be brief to the system prompt, ships it, and the next week the replies are a little shorter and shaped exactly the same. The ticket gets closed. The session logs still show most new users sending a single message.
What a 640-word answer costs
The first cost is activation. If your activation event is a user acting on an answer, a reply where the answer is bullet nine delays that action or kills it. Product-led SaaS benchmarks put activation at 20 to 40 percent for most products, and a ten point improvement typically drives a 15 to 25 percent increase in free-to-paid conversion. A first answer that buries the point is one of the cheapest places to lose those points.
The second cost is margin. Every word the model writes is output you pay for, read or not. A 640-word reply costs roughly ten times the output of a sixty-word one that says the same thing first. AI product builders average around 52 percent gross margin, against 70 to 80 percent for traditional software, because inference sits in cost of goods. A product that writes memos by default is paying to produce paragraphs nobody reads, on every question, from every user, including the ones who never come back.
The third cost is the follow-up. People who get a long answer tend to rephrase and ask again, hoping for the short version. That is another full reply, billed again, for a question that was already answered.
A format spec you can paste into the system prompt this week
Treat output format the way you treat a component in your design system. Write it down once, give it rules, and make the model follow it. Here is a starting spec. Change the numbers to fit your product.
Answer in the first sentence. The first sentence answers the question in the user's own words, with the key number or name in it. For the churn example: Churn rose in March mainly because the Starter plan went from $19 to $29 on March 4.
Set length by question type. Why and what happened questions get one to three sentences, then one line of evidence. How do I questions get numbered steps, six at most. Compare questions get a table only if there are three or more items and two or more things to compare. Draft me questions get the draft and nothing around it.
No headings under 150 words. A short answer with a heading on it reads like a form letter. Headings are for replies the user asked to be long.
One caveat, only when it is specific. Not this analysis may not reflect all factors. Instead: March data is missing refunds from the last two days of the month. If there is nothing specific to say, say nothing.
End with one offer, not a summary. A single line such as Want this broken down by plan? gives the user a next step. A closing summary repeats what they just read.
Write the spec in plain prose. Describe the shape you want in sentences, with one short example reply. A system prompt written as nested bullets invites bulleted answers.
This spec is the design work. It decides what the user sees first, how much they have to read and where their next click goes. Those are the same decisions you would make for any other screen, which is why we put them alongside the rest of AI product UX rather than leaving them to whoever owns the prompt.
Markdown is a promise your renderer has to keep
Every formatting rule in the prompt has a twin in the front end. If the model is allowed to write a table, your chat panel needs to render a table that works at 375 pixels wide, or turn it into stacked cards. If the model is allowed bold, your renderer has to show bold and not two literal asterisks on each side of a word. Raw asterisks in a reply are a small thing that makes the whole product look unfinished.
So make the list from the screen backwards. Open the chat panel on a phone. Write down what it can display well: maybe paragraphs, bold, numbered lists and code blocks. Everything else goes on a short not allowed list in the system prompt, phrased as what to use instead. Use a numbered list rather than a table when there are fewer than three items. Use plain sentences rather than headings in short replies.
If your product has more than one surface, such as a chat panel, an email digest and a Slack message, each surface needs its own rules. The model cannot know which one it is writing for unless you tell it. Pass the surface in with the request and give each one its own spec.
Test the shape on thirty real questions
You cannot judge format from one good demo. Pull thirty real first questions from your logs, the ones new users typed in their first session. Run them through the new prompt and put the replies side by side with the old ones. Then score each reply on four checks:
Does the first sentence answer the question, with the key fact in it?
How many words come before the first number or name the user needed?
Does it render cleanly on a 375 pixel screen, with no sideways scrolling?
Would the user know what to ask or click next?
The first two checks are easy to automate, so turn them into an eval that runs whenever someone changes the system prompt. Format drifts every time a prompt is edited or a model version changes underneath you, and nobody notices until a customer screenshots a reply that is mostly headings. This connects directly to what you count as activation for an AI product: if the answer is acted on, the format worked.
We wrote separately about the first response in a chat interface, the one reply that decides whether a second message gets sent. The spec above is about every reply after that one too, because a product that answers well once and then slides back into memos teaches users to stop asking.
You are the wrong reader for your own answers
The founder in this example is not careless. He reads every reply knowing what it should say, so his eye goes straight to bullet nine and the other thirteen bullets cost him nothing. A new user does not know which bullet matters. They read from the top, and the top is a heading and a restated question. The format looks fine to the one person who could find the answer without it.
Studio Maydit designs the screens where AI products win or lose a new user, and in a chat product the reply is that screen. We work with AI and SaaS founders in the US, UK and Europe on the first run, the first answer and the step after it, and our fixed-scope projects of three to four weeks end with a diagnosis of where the product is leaking. If your answers are right and people still stop after one question, book a 30-minute call with Studio Maydit.
Frequently Asked Questions
Continue Reading

ChatGPT Is Introducing Ads. Here’s the UX Risk Nobody Is Talking About
As ChatGPT prepares to introduce ads, most conversations focus on revenue and scale. But the bigger question is how monetization reshapes user trust, cognitive flow, and product intent. This Studio Notes piece explores the hidden UX risks product teams should pay close attention to.

Siddarth Ponangi

Why designing for power users too early breaks SaaS products
Many SaaS products become difficult to use not because they lack features, but because they introduce complexity before users are ready for it. Designing for power users too early often feels like progress, but it quietly undermines adoption for everyone else.

Siddarth Ponangi

Why second-use experience matters more than first impressions in SaaS
Many SaaS products spend enormous effort optimizing first impressions. What often gets overlooked is what happens when users come back for the second time, which is usually where real adoption either starts or quietly falls apart.

Siddarth Ponangi

