Yes, show AI reasoning to users, but not the raw transcript. Show a short, edited account of what the system checked, collapsed under the answer by default, and hide it entirely when the answer is low stakes or obvious.
Reasoning models made this a live choice, because the thinking now exists as text you could pipe straight to the screen. You know what that text means, since you watched it get better for months. A new user does not. To them, a model arguing with itself in public looks like a product that is not sure, and the hedging lands on the one screen where they decide whether to trust you.
A grey block of 600 words above a home office answer
Picture an AI tax assistant for US freelancers, used here as an example. A user types: I run my design business from the spare bedroom, but my exercise bike is in there too. Can I deduct the room?
Version one answers in two parts. First comes a grey box labelled Model reasoning, open by default, about 600 words long. It reads like a scratchpad because it is one. Let me think about the home office rules. The user mentions a bike. Hmm, wait, that might be wrong, the regular use test may not matter here. Let me reconsider. Actually the key issue is exclusive use. The box goes on like that for four screens on a phone, then the answer appears at the bottom in a white card.
The answer itself is right. The room as a whole probably does not qualify, because the IRS wants the space used only for business. A marked off part of the room that holds nothing but the desk can still count. Under the simplified method that part is worth 5 dollars per square foot, up to 300 square feet, as the IRS home office topic sets out.
Version two runs the same model and gives the same answer, but the answer comes first. Under it sits one grey line with a small arrow: Checked 3 IRS rules. 1 not met: exclusive use. Tap it and a numbered trail of four short lines opens. Exclusive use: not met for the whole room, because the bike is personal. Regular use: met, you work there most days. Main place of business: met, you have no other office. What would change this: move the bike out, or measure only the desk area.
What a user does with the line that says wait
Watch the session recordings for version one and two behaviours show up. Most users scroll past the grey box without reading it, so the 600 words cost them four screens of thumb work and nothing else. A smaller group stops, and they stop at the worst line in the box: wait, that might be wrong.
That group does the expensive things. They ask the same question again in new words to see if the answer holds. They screenshot the hedge and send it to support with a one-line question asking whether the answer is wrong. Some of them leave and ask a human accountant, which is the exact cost the product exists to remove. Nobody files a ticket saying the reasoning box made them doubt the answer. They just stop relying on it.
The founder sees the support volume and reads it as a model quality problem. So the team spends a sprint on the prompt to make the reasoning sound more confident. The answers were already right. The problem was never the model. It was the decision to publish a draft as if it were a receipt.
Why the raw transcript is the wrong evidence
Most advice on AI transparency says to make the thinking visible. Our own guide to agentic UX says it too, and it is right about the goal. Where that advice goes wrong is the form. A raw chain of thought is a working draft. It contains every dead end, every false start and every self correction, because that is how the model gets to a good answer. Drafts are written for the writer, not the reader.
There is also a harder problem. The transcript is not a reliable account of how the answer was reached. In research from Anthropic, reasoning models given a hint mentioned it in their visible reasoning only 25 percent of the time for Claude 3.7 Sonnet and 39 percent for DeepSeek R1. The unfaithful transcripts were longer than the faithful ones. So the long, hedged block that looks most like honesty can be the least honest thing on the screen.
The model providers made the same call for their own products. OpenAI's reasoning guide says it does not expose the raw reasoning tokens, and offers a summary instead. If the people who trained the model ship a summary, a product team piping the raw text to a tax question has skipped a design decision rather than made one.
Checked 3 IRS rules, and why that line reads as rigour
The collapsed line works because it says what was checked, not what was thought. Checked 3 IRS rules is a count of real work. 1 not met tells the user where the answer turns before they open anything. A careful user taps it. A busy user reads the answer and moves on. Both leave trusting the product more than when they arrived.
The trail under it follows four rules. Each line names one check the system ran, in the user's words, not the model's. Each line gives a verdict: met, not met, or could not check. One line says what would change the answer, which turns the trail into advice instead of a defence. And nothing in the trail is written after the fact to sound smart. Every line maps to a step the system actually performed.
That last rule matters most. The trail should come from your pipeline, not from asking the model to narrate itself again. If your system looks up a rule, runs a check or pulls a figure from the user's data, log it as a step with a name and a result. The trail is those logs, rewritten in plain words. If the system cannot report a step, the screen cannot claim it. That is the same rule that keeps a good loading state honest, applied to the answer instead of the wait.
The reply on the launch post, and what it costs
In this example, the reasoning box was a reaction. The founder launched, and the top reply under his post said it was a chat model with a tax prompt on top. He knew what sat behind the product: a rules layer mapped to IRS guidance, an eval suite of hundreds of real freelancer questions, a check that refuses to answer when the user's state changes the outcome. None of it was on the screen. So he turned on the raw reasoning, because it was the only depth he could show in a day.
It showed the wrong depth. The raw text makes the product look like the thing he was accused of being: a model thinking out loud. The edited trail shows the opposite. Checked 3 IRS rules is the rules layer, visible. Could not check: your state rules is the refusal logic, visible. That is the answer to the wrapper reply that does not require him to argue with a stranger.
The costs are real. The investor question is now whether the company would still have a reason to exist if a foundation model provider shipped something ten times better tomorrow. A trail that shows checks the model alone does not run is evidence for yes. Gross margin is the second cost. AI product builders average around 52 percent gross margin against 70 to 80 percent for traditional software, so every repeat question from a doubting user is inference paid twice. The third is activation. Product-led SaaS benchmarks put activation at 20 to 40 percent for most products, and a user who goes to an accountant instead has not activated.
The same logic works on the marketing page. When we built the site for Chariot, a speech reasoning lab, the page shows hard written lines, a drug dose, an invoice amount, a flight time, next to how the model speaks them. A visitor sees the work on the inputs that are hard, which no claim of depth can do.
When to hide AI reasoning entirely
Not every answer needs a trail. Showing one on everything trains users to ignore it, the same way they ignore cookie banners. Use a simple rule based on what the user will do next.
- Hide it when the output is easy to check by looking: a renamed file, a summary of a short email, a draft the user will edit anyway. The result is its own proof.
- Show the collapsed line when the user will act on the answer with money, law or health, or when the answer is a no. A no without reasons reads as the product being lazy.
- Open the trail by default only when the answer contradicts what the user said, such as a deduction they expected and do not get. That is the moment they will look for the reason anyway.
- Keep the raw transcript out of the main screen for everyone. If your users are developers who debug prompts, put it behind a Debug or Raw reasoning tab in settings, not under the answer.
Rewrite one answer screen this week
You can do this without a redesign. Pick the one question type your users act on most, and work through it in this order.
- List what your system actually does for that question. Lookups, rule checks, data pulls, refusals. Ask an engineer for the real steps, not the prompt.
- Give each step a name a user would recognise. Checked exclusive use rule, not Evaluated constraint set 2.
- Write the collapsed line as a count plus the turning point. Checked 3 rules. 1 not met. Keep it under ten words.
- Write the expanded trail as one line per step, each with a verdict. Add one line on what would change the answer.
- Put the answer first and the trail under it. Collapse the trail by default.
- Move any raw reasoning out of the answer view. Keep it in logs, or in a debug tab for technical users.
- Watch five new users ask that question. Count how many re-ask, how many open the trail, and how many contact support. Compare with last month.
If re-asks and support tickets on that question drop, roll the pattern out to the next question type. If nobody opens the trail, keep it. Its job is partly to sit there as proof.
At Studio Maydit we design the screens where a new user decides whether to rely on an AI product: the first answer, the first run, the moments that need proof. Fixed-scope projects take three to four weeks and end with a diagnosis of what is leaking in the product, and you can see how that work is scoped on our product design page, or compare studios in our ranking of design agencies for products that look generic. If your product does real work that nobody can see on screen, book a 30-minute call.





