To measure AI output quality, watch what people do with each output inside your product. Count how often they use it, how much of it they change first, and how often they ask for another one. Those three numbers are acceptance rate, edit distance and regenerate rate, and together they tell you more than any eval score.
Evals still matter. They tell you the model follows the rules your team wrote down. But you wrote those rules, and you know your product better than anyone who will ever use it. A test suite you built can only fail in ways you already thought of. The person deciding whether to keep paying grades the output on something else, and the only place that grade gets recorded is in their clicks.
The reply draft that passed all 386 eval cases
Picture an AI tool for small hotels. A guest email lands in the shared inbox: arriving on the 14th at 9 in the morning after a night flight, is there any chance of checking in early? Under the email sits a blue button, Draft reply. Click it and, after about four seconds, a draft fills the reply box. Three buttons sit under it: Send, Try again and Discard.
The draft reads well. It opens with Thank you for choosing to stay with us. It explains that standard check-in is from 3 in the afternoon and that early check-in is subject to availability on the day. It suggests the guest ask at reception on arrival. It is 104 words, warm, and it never promises a room.
In the team's eval dashboard, under Evals, then Guest replies, then Latest run, this draft type passes. The suite has 386 cases and checks four things: the tone is courteous, no room or time is promised, the check-in policy is stated, and the reply stays under 150 words. The pass rate on the last run reads ninety-four percent, in green.
Now watch the reservations manager. Her room screen shows that room 12 is empty the night before. She deletes the policy paragraph, deletes the suggestion to ask at reception, and types: Room 12 will be ready from 10, and you are welcome to leave your bags with us before that. Then she presses Send. Of the 104 words the model wrote, the first line survived.
Across the product, in this example, sixty-one percent of sent drafts were changed by more than half before they went out. Eighteen percent were sent back with Try again at least twice. Neither number appears anywhere the founder looks. And the rule that kept the eval score high, never promise a room, is exactly why the draft was useless to her.
What the founder looks at instead
The founder opens the eval dashboard after every prompt change. When the team moved to a newer model version last month, the pass rate went from ninety-one to ninety-four, and he posted the screenshot in the team channel with a short note: evals up again. The investor update that month carried a line about draft quality at ninety-four percent.
Weekly active hotels did not move. Two of the larger groups asked on renewal calls whether the drafts could sound less like a call centre. He read that as a tone problem, added a rule to the system prompt, and wrote twenty new cases to check it. The suite passed. Someone on the team has suggested pulling a hundred sent emails and laying them next to the drafts that started them. That card has sat on the sprint board for three weeks under Nice to have.
None of this is careless. He measures what he can see. But every case in the suite was written by people who know how the product is meant to behave. The manager knows something the suite cannot: room 12 is free. That gap between what the maker knows and what the user knows is what in-product numbers measure.
Why a green eval run says little about who stays
Most advice on measuring AI output quality stops at evals: build a golden set, add a model as judge, track the pass rate. Keep doing that. Our claim is less popular. Past a basic bar, the eval pass rate tells you almost nothing about whether people will keep using the product.
Evals grade the output against a rubric. Users grade it against their afternoon. The rubric asks whether the draft is polite and short. The manager asks whether it saved her typing. A draft can pass the first and fail the second, and the suite will never notice, because it cannot see the room screen.
Insiders joke that their evals are vibes, meaning someone reads a sample and nods. Even a careful suite is the founder's taste written down. Behaviour inside the product is the only grade written by the person who pays.
The cost lands in the margin. Every Try again is a second full generation, paid for in tokens. Every draft rewritten from scratch is inference spent on words that got deleted. AI product builders average around 52 percent gross margin, against 70 to 80 percent for traditional software, so waste in the output line comes out of a margin that is already thin. And if activation is counted as first draft sent, the dashboard marks a hotel as activated the moment the manager sends a reply she typed herself.
Three numbers your product already produces
Every AI output in your product ends in one of four ways. The person uses it as it is, changes it and then uses it, asks for another, or throws it away. Count those endings and you have the three numbers.
Acceptance rate is outputs used, divided by outputs shown. Used means the thing the output exists for actually happened: the reply was sent, the summary was shared, the code was committed. Copying is a weaker signal, so count it separately.
Edit distance is how much of the output the person changed before using it. The simple version compares the draft with what was sent, counts the characters that differ, and divides by the length of the draft. Zero means sent untouched. One means rewritten. Sort the results into three bands: light edits under a fifth, heavy edits up to half, and rewritten above that.
Regenerate rate is the share of outputs where the person asked for another try. Track a second line for two or more tries in a row. One retry can be curiosity. Two in a row is the person telling you the answer missed and they had no way to say how.
Read together, they point at different problems. High acceptance with heavy edits means people like the starting point but replace the facts, so the model is missing context. Low acceptance with repeated retries means people keep hoping for a better roll, so the model is missing an instruction the screen gives them no way to send. Both are product problems, and neither shows up in an eval run.
How to log them before Friday
You need four events and an afternoon of reading, not a new tool.
- Give every output an ID the moment it appears. Log output_shown with that ID, the prompt version, the model version and the length in characters.
- Log output_regenerated with the same ID and the attempt number, so a third try is visible as a third try.
- When the output is used, log output_used with the final length and the edit ratio. Work out the ratio on your server by comparing the stored draft with what was sent. Keep the customer's text out of your analytics tool.
- Log output_discarded when the person clears the box or leaves with the draft unused.
- Tag each event with the account and the account's age in weeks, so you can split every number by customer and by how long they have been around.
- Then pull fifty used outputs and set each one beside its draft. The numbers tell you where to look. Reading tells you why.
If nobody on the team has time for the reading, that is the part to hand to someone outside. A UX audit studio that works on AI products will do it in a week, and they will not have written the rubric.
One weekly number next to the eval score
Three numbers are too many for the first row of a dashboard. Fold them into one: the share of outputs used on the first try with light edits. In the hotel tool, that is drafts sent with no retry and less than a fifth changed, divided by all drafts shown. Call it first-try usable. Put it beside the eval pass rate and report both in the investor update.
Then add a rule. A prompt or model change ships when the evals pass and first-try usable holds for a week on a slice of accounts. Being eval-gated is the floor. The behaviour number is the gate that counts. When the two disagree, believe the people using it.
Finally, test the claim in this post on your own data. Split accounts by their first-try usable in month one and check who is still active in month three. If the line holds, you have found the quality number that tracks retention. It also belongs inside your activation event, as defining activation for an AI product argues: count the output used, not the output seen.
What the numbers point at on the screen
The numbers do not fix anything. They point at a screen. In the hotel tool, nearly all the heavy edits hit the same two paragraphs, the policy line and the advice to ask at reception. That says the drafter cannot see what the manager sees. Part of the fix is context and part is design: a line above the draft reading Using: booking #4471, room status for the 14th, last two emails. Now she can tell at a glance what the model read and what it missed.
The double retries point at a missing control. Try again gives the model no hint about what was wrong, so the next draft misses the same way. Replace it with three chips under the draft, Shorter, Drop the policy line and Confirm early check-in, chosen from the edits people make most. Each chip is one of her deletions turned into a button. The broader set of recovery moves is in designing for AI errors; this is the measured version, where the options come from your edit logs instead of a guess.
We took the same route with Dualite, an AI builder: the product work went into the moments people stall, such as the first run, the empty project and the failed build, rather than into the model behind them.
Where Studio Maydit fits
Studio Maydit designs the screens inside AI products where people decide whether an output is worth keeping: the draft box, the retry, the first run. Our fixed-scope product design projects run three to four weeks and end with a diagnosis of what is leaking in the product, which for a drafting tool usually means reading edit logs beside the screens that produce them. If your evals are green and your users are still rewriting, book a 30-minute call with Studio Maydit.





