Reading time:
10 min read
Last updated:
North Star Metric Examples for AI Products
North star metric examples for AI products: why messages and tokens mislead, seven outcome-based metrics by product type, and the guardrail to pair.
A good north star metric for an AI product counts outcomes the user accepted, not work the model did. Examples: support tickets resolved without a human, drafts sent with light edits, code suggestions merged and still there a week later. Messages, tokens, runs and conversations handled are not north stars. They are your cost line wearing a growth costume.
This matters because the north star decides what your team ships for a year. Point it at volume and every sprint makes the model do more. Point it at accepted outcomes and every sprint makes the model more useful. Those are different roadmaps, and only one of them shows up in renewals.
The tile that says Conversations handled: 18,400
Picture an AI support agent sold to online stores. The founder's dashboard opens on one big tile: Conversations handled, 18,400 this month, with a green arrow pointing up beside it. Under it sit two smaller tiles, Avg response time 4s and CSAT 4.2. That is the whole first row.
Now watch one of those 18,400 from the shopper's side. She opens the chat bubble on an order page and types: my order says delivered but nothing came. The agent replies in four seconds with a link titled Shipping and delivery FAQ and the line Did this answer your question? with Yes and No buttons. She ignores both and types agent. The agent offers the same article. She types talk to a person. A grey button appears, Connect me with the team, and a human picks it up eleven minutes later.
On the dashboard, that is one conversation handled. On a report nobody opens, under Insights, then Resolution, then By intent, there is a number for how often the agent closed an issue with no human stepping in. In this example it reads twenty-two percent. Nobody has put it on the first row, because nobody asked for it.
What the founder does with the big number
He puts 18,400 in the monthly investor update, next to last month's 14,000. The team celebrates the jump in standup. To push it further, an engineer ships a change that opens the chat bubble automatically on the order status page. Conversations handled goes up again, because more people now talk to the agent. Most of them were only checking a tracking number.
Meanwhile the support lead at one of his biggest customers keeps the same number of human agents on shift. At the renewal call she asks one question: how many tickets did it actually take off my team? He opens the dashboard to answer and the first row has nothing to say. So he promises a resolution report next quarter, and goes back to tuning prompts so the agent can answer even more.
This is the trap in plain terms. He knows his product works because he has seen it work. The metric he chose agrees with him, because it counts every time the product tried. The one person whose budget pays for it is counting something else.
Why messages, tokens and runs make bad north stars
Most north star advice says pick the number that best reflects engagement. In a normal app that is fair, because engagement is cheap and people only come back when they get value. In an AI product the model can create engagement on its own. An agent that answers badly makes the user ask again, and that counts twice.
Each volume metric fails in its own way.
Messages sent. A confused user sends more messages than a happy one. The worse the answer, the higher the number.
Tokens generated. Longer output scores better, even when the user wanted three lines. This one is your bill, not your growth.
Agent runs or tasks started. A run that fails and gets retried counts twice. A run nobody looks at counts the same as one that saved an afternoon.
Conversations handled. Every chat the bot touched, including the ones a human finished. It counts attempts and calls them wins.
There is a money reason this hurts more for AI companies. Every one of those messages and tokens cost inference, and inference sits in cost of goods. AI product builders average around 52 percent gross margin against 70 to 80 percent for traditional software. A volume north star asks the team to grow the exact line that eats your margin, whether or not anyone got helped.
North star metric examples by AI product type
First, one line to keep two ideas apart. Activation is the first time a new user acts on an output, and our guide on how to define activation for an AI product covers how to find it. The north star is the same kind of accepted outcome, counted across every user, every week, as the single number the whole company steers by.
Here is what an outcome-based version looks like for common AI products. Each one counts the moment the user took the output and used it, and each has a time window so a quick undo does not count.
AI support agent. Issues resolved with no human takeover, where the customer did not reopen within 72 hours.
AI coding assistant. Suggested changes merged into the main branch and not reverted within seven days.
AI writing tool. Drafts the user sent or published, where they kept most of the text instead of rewriting it.
AI sales email agent. Emails sent from the agent's draft that got a reply from the prospect.
AI research or analyst tool. Reports or charts the user exported, shared or pinned to a dashboard.
AI image or video tool. Assets downloaded or placed into a project, not just generated.
AI agent that takes actions. Tasks the user approved and did not undo, such as a booked meeting that stayed on the calendar.
The pattern is the same in every row. The model's output has to leave the chat and do a job somewhere the user cares about. That last case, where the agent acts on its own, has its own design problems around approval and undo, which we cover in designing for AI agents.
The market is already moving this way. Intercom prices its Fin support agent at $0.99 per outcome, and counts an outcome when a customer confirms the issue is resolved, does not ask for more help, or Fin completes a set workflow. If your buyer pays for outcomes, a north star that counts conversations is measuring something nobody is buying.
The guardrail number that sits beside it
Any single metric can be gamed, including a good one. A support agent can raise its resolution count by making the Connect me with the team button hard to find. People give up, do not reopen, and count as resolved. So every north star needs one guardrail metric beside it, a number that goes bad when the north star is being gamed.
For the support agent, the guardrail is human takeover rate plus reopen rate. For the coding assistant, it is revert rate. For the writing tool, it is the share of drafts deleted unsent. Put the guardrail on the same row as the north star, at the same size. If the north star climbs and the guardrail climbs with it, you have not improved anything. You have moved the failure somewhere the main tile cannot see.
A second guardrail is worth having for AI products in particular: inference cost per accepted outcome. If resolutions go up and cost per resolution doubles, the model is brute-forcing its way to the number.
How to pick your north star this week
You do not need a data team. You need one export of last month's outputs and an afternoon.
Write down the one job a paying customer would say your product does, in their words. For the support agent: it takes tickets off my team.
List every place an output can go after the model makes it. Sent, copied, merged, exported, approved, shared, ignored, regenerated, deleted.
Pick the destination that matches the job. Mark it as the accepted outcome. Everything else on the list is either a step toward it or a sign of failure.
Add a time window that rules out quick undos: 72 hours for a ticket, seven days for code, one send for an email.
Pull last month's count. Expect it to be much smaller than your volume tile. In our example, 18,400 conversations might become around 4,000 resolutions. That smaller number is the real one.
Choose one guardrail from the failure signs on your list, usually the undo, the reopen or the human takeover.
Move both to the first row of the dashboard. Keep the volume tile, but rename it to what it is, such as Conversations attempted, and move it down.
Then send the next investor update with both numbers and one sentence on why you changed them. A founder who can explain why his headline number got smaller usually gets better questions than one whose number only goes up.
What changes in the product once the number changes
This is the part most teams do not expect. Change the north star and the roadmap changes on its own. In the support example, the team stops asking how to get more people into the chat. They start asking why the agent sends a FAQ link when the shopper said the parcel never arrived. The fix is a product fix: the agent checks the carrier status, says the parcel shows delivered at 2:14pm, and offers Report missing parcel as a button. That one screen does more for resolutions than any prompt tuning did for volume.
It also changes who the team designs for. A volume metric is designed for the model, because it rewards output. An outcome metric is designed for the person reading the output, because only they can accept it. That is the same shift product-led companies make when the product has to do the selling, which our piece on product-led growth design walks through.
Why the founder is the last to see it
He reads every output as someone who knows what a good one looks like. When the agent sends the FAQ link, he sees a correct answer, because it is correct. The shopper sees a link to a page she already read. He knows his product and nobody else does, so the metric he picks counts what he sees, not what she does next.
Studio Maydit designs the screens where an AI output gets accepted or ignored: the reply, the buttons under it, and the path from an answer someone reads to one they act on. Our AI product design work runs as fixed-scope projects of three to four weeks, and each one ends with a diagnosis of where the product is leaking. If your biggest tile keeps growing and your renewals do not, book a 30-minute call with Studio Maydit.
Frequently Asked Questions
Continue Reading

ChatGPT Is Introducing Ads. Here’s the UX Risk Nobody Is Talking About
As ChatGPT prepares to introduce ads, most conversations focus on revenue and scale. But the bigger question is how monetization reshapes user trust, cognitive flow, and product intent. This Studio Notes piece explores the hidden UX risks product teams should pay close attention to.

Siddarth Ponangi

Why designing for power users too early breaks SaaS products
Many SaaS products become difficult to use not because they lack features, but because they introduce complexity before users are ready for it. Designing for power users too early often feels like progress, but it quietly undermines adoption for everyone else.

Siddarth Ponangi

Why second-use experience matters more than first impressions in SaaS
Many SaaS products spend enormous effort optimizing first impressions. What often gets overlooked is what happens when users come back for the second time, which is usually where real adoption either starts or quietly falls apart.

Siddarth Ponangi

