Reading time:

7 min read

|

Last updated:

Designing AI Voice Products Without a Screen

How to design an AI voice product: what to say while the model thinks, barge-in, recovering from mishearing, and what goes on the companion screen.

We design websites and products that make AI companies more money.

Siddarth Ponangi

Founder, Studio Maydit

We design websites and products that make AI companies more money.

Web and product design for AI companies

We help AI companies build fast, clean, and conversion-focused websites and products.

Designing an AI voice product means designing time instead of screens. The user cannot see a spinner, scroll back, or tap a button, so the product has to fill the wait while the model thinks, stop talking the moment the user cuts in, and confirm what it heard before it acts. Most voice products still need a small screen for setup, history, and anything the user must read or check later.

On a screen, a mistake sits there until someone fixes it. In voice, a mistake is gone the second it is spoken, and the user has to remember it. Every design choice below comes from that one difference.

Latency: what to say while the model is thinking

In a normal talk, a pause means something. A few seconds of silence on a call makes people ask if the line dropped. A voice product that waits for the full answer before it speaks will feel broken even when it works.

There are three ways to fill the gap. Start speaking as soon as the first words are ready, instead of waiting for the whole reply. Use a short spoken cue, such as one moment while I check that, when a task will take longer. And for work that takes more than a few seconds, say what is happening, like looking up your last order.

Be careful with filler. The same phrase before every answer gets old by the third turn. Save cues for the slow cases, and vary them. If a device has a light or a small display, a clear thinking state there does the same job without words.

Barge-in: letting people interrupt

People interrupt each other all the time. They cut in to correct, to skip ahead, or because they already heard what they needed. This is called barge-in, and a voice product has to allow it.

When the user starts talking, the product should stop speaking almost at once. It should not finish its sentence first. It also has to avoid hearing its own voice through the speaker and treating that as the user, which is an engineering problem the design depends on.

Decide what an interruption means. Sometimes it means stop. Sometimes it means the user is correcting one detail. A short mm or yes from the user may not be an interruption at all. Keep spoken answers short, and offer more only if asked. Short answers leave less to interrupt.

Recovering from a misheard request

Speech recognition will get things wrong, especially names, numbers, and noisy rooms. Without a screen, the user cannot see the mistake. So the product has to say it back.

Match the check to the risk. For low risk actions, repeat the key detail inside the reply, like setting a timer for ten minutes. The user hears it and can correct it. For actions that cost money, send a message, or cannot be undone, ask a direct yes or no question first.

When something is unclear, ask only about the part that is unclear. If the product heard the day but not the time, ask for the time. Do not make the user repeat the whole request. After two failed tries, change the approach. Offer two choices, send a link by text, or pass the user to a person.

Our post on designing for AI errors covers the wider question of how to show a model is wrong. Voice makes the same problem harder, because the error leaves no trace.

When a voice product needs a screen anyway

Voice is good for quick requests and hands-busy moments. It is poor at some jobs, and a companion screen handles those.

Setup is one. Connecting accounts, granting access, and choosing preferences are easier with taps than with talk. Long lists and comparisons are another, because nobody can hold six options in their head from one hearing. Anything the user must copy or check exactly, such as a code, an address, or a price, belongs on a screen. So does history, when the user wants to know what the product did yesterday.

If your voice agent takes phone calls for a business, the screen often belongs to a different person. The caller only hears the voice. The business owner needs a place to read transcripts, see outcomes, and catch calls that went wrong.

What to put on the screen that does exist

Keep it small and useful. Show the state clearly: listening, thinking, speaking, or off. Show a live transcript of what the product heard, so a misheard word is visible before the product acts on it. Let the user tap to fix that word instead of saying it again.

Turn each finished answer into a card the user can keep, with the key facts in text. Give a clear way to stop or mute at any time. For call agents, lead the owner's screen with the calls that need attention, not with totals.

The screen you do not own is your website. A voice product has little to show in a static image, so the landing page has to let a visitor hear it. Our list of landing page design agencies for AI voice products compares studios that have faced that problem.

Where to go next

If you are building the first version, our shortlist of MVP design agencies for AI voice products compares studios on what they publish. If your product mixes voice with chat, our guide to conversational UI design covers how the first reply sets the tone.

Studio Maydit designs websites and products for AI founders in the US, the UK, and Europe, building in Framer, Webflow, or custom code and carrying on into product design. Recent clients include Wave, PixelFlow, Mi-VAD, and 15 other AI and SaaS teams. A fixed-scope project runs three to four weeks and ends with a diagnosis of what is leaking in the product, and a monthly retainer is there with no long lock-in. If your voice product works in a demo and loses people on real calls, book a 30 minute call and play us a recording.

Frequently Asked Questions

Table of Contents
Scroll to view headings
0%