The short answer
Voice mode is a real-time spoken conversation: you talk, the model answers out loud, and you can interrupt it. Every paid plan includes voice minutes: 10 per month on Starter, 80 on Pro, and 500 on Powerhouse.
The session produces a transcript that lands in your chat history, so a spoken conversation is re-readable afterwards like a typed one. Voice minutes are counted against their own meter and do not spend credits.
Starting and controlling a session
Open voice mode from the chat, and the session connects and begins listening. There is no push-to-talk step: you just start talking. Replies come in one preset voice, a warm voice named Sulafat, and voice runs on the Gemini Live API: Whizi mints a short-lived session token and your browser streams audio straight to Google, so no audio passes through Whizi servers.
The controls during a session are deliberately minimal. You can mute the microphone, which stops it hearing you without ending the session, and you can end the session entirely. The state indicator shows whether Whizi is currently listening or speaking, which matters more than it sounds, because knowing whether you have been heard is most of what makes a voice interface tolerable.
Because the conversation is real-time rather than turn-based, you can cut in while the model is still speaking, the way you would with a person. Without that, a long wrong answer has to be waited out before you can redirect it.
The transcript can be opened in the chat at any point, including mid-session. That is the escape hatch when something needs to be precise: read what was actually captured rather than assuming. Speech recognition on names, numbers and technical terms is the weakest part of any voice interface, so a mid-session look at the transcript is the fastest way to catch a mishearing before it shapes the next answer.
How the minutes are counted
Minutes count session time, and the allowance is counted per account and resets with your billing period. A single session runs at most 15 minutes: the start call reserves time from your allowance and tells the session how long it may run, so when your remaining minutes are lower than that, the session is shortened to fit.
The allowance follows the plan: 10 minutes a month on Starter, 80 on Pro, and 500 on Powerhouse. Free accounts get none. Powerhouse works out to a little over eight hours per month, or roughly twenty minutes on every working day. Voice conversations are naturally short, the format suits five minute exchanges rather than hour-long sessions, so the higher tiers are rarely a binding constraint.
Muting the microphone does not stop the clock. Ending the session does. If you are stepping away, end it rather than muting.
Voice minutes are their own allowance and draw on nothing else. Voice minutes, credits and image generations are separate meters that do not share a pool, so a long spoken conversation does not reduce the number of Claude or GPT messages you can send that month. It also means an account that has run its credits down still has its voice minutes, and an account that has talked its voice allowance flat can still type.
The transcript a session leaves behind is an ordinary chat, so it costs nothing further to keep and sits in your history alongside everything you typed.
What voice is actually better at
Voice is not a faster version of typing, and using it that way is disappointing. Three cases where it genuinely wins:
Thinking out loud. Talking through a problem you have not structured yet is much easier spoken than written, because you do not have to commit to a sentence before you know where it ends. The transcript afterwards is usually a better description of the problem than anything you would have typed.
Hands busy or eyes elsewhere. Cooking, driving, walking, or working through something physical. This is the obvious one and it is real.
Practising a conversation. Interview answers, a difficult message to a client, a pitch. Saying it out loud and hearing a response surfaces the awkward phrasing that reads fine on a page.
Where it loses: anything involving code, exact numbers, names that are hard to spell, or long structured output you will want to copy. Type those.
| What you are doing | Spoken or typed | Why |
|---|---|---|
| Thinking through a problem you have not structured yet | Spoken | You do not have to commit to a sentence before you know where it ends |
| Cooking, driving, walking, or working through something physical | Spoken | Hands are busy and eyes are elsewhere, and the transcript is in your history afterwards |
| Practising interview answers, a difficult message, or a pitch | Spoken | Saying it out loud surfaces the awkward phrasing that reads fine on a page |
| Code, exact numbers, hard-to-spell names, long output you will copy | Typed | Speech recognition is weakest on names, numbers and technical terms |
- Every paid plan includes voice: 10 minutes on Starter, 80 on Pro, 500 on Powerhouse
- A single session runs at most 15 minutes
- Voice does not spend credits, it has its own meter
- Muting does not stop the clock, ending the session does
- The transcript lands in chat history and stays re-readable
- Use typing for code, exact numbers and anything you will copy
Frequently asked questions
Which plan includes voice mode?
Every paid plan includes voice mode, and free accounts get no voice minutes at all. The allowance is what changes with the tier: Starter includes 10 voice minutes per month, Pro includes 80, and Powerhouse includes 500, resetting with the billing period.
Do voice minutes come out of my credits?
No. Voice has its own meter of minutes and spends no credits, so a long spoken conversation does not reduce the number of Claude or GPT messages you can send. Voice minutes, credits and image generations are separate allowances that do not share a pool.
Can I see what I said in a voice session?
Yes. Every session produces a transcript that goes into your chat history, and you can open it during the session as well as after it. That mid-session look is worth taking whenever precision matters: names, numbers and technical terms are what speech recognition gets wrong most.
Can I interrupt the model while it is speaking?
Yes, the conversation is real-time rather than turn-based, so you can cut in the way you would with a person, and the model stops instead of finishing a long wrong answer you would otherwise have to wait out.