Talking to Whizi in voice mode

The short answer

Voice mode is a real-time spoken conversation: you talk, the model answers out loud, and you can interrupt it. It is a Powerhouse feature, and the plan includes 500 minutes per month.

The session produces a transcript that lands in your chat history, so a spoken conversation is searchable and re-readable afterwards exactly like a typed one. Voice minutes are counted against their own monthly meter and do not spend credits.

Starting and controlling a session

Open voice mode from the chat, and the session connects and begins listening. There is no push-to-talk step: you just start talking.

The controls during a session are deliberately minimal. You can mute the microphone, which stops it hearing you without ending the session, and you can end the session entirely. The state indicator shows whether Whizi is currently listening or speaking, which matters more than it sounds, because knowing whether you have been heard is most of what makes a voice interface tolerable.

The transcript can be opened in the chat at any point, including mid-session. That is the escape hatch when something needs to be precise: read what was actually captured rather than assuming.

How the 500 minutes are counted

Minutes count session time, and the allowance resets monthly along with everything else on the account. It is shared across web and mobile, so time spent talking on the phone comes out of the same 500 minutes as time spent on the desktop.

500 minutes is a little over eight hours per month, or roughly twenty minutes on every working day. For most people this is not a binding constraint, because voice conversations are naturally short: the format suits five minute exchanges rather than hour-long sessions.

Muting the microphone does not stop the clock. Ending the session does. If you are stepping away, end it rather than muting.

What voice is actually better at

Voice is not a faster version of typing, and using it that way is disappointing. Three cases where it genuinely wins:

Thinking out loud. Talking through a problem you have not structured yet is much easier spoken than written, because you do not have to commit to a sentence before you know where it ends. The transcript afterwards is usually a better description of the problem than anything you would have typed.

Hands busy or eyes elsewhere. Cooking, driving, walking, or working through something physical. This is the obvious one and it is real.

Practising a conversation. Interview answers, a difficult message to a client, a pitch. Saying it out loud and hearing a response surfaces the awkward phrasing that reads fine on a page.

Where it loses: anything involving code, exact numbers, names that are hard to spell, or long structured output you will want to copy. Type those.

Workflow checklist
  • Voice mode requires the Powerhouse plan
  • Powerhouse includes 500 voice minutes per month
  • Minutes are shared across web and mobile
  • Voice does not spend credits, it has its own meter
  • Muting does not stop the clock, ending the session does
  • The transcript lands in chat history and stays searchable
  • Use typing for code, exact numbers and anything you will copy
Common questions

Frequently asked questions

Which plan includes voice mode?

Powerhouse only, at $49.99 per month or $34.99 per month billed yearly, and it includes 500 minutes per month. Voice is one of four things that live exclusively on Powerhouse, alongside AI video generation, AI music generation and side-by-side model chats.

Do voice minutes come out of my credits?

No. Voice has its own monthly meter of 500 minutes and spends no credits, so a long spoken conversation does not reduce the number of Claude or GPT messages you can send. Credits, image generations, voice minutes and video generations are four separate allowances that do not share a pool.

Can I see what I said in a voice session?

Yes. Every session produces a transcript that goes into your chat history, and you can open it during the session as well as after it. This is worth doing whenever precision matters, because speech recognition on names, numbers and technical terms is the weakest part of any voice interface.

Can I interrupt the model while it is speaking?

Yes, the conversation is real-time rather than turn-based, so you can cut in the way you would with a person. This matters more than it sounds: without it, a long wrong answer has to be waited out, which is the single most frustrating thing about voice interfaces that lack it.