How to generate images inside a Whizi chat

Quick answer

To generate images in a Whizi chat, open the model picker, switch it from Text to Image & Video, pick Whizi Image, Nano Banana, Flux, or Stable Diffusion, and describe the image. Image generation is included on Pro (100 per month) and Powerhouse (500). Iterate in the same thread, then download the result.

The short answer

Open the model picker, switch the toggle at the top from Text to Image & Video, pick an image model, and describe the image. The composer hint changes to "Describe the image you want to generate...", and the result streams back into the same conversation.

ModelWhat it is for
Whizi ImageMaking any image; the first row in the creative picker
Nano BananaEditing a photo you attach
FluxSharp photoreal art
Stable DiffusionStability AI's image model

Image generation is included on Pro, at 100 images per monthly period, and Powerhouse, at 500. Free and Starter accounts have an image allowance of zero. An image prompt can be 1 to 4,000 characters, and generated images are kept for 30 days, with a download control on each one. If a generation errors instead of rendering, the wording names the cause; see image generation failed.

Why generating in the chat is different

The usual image workflow is a separate product, a separate subscription, and a separate blank box that knows nothing about what you are working on. Generating inside the conversation changes two practical things.

First, the model already has the context. If you have just spent twenty messages developing a campaign concept, "generate the hero image for this" carries all of it. You are not re-describing the brief to a different tool.

Second, you can use a text model to write the image prompt. This is the trick most people miss: ask Claude or GPT to turn your rough idea into a properly structured image prompt, then run it. Language models are considerably better at writing image prompts than most people are, because the format is a known one and they have seen a great many examples of it.

Prompt: let the text model write the image prompt

Write an image generation prompt for: [rough idea]. It will be used for [purpose and placement]. Structure it as: subject, action or state, setting, lighting, composition and framing, medium or lens, colour treatment, mood. Add a negative list of what must not appear. Give me three variations that differ in composition rather than in adjectives.

The prompt structure that works

Nearly every good image prompt has the same seven parts in roughly this order. Missing parts get filled in by the model with its defaults, and the defaults are why generic prompts produce generic stock imagery. If this is your first time writing one, how to create AI images covers the same ground from scratch.

PartWhat to sayEffect if you omit it
SubjectThe specific thing, with the detail that mattersYou get the most generic version of the noun
Action or stateWhat it is doing, or how it sitsStatic, posed, lifeless
SettingWhere, and how much of it is visibleA blank or cluttered default background
LightingDirection, quality, time of dayFlat, evenly lit, characterless
CompositionAngle, distance, where the subject sits in frameCentred medium shot every time
Medium or lensPhotograph and focal length, or illustration styleDefaults toward glossy digital art
Colour and moodPalette and feelingOversaturated, high contrast

A worked example. Weak: a person working in an office. Strong: A woman in her fifties reviewing printed drawings at a standing desk, late afternoon light from a window to her left, shot from slightly behind her shoulder at eye level, 50mm, muted greens and warm neutrals, calm and unhurried. Negative: stock-photo smiling at camera, cluttered desk, visible brand logos, text.

The negative list matters more than people expect. Most of what makes generated images look generated is a small set of recurring artefacts, and naming them removes them: stock-photo smiling, over-saturated, text, watermark, extra fingers, symmetrical corporate composition.

Choosing the model

  • Whizi Image for making any image. It is the house generator and the first row in the creative picker.
  • Nano Banana for editing a photo you already have. Attach the picture and describe the change; the composer switches its hint from generating to editing as soon as an image is attached.
  • Flux for sharp photoreal art and anything that has to sit alongside real photography.
  • Stable Diffusion, Stability AI's model, as the second opinion: it rewards specific prompting and produces a different look from the same words.

You can switch between them inside the same chat, which is the fastest way to answer the only question that matters: run the same prompt through two models and look at them next to each other. Model quality varies enormously by subject, and the winner on portraits is often not the winner on interiors or on flat illustration.

Iterating without going in circles

The common failure is regenerating the same prompt over and over hoping for a better roll. Some of that is unavoidable, but most wasted attempts come from changing several things at once and losing track of what helped.

  • Change one variable per attempt. Lighting, or framing, or palette. Not all three.
  • Keep what worked in words. When an image is close, describe what is right about it in your next prompt rather than assuming the model remembers.
  • Fix composition before detail. Getting the framing and the light right first is far more efficient than perfecting a subject that is in the wrong part of the frame.
  • Generate at the final aspect ratio. Cropping a square into a wide banner throws away the composition you asked for. Ask for the ratio you actually need up front.
  • Save the prompts that worked. Eight reliable house-style prompts you can paste again are worth more than any single image, and reusing them is what makes a set of assets look like a set.

For editing an image you already have, upload it and describe the change. This works well for adjustments to background, lighting, and palette, and less well for precise structural edits, which are still faster in a real image editor.

What image models still get wrong

Text. Rendered words remain the weakest area across every model. Short words sometimes come out; a headline, a logo, or a label usually does not. Generate the image clean and add text in your design tool.

Hands, counts, and small repeated details. Fingers, teeth, chair legs, and windows in a building are all places where models lose count. Zoom in before you use anything.

Precise spatial instructions. "Exactly three objects, the red one on the left" is unreliable. Compose in your editor if the arrangement genuinely matters.

Consistency across images. Getting the same character or product to appear identically in a series is difficult. Detailed, repeated descriptions help; identical results are not guaranteed.

On usage: check your rights before anything generated goes into a paid placement, and take particular care with anything resembling a real person, a recognisable location, or a distinctive style associated with a living artist. Provider terms differ, and the commercial exposure sits with whoever publishes the image.

Workflow checklist
  • Ask a text model to write the image prompt from your rough idea
  • Include all seven parts: subject, action, setting, lighting, composition, medium, mood
  • Add a negative list naming the artefacts you never want
  • Generate at the aspect ratio you will actually use
  • Change one variable per iteration
  • Run the same prompt through two models and compare
  • Add text in your design tool, not in the image model
  • Save the prompts that worked as reusable templates
Common questions

Frequently asked questions

How many images can I generate per month?

Pro includes 100 image generations per monthly period and Powerhouse includes 500. Free and Starter accounts have an image allowance of zero, so Pro is the lowest plan with image generation. On a weekly billing cycle the allowance is the monthly figure divided by four, rounded up. Since iteration is where generations are spent, writing a structured prompt first (ideally with a text model) noticeably reduces how many attempts a usable image takes.

Can I edit an existing image?

Yes. In Image & Video mode, attach the photo and describe the change; the composer switches its hint from generating to editing as soon as a picture is attached, and Nano Banana is the model built for photo edits. This works well for background, lighting, palette, and style adjustments. It works less well for precise structural edits such as moving an object a specific distance or changing text, which are still faster in a normal image editor.

Which image model should I use?

Whizi Image for making any image, Nano Banana for editing a photo you attach, Flux for sharp photoreal art that sits next to real photography, and Stable Diffusion for a different look from the same prompt. Because model quality varies a lot by subject, the fastest way to decide is to run the same prompt through two of them in the same chat and compare the results.

Why does the text in my image come out wrong?

Rendered text is still the weakest area across every image model, and no prompt reliably fixes it. Short single words sometimes work; headlines, logos, and labels usually do not. Generate the image without text and add typography in your design tool, which also gives you the correct font and proper control over placement.

Do I need a separate image subscription alongside Whizi?

For most campaign, content, and concept work, no. Image generation is included on Pro (100 per month) and Powerhouse (500), which covers the same ground a standalone subscription would for typical marketing and social output. A dedicated tool still makes sense for teams with very specific stylistic requirements or a deep existing prompt library they do not want to rebuild.