The short answer
The cheapest priced row in the Whizi cost index is Ling from inclusionAI, at $0.000053 to answer once. It charges one credit per message, as does every other row at the bottom of the ladder, and the floor on a turn guarantees nothing is ever charged less than one.
One cheap row per plan, as a lead into the full tables below:
| Model | Provider | Cost per answer | Context | Credits | Plan |
|---|---|---|---|---|---|
| Ling-3.0-flash | inclusionAI | $0.000053 | 262K | 1 | Powerhouse |
| Llama 3.3 70B Instruct | Meta | $0.00026 | 131K | 1 | Pro |
| GPT-5.6 Luna | OpenAI | $0.0008 | 1M | 1 | Starter |
A standard answer in the index is 1,000 input tokens plus 500 output tokens, so list rates from different providers become comparable. The source prices were fetched from OpenRouter on 2026-08-20. The index prices 100 rows across 29 providers, and 89 of those rows carry a Whizi credit charge. The full dataset is published as the AI model cost index.
This page is the cheap end of that index. For the catalogue by plan see the model list, and for the credit ladder itself see the credits reference.
Rows under a tenth of a cent per answer
A selection of the rows that answer once for less than $0.001 at the standard answer size. All but one of the rows printed here charge a single credit.
| Model | Provider | Cost per answer | Context | Credits | Plan |
|---|---|---|---|---|---|
| Ling-3.0-flash | inclusionAI | $0.000053 | 262K | 1 | Powerhouse |
| Nex-N2-Mini | Nex Agi | $0.000075 | 262K | 1 | Powerhouse |
| Solar Pro 4 | Upstage | $0.00009 | 524K | 1 | Powerhouse |
| Qwen3.7 Flash | Qwen | $0.000095 | 1M | 1 | Powerhouse |
| Granite 4.1 8B | IBM | $0.0001 | 131K | 1 | Powerhouse |
| Laguna XS 2.1 | Poolside | $0.00012 | 262K | 1 | Powerhouse |
| Phi 4 | Microsoft | $0.00014 | 16K | 1 | Powerhouse |
| Nemotron 3.5 Lightning | NVIDIA | $0.00018 | 262K | 1 | Powerhouse |
| Llama 3.3 70B Instruct | Meta | $0.00026 | 131K | 1 | Pro |
| DeepSeek V4 Flash 0731 | DeepSeek | $0.00028 | 1.31M | 1 | Powerhouse |
| Gemini 2.5 Flash Lite | $0.0003 | 1M | 1 | Pro | |
| DeepSeek V3.2 | DeepSeek | $0.000469 | 164K | 1 | Pro |
| Qwen3 Coder Next | Qwen | $0.00052 | 262K | 1 | Powerhouse |
| Llama 4 Maverick | Meta | $0.0006 | 1M | 1 | Pro |
| DeepSeek V3.1 | DeepSeek | $0.000725 | 164K | 1 | Pro |
| Codestral 2508 | Mistral | $0.00075 | 256K | 1 | Powerhouse |
| Qwen3.6 Flash | Qwen | $0.00075 | 1M | 1 | Powerhouse |
| MiniMax M2 | MiniMax | $0.000765 | 205K | 1 | Powerhouse |
| GPT-5.6 Luna | OpenAI | $0.0008 | 1M | 1 | Starter |
| GPT-5.4 Nano | OpenAI | $0.000825 | 400K | 2 | Powerhouse |
| MiniMax M3 | MiniMax | $0.0009 | 1M | 1 | Powerhouse |
| Qwen3.7 Plus | Qwen | $0.00096 | 1M | 1 | Pro |
Several of the cheapest rows carry a million tokens or more, and the largest window in the whole price index belongs to a row near the bottom of the cost list. Cheap and small are not the same axis here. The smallest window printed above is 16K.
Most of these rows sit on the most expensive plan, for the reason set out below.
The next band, up to the index median
The median priced row is a Kimi reasoning row at $0.00185 per standard answer. A sample of the rows between a tenth of a cent and that median:
| Model | Provider | Cost per answer | Context | Credits | Plan |
|---|---|---|---|---|---|
| Gemini 3.1 Flash Lite | $0.001 | 1M | 1 | Powerhouse | |
| Muse Glimmer 30B | Meta | $0.0011 | 131K | 2 | Powerhouse |
| Mistral Large 3 2512 | Mistral | $0.00125 | 262K | 2 | Pro |
| GLM 4.7 | Z.ai | $0.001275 | 205K | 2 | Powerhouse |
| Gemini 3.7 Flash | $0.001313 | 1M | 2 | Pro | |
| Mistral Medium 3.1 | Mistral | $0.0014 | 131K | 2 | Pro |
| GLM 4.6 | Z.ai | $0.0015 | 205K | 1 | Powerhouse |
| Gemini 2.5 Flash | $0.00155 | 1M | 2 | Pro | |
| Gemini 3.5 Flash Lite | $0.00155 | 1M | 2 | Powerhouse | |
| GLM 5 | Z.ai | $0.00156 | 205K | 2 | Pro |
| Kimi K2 0711 | Moonshot | $0.00172 | 131K | 10 | Powerhouse |
| Kimi K2 Thinking | Moonshot | $0.00185 | 262K | 10 | Powerhouse |
The credit column and the dollar column do not cover the same set of rows. Eleven of the 100 priced rows carry no credit charge at all: they are priced for comparison but are not offered as their own metered row in the product. A dollar figure in the index therefore does not by itself mean the row can be opened in the picker.
Why most of the cheapest rows are Powerhouse
The plan column above is not a pricing decision, it is a fallback. Model access resolves by identifier: the four Starter identifiers first, then the agent personas, then an explicit list of 37 further identifiers for Pro, and anything else returns Powerhouse. A cheap row that nobody ever promoted into the Pro list therefore sits on the top tier by default, which is why the bottom of the cost list reads as a Powerhouse column. The fallback itself is documented on the model list.
| Plan | What it reaches at the cheap end |
|---|---|
| Free | No credit allowance. Seven messages, lifetime, on any text model |
| Starter | Four identifiers in total: Auto, Whizi AI, GPT-5.6 Luna, Gemini 3 Flash. All of them cost 1 credit |
| Pro | Starter plus 37 more, including the DeepSeek, Llama, Gemini Flash, Mistral and Qwen rows marked Pro above |
| Powerhouse | Everything else in the catalogue, which is where the smaller-lab rows at the very bottom of the cost list sit |
The free tier is the odd one out: free accounts are not model-gated. The chat path returns before the model gate for free accounts, so a prospect can try any text model inside a lifetime allowance of seven messages that never resets. The cheapest row in the index and the dearest row in the catalogue are equally reachable on it.
Starter is the narrow one. It reaches four identifiers, every one of them charges a single credit, and its 400 credit allowance therefore equals 400 messages. There is no Anthropic model on Starter, and no access to the cheap smaller-lab rows in the tables above.
What one credit buys
A credit is defined as one message on the house model, which runs on the base OpenAI model pinned at one credit. A build test fails if either of them ever leaves that rung, because moving either one re-denominates the entire ladder.
The charge is a fixed integer per model id. At this end of the catalogue that means a 1 credit row costs one credit whatever the input or output length: paste a long document into it, let the reply run long, and the turn is still charged one.
| Row | Cost per answer | Cost per 1,000 answers | Credits |
|---|---|---|---|
| Nemotron 3.5 Lightning | $0.00018 | $0.18 | 1 |
| Llama 3.3 70B Instruct | $0.00026 | $0.26 | 1 |
| DeepSeek V4 Flash 0731 | $0.00028 | $0.28 | 1 |
| Gemini 2.5 Flash Lite | $0.0003 | $0.30 | 1 |
| DeepSeek V3.2 | $0.000469 | $0.469 | 1 |
| Llama 4 Maverick | $0.0006 | $0.60 | 1 |
| GPT-5.6 Luna | $0.0008 | $0.80 | 1 |
| Qwen3.7 Plus | $0.00096 | $0.96 | 1 |
Turned into an allowance, the bottom rung is the one place where the credit figure and the message figure are the same number: every turn costs one, so a plan allowance of N credits is N messages for as long as you stay on 1 credit rows. The allowance per plan, the reset and the rollover rule are all in the credits reference.
One more mechanic that matters at the cheap end: a turn that does not fit in the remaining balance is refused whole rather than part-charged. Twenty credits on three remaining credits is rejected, not discounted. On a 1 credit model that almost never bites.
Where a credit rung comes from, and when owner policy overrides it
The cost index measures a standard answer of 1,000 input tokens plus 500 output tokens against provider list rates. The credit ladder is derived from a different reference turn, 3,000 input tokens and 800 output tokens, divided by what that same turn costs on the anchor model, which is $0.00156 of provider spend. The result is then rounded up to a legal rung. The ladder holds eighteen of them: 1, 2, 3, 4, 5, 6, 8, 10, 12, 15, 20, 25, 30, 40, 50, 60, 80, 100. Fourteen carry at least one model today, and the empty four are kept deliberately, because a missing rung would round the next expensive model up to the one above it. Rounding goes up when a model falls between two rungs.
On top of that, some rows are pinned by owner policy rather than arithmetic. Three examples visible in the tables above and in the wider catalogue:
- The Kimi family is charged 10 credits against cost-true rates of 2, 3, 4 and 12 depending on the row. That is why two Moonshot rows sit within a hair of each other in dollars, at $0.00172 and $0.00185, yet both charge 10 credits.
- The Gemini row on Starter is charged 1 credit against a cost-true 2, which is what keeps Starter a flat one credit per message.
- The GPT Terra row is charged 4 credits against a cost-true 10.
Search-grounded rows are priced differently again, because the provider charges a flat fee per search that does not scale with the per-token rate. One search at $0.005 is worth 3.2 credits, so a single search costs more than an entire message on most of the catalogue. Those rungs are floors that assume one search per answer.
A model identifier with no rung at all is charged the top of the ladder, 100 credits, deliberately: an unpriced model should surface as a support ticket rather than quietly eat margin. A build test fails if any catalogue model or agent identifier reaches production without a rung.
Finally, credit pricing is enabled per platform. A platform on the allowlist sees real multipliers in the picker and is charged them against the credit allowance. A platform that is not sees a flat one credit per turn and keeps the older message allowance instead, which is 400 messages a month on Starter, 800 on Pro and 5,000 on Powerhouse. The web client is on the credit system today.
- The cheapest priced row in the index answers once for $0.000053
- A standard answer is 1,000 input tokens plus 500 output tokens, priced from OpenRouter list rates on 2026-08-20
- Every charge is floored at one credit, so no message costs less than one
- Most of the very cheapest rows sit on Powerhouse, because anything not on the Starter or Pro list falls through to the top tier
- Starter reaches four identifiers and all of them cost 1 credit, so 400 credits is 400 messages
- Free accounts have no credit allowance, but they are not model-gated within seven lifetime messages
- The credit rung is derived from a different reference turn than the index figure, and some rungs are pinned by policy
Frequently asked questions
What is the cheapest AI model on Whizi?
By cost per standard answer, the cheapest priced row in the Whizi cost index is Ling from inclusionAI at $0.000053, followed by Nex from Nex Agi, Solar Pro from Upstage and the Qwen fast tier. All of them charge one credit per message. By what you actually pay, they are tied with every other one credit row, including the Whizi house model, the base GPT model on Starter and several DeepSeek, Llama and Gemini rows.
Which plan do I need to reach the cheapest models?
It depends on the row. Several of the very cheapest rows require Powerhouse, because model access falls through to the top tier for any identifier that is not on the four-item Starter list or the 37-item Pro list. The cheap rows reachable on Pro include the DeepSeek, Llama, Mistral, Qwen and Gemini Flash entries marked Pro in the tables on this page.
Does a cheaper model always cost fewer credits?
No. The two figures are measured differently. The index prices a 1,000 in, 500 out standard answer against provider list rates, while the credit rung comes from a 3,000 in, 800 out reference turn rounded up to a legal rung on a fixed ladder, with some rows pinned by policy above or below their arithmetic. That is why a row at $0.0015 per answer can charge one credit while a row at $0.000825 charges two.
What happens to the credit if an answer fails partway through?
It comes back. A generation that fails or is aborted is refunded by deleting its usage row, so the balance returns to where it was. The refund is keyed to a server-minted request id rather than anything the client supplies, so applying the same refund twice restores exactly one turn and not two. One thing that is not refunded: a turn rejected for being out of quota still spends its rate-limit tokens, otherwise an out-of-credit client could retry without limit.
Can I try the cheap models without a subscription?
Yes, inside the seven message lifetime allowance described above, which applies to any text model. That allowance never resets, and claiming a guest account carries its usage into the registered account, so starting as a guest and signing up afterwards does not hand back the seven.