
The median LLM API charges $1.60 per million output tokens
On OpenRouter on 6 October 2026, half of all paid models charge between $0.15 and $1.25 per million input tokens, and between $0.50 and $6 per million output. The top tenth charges $3 or more for input and $14 or more for output.
| Statistic | Figure | Source | Date |
|---|---|---|---|
| Median input price, paid text models | $0.40 per million tokens | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Median output price, paid text models | $1.60 per million tokens | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Output price as a multiple of input, median | 4x | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Models that charge more for output than input | 95.8% of 335 | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Median cost of one standard answer (1,000 tokens in, 500 out) | $0.00125 | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Cheapest to priciest standard answer | $0.000034 to $0.45, a 13,235x spread | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Cheapest to priciest across the 126 models Whizi tracks | 2,000x | Whizi AI Model Cost Index | 20 Aug 2026 |
| Indexed models listed at a new price or delisted within 47 days | 29 of 100 | Whizi index against the OpenRouter list | 20 Aug to 6 Oct 2026 |
| Batch endpoints priced at exactly half the normal rate | 66 of 73 | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Cached input price as a share of normal input, median | 12% | Whizi count of the OpenRouter model list | 6 Oct 2026 |
| Annual price fall for a fixed level of capability | 9x to 900x | Epoch AI | Mar 2025 |
| Fall in inference cost for GPT-3.5 level performance | over 280x, Nov 2022 to Oct 2024 | Stanford AI Index 2025 | Apr 2025 |
Prices at the top of the market didn't fall. Anthropic cut the output price of Claude Opus from $75 to $25 per million tokens with Opus 4.5 in November 2025, and then to $20 with Opus 5.5 in September 2026. But it put a new tier above Opus. Fable 5 arrived in June 2026 at $50, and Fable 5.1 kept that price. OpenAI's priciest standard GPT climbed from $10 with GPT-5 in August 2025 to $50 with GPT-6 Astra in September 2026.
Each row compares the list price in August 2025 with the list price in September 2026 for the same slot in a lab's lineup. Claude Fable 5.1 and GPT-6 Astra both charge $10 in and $50 out. So the priciest standard models from the two biggest labs now cost the same.
How much does an LLM API cost per answer?
A typical answer from an LLM API costs about an eighth of a cent. Whizi prices one standard answer as 1,000 tokens in and 500 tokens out, and the median for the 335 paid text models on OpenRouter was $0.00125 on 6 October 2026. That's $1.25 for a thousand of them.
These are the flagship list prices that each lab put on its own pricing page, as we read them on 6 October 2026, in US dollars per million tokens:
| Model | Input | Output | Cached input | One standard answer |
|---|---|---|---|---|
| GPT-6 Astra (OpenAI) | $10.00 | $50.00 | $1.00 | $0.035 |
| GPT-5.5 (OpenAI) | $5.00 | $30.00 | $0.50 | $0.020 |
| GPT-6.1 Sol (OpenAI) | $2.00 | $10.00 | $0.10 | $0.007 |
| GPT-6 Luna (OpenAI) | $0.10 | $0.50 | $0.01 | $0.00035 |
| Claude Fable 5.1 (Anthropic) | $10.00 | $50.00 | $0.25 | $0.035 |
| Claude Opus 5.5 (Anthropic) | $4.00 | $20.00 | $0.20 | $0.014 |
| Claude Sonnet 5.5 (Anthropic) | $2.00 | $10.00 | $0.20 | $0.007 |
| Claude Haiku 4.5 (Anthropic) | $1.00 | $5.00 | $0.10 | $0.0035 |
| Gemini 3.1 Pro Preview (Google, up to 200K tokens) | $2.00 | $12.00 | $0.20 | $0.008 |
| Gemini 3.8 Flash (Google, to 31 Dec 2026) | $0.75 | $3.75 | $0.075 | $0.0026 |
| Gemini 3.1 Flash-Lite (Google) | $0.25 | $1.50 | $0.025 | $0.001 |
| DeepSeek V4 Pro (DeepSeek, peak hours) | $1.32 | $3.96 | $0.044 | $0.0033 |
The gap is wide even within one lab. An answer from GPT-6 Astra costs 100 times as much as one from GPT-6 Luna, and Claude Fable 5.1 costs 10 times as much as Claude Haiku 4.5. Gemini 3.1 Pro Preview also gets dearer once a prompt runs past 200,000 tokens, at $4 in and $18 out, as Google's Gemini API pricing page shows.
Whizi's AI Model Cost Index runs the same sum for each of the 126 models that Whizi tracks. On its 20 August 2026 snapshot, the cheapest answer was from Ling-3.0-flash at $0.000053, and the priciest was from Claude Opus 4.7 in its Fast mode at $0.105. That's a spread of 2,000x. The model in the middle, Kimi K2 Thinking, cost $0.00185 an answer. If you want two vendors side by side with their quality scores, read ChatGPT API vs Claude API pricing.
Are LLM API prices going down?
Yes, for the same level of capability, and they're falling fast. Epoch AI found in March 2025 that the price of reaching a given benchmark score fell by between 9x and 900x a year, depending on which task it was. For the score GPT-4 got on PhD level science questions, the price fell 40x a year, according to Epoch AI's analysis.
a16z's Guido Appenzeller put it in one line in November 2024: "For an LLM of equivalent performance, the cost is decreasing by 10x every year." In his example, GPT-3 cost $60 per million tokens in November 2021, and Llama 3.2 3B matched its MMLU score for $0.06. Stanford's 2025 AI Index found that the cost of GPT-3.5 level performance fell by more than 280 times between November 2022 and October 2024.
That doesn't mean new models are cheap. The 84 text models that OpenRouter listed in the 90 days up to 6 October 2026 have a median output price of $2.50 per million tokens. The 132 it listed more than a year before that sit at $1.01. Old capability gets cheaper. New capability starts high. Google has even put a date on a rise. Its pricing page lists Gemini 3.6, 3.7 and 3.8 Flash at $0.75 in and $3.75 out through 31 December 2026, and then at $1.50 and $7.50 from 1 January 2027 on.
Listed prices also move after a model is out, though not always because a lab moved them. Of the 100 models in Whizi's 20 August 2026 snapshot, 9 had left OpenRouter by 6 October, and 20 of the rest showed a new price: 11 lower and 9 higher at our 05:49 UTC pull. NVIDIA's Nemotron 3 Ultra fell from $0.60 in and $3.60 out to $0.50 and $2.20. Some moves are about how OpenRouter reports a price: its rate for DeepSeek V4 Pro follows DeepSeek's clock, showing the off-peak $0.66 in and $1.98 out at our 05:49 UTC pull and the peak $1.32 and $3.96 after 06:00. Gemini 3.7 Flash went the other way, from $0.375 and $1.875 (Google's batch rate) to the standard $0.75 and $3.75. For open models, the price also follows the hosts OpenRouter sends traffic to.
Why do output tokens cost more than input tokens?
Output costs more because the model writes it one token at a time, while it reads the whole prompt in one parallel pass. On 6 October 2026, 95.8% of the 335 paid text models on OpenRouter charged more for output than for input. The median charged 4x. The middle half ran from 3x to 5x.
The flagships all follow the pattern. Claude Opus 5.5 and GPT-6 Astra charge 5x as much for output, GPT-5.5 and Gemini 3.1 Pro Preview charge 6x, and DeepSeek V4 Pro charges 3x. So an app that writes long answers costs more than one that reads long files, token for token. If you're weighing that against a plan, ChatGPT API vs subscription works out where the API passes $20 a month, and AI subscription statistics covers what people pay for the plans themselves.
What is the cheapest LLM API in October 2026?
The cheapest paid LLM API on OpenRouter on 6 October 2026 is Mistral Nemo, at $0.019 per million tokens in and $0.03 per million out. That's $0.000034 for a standard answer, or 3.4 cents for a thousand of them. Of the three biggest labs, Google and OpenAI sell the cheapest models. Gemini 2.5 Flash-Lite costs $0.10 in and $0.40 out, and GPT-6 Luna costs $0.10 in and $0.50 out, so both come to about $0.0003 an answer. The cheapest current model from Anthropic is Claude Haiku 4.5, at $1 in and $5 out.
There are free models, but not many of them. Only 17 of the 352 text endpoints on OpenRouter were free on 6 October 2026, which is 4.8%, and 16 of those were free variants of a model. Google's Gemini API also has a free tier on most of its models.
At the top, OpenAI's o1-pro still lists at $150 in and $600 out, which is $0.45 an answer and 13,235 times the price of Mistral Nemo. GPT-5.5 Pro and GPT-5.4 Pro charge $30 in and $180 out. Across all of the 335 paid models, 37% charge less than $1 per million output tokens, and 19% charge $10 or more.
How much do batch and caching cut an LLM API bill?
Batch processing cuts the price in half at each of the big labs. Anthropic, OpenAI and Google all charge 50% less for jobs that can wait. On 6 October 2026, 66 of the 73 batch endpoints on OpenRouter were at exactly half of the normal rate.
Caching cuts the input side even more. Across the 214 models that list a price for cache reads, the median cached token costs 12% of what a normal input token does. The flagships go further than that. Claude Opus 5.5 and GPT-6.1 Sol take 95% off cached input, and Claude Fable 5.1 takes off 97.5%. GPT-6 Astra and Gemini 3.1 Pro Preview take off 90%.
DeepSeek adds a clock. Its API charges half price outside peak hours, which DeepSeek sets at 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, per DeepSeek's pricing page. V4 Pro drops from $1.32 in and $3.96 out to $0.66 and $1.98.
How we compiled these numbers
Whizi Research pulled the public model list from OpenRouter (api/v1/models, which needs no key) at 05:49 UTC on 6 October 2026. We kept the 449 entries that output only text. Then we took out 24 routers and meta endpoints, as well as each free or batch variant of a model. That left 336 models from 52 providers, and 335 of them were paid. For the percentiles we used linear interpolation.
For the 47 day comparison we matched the 100 snapshot rows to the 6 October list by OpenRouter id. Three of the 9 that are gone have successors under new ids: Nex N2 Mini and Pro became Nex N2.5 Mini and Pro, and Qwen3.8 Max became qwen3.8-max-0902. At least 6 of the 20 changes are listing artifacts, such as a batch, off-peak or cached rate shown as the main price, rather than a lab repricing a model.
A standard answer is 1,000 tokens in and 500 out at list price. It's the same sum that's behind the Whizi AI Model Cost Index, and all 126 of its rows are free to download from the cost index page. Neither of these figures counts caching, batch or any negotiated discount. We took the flagship prices from the pricing pages of OpenAI, Anthropic, Google and DeepSeek on the same day.
We left out any figure that we could only trace to a stats aggregator. The price OpenRouter shows for an open model depends on the hosts behind it, and the date it lists a model stands in for the day that model came out. Trackers that weight their sample toward flagship models report higher medians than ours. So when you quote a median, quote the sample with it.
How to cite these LLM API pricing statistics
You're free to use every figure on this page, as long as you link back to it. A plain credit line like this one works:
```text
Source: Whizi Research, "LLM API pricing statistics 2026", whizi.io/resources/llm-api-pricing-statistics, updated October 2026.
```
There are also other pages from the same research batch, which went live on the same day: AI chatbot market share, ChatGPT statistics and AI hallucination statistics.
- Check the date: 29 of 100 indexed models showed a new price or left OpenRouter in 47 days
- Say whether a figure is input, output or a blended answer
- Name the sample, because a median over 335 models is lower than one over flagships
- Note whether batch, caching or off-peak discounts are in the figure
- For a trend claim, say whether capability was held fixed
Frequently asked questions
How much does an LLM API cost per million tokens?
The median paid LLM API charged $0.40 per million input tokens and $1.60 per million output tokens on 6 October 2026, across the 335 text models on OpenRouter. The flagships cost more than that. Claude Opus 5.5 lists at $4 in and $20 out, and GPT-6 Astra at $10 in and $50 out.
Are LLM API prices dropping in 2026?
Yes, for the same level of capability: Epoch AI measured falls of 9x to 900x a year. But not at the very top. OpenAI's priciest standard GPT went from $10 to $50 per million output tokens between August 2025 and September 2026, and Anthropic's Fable tier also charges $50, even though Anthropic cut Claude Opus from $75 to $20.
Why is output more expensive than input in LLM APIs?
Output costs more because the model generates it one token at a time, while it reads a prompt in one parallel pass. On 6 October 2026 the median paid model on OpenRouter charged 4x as much for output, and 95.8% charged more for output than input.
Is there a free LLM API?
Yes, a few. On 6 October 2026, 17 of 352 text endpoints on OpenRouter were free, about 4.8%. Google's Gemini API also offers a free tier on most of its models, with paid rates once you move past it.