LLM API pricing statistics 2026: output costs 4x input, and the median model charges $1.60 per million

In brief

As of October 2026, the median LLM API charges $0.40 per million input tokens and $1.60 per million output tokens. Output costs 4x input at the median, across 335 paid text models Whizi counted on OpenRouter on 6 October. Twenty of the 100 models in Whizi's 20 August 2026 snapshot showed a new OpenRouter price 47 days later.

Steel coin stacks falling from one tall stack to many short ones, with the OpenAI, Claude and Gemini logos on top

The median LLM API charges $1.60 per million output tokens

On OpenRouter on 6 October 2026, half of all paid models charge between $0.15 and $1.25 per million input tokens, and between $0.50 and $6 per million output. The top tenth charges $3 or more for input and $14 or more for output.

StatisticFigureSourceDate
Median input price, paid text models$0.40 per million tokensWhizi count of the OpenRouter model list6 Oct 2026
Median output price, paid text models$1.60 per million tokensWhizi count of the OpenRouter model list6 Oct 2026
Output price as a multiple of input, median4xWhizi count of the OpenRouter model list6 Oct 2026
Models that charge more for output than input95.8% of 335Whizi count of the OpenRouter model list6 Oct 2026
Median cost of one standard answer (1,000 tokens in, 500 out)$0.00125Whizi count of the OpenRouter model list6 Oct 2026
Cheapest to priciest standard answer$0.000034 to $0.45, a 13,235x spreadWhizi count of the OpenRouter model list6 Oct 2026
Cheapest to priciest across the 126 models Whizi tracks2,000xWhizi AI Model Cost Index20 Aug 2026
Indexed models listed at a new price or delisted within 47 days29 of 100Whizi index against the OpenRouter list20 Aug to 6 Oct 2026
Batch endpoints priced at exactly half the normal rate66 of 73Whizi count of the OpenRouter model list6 Oct 2026
Cached input price as a share of normal input, median12%Whizi count of the OpenRouter model list6 Oct 2026
Annual price fall for a fixed level of capability9x to 900xEpoch AIMar 2025
Fall in inference cost for GPT-3.5 level performanceover 280x, Nov 2022 to Oct 2024Stanford AI Index 2025Apr 2025

Prices at the top of the market didn't fall. Anthropic cut the output price of Claude Opus from $75 to $25 per million tokens with Opus 4.5 in November 2025, and then to $20 with Opus 5.5 in September 2026. But it put a new tier above Opus. Fable 5 arrived in June 2026 at $50, and Fable 5.1 kept that price. OpenAI's priciest standard GPT climbed from $10 with GPT-5 in August 2025 to $50 with GPT-6 Astra in September 2026.

Both labs' top models now charge $50 per million output tokens, after Anthropic cut Opus from $75 to $20
$0$10$20$30$40$50$60$70per million output tokensClaude Opus$75$20Aug 2025Sep 2026Opus 4.1 to 5.5Anthropic top tier$75$50Aug 2025Sep 2026Opus 4.1 to Fable 5.1Top standard GPT$10$50Aug 2025Sep 2026GPT-5 to GPT-6 Astra

Each row compares the list price in August 2025 with the list price in September 2026 for the same slot in a lab's lineup. Claude Fable 5.1 and GPT-6 Astra both charge $10 in and $50 out. So the priciest standard models from the two biggest labs now cost the same.

How much does an LLM API cost per answer?

A typical answer from an LLM API costs about an eighth of a cent. Whizi prices one standard answer as 1,000 tokens in and 500 tokens out, and the median for the 335 paid text models on OpenRouter was $0.00125 on 6 October 2026. That's $1.25 for a thousand of them.

These are the flagship list prices that each lab put on its own pricing page, as we read them on 6 October 2026, in US dollars per million tokens:

ModelInputOutputCached inputOne standard answer
GPT-6 Astra (OpenAI)$10.00$50.00$1.00$0.035
GPT-5.5 (OpenAI)$5.00$30.00$0.50$0.020
GPT-6.1 Sol (OpenAI)$2.00$10.00$0.10$0.007
GPT-6 Luna (OpenAI)$0.10$0.50$0.01$0.00035
Claude Fable 5.1 (Anthropic)$10.00$50.00$0.25$0.035
Claude Opus 5.5 (Anthropic)$4.00$20.00$0.20$0.014
Claude Sonnet 5.5 (Anthropic)$2.00$10.00$0.20$0.007
Claude Haiku 4.5 (Anthropic)$1.00$5.00$0.10$0.0035
Gemini 3.1 Pro Preview (Google, up to 200K tokens)$2.00$12.00$0.20$0.008
Gemini 3.8 Flash (Google, to 31 Dec 2026)$0.75$3.75$0.075$0.0026
Gemini 3.1 Flash-Lite (Google)$0.25$1.50$0.025$0.001
DeepSeek V4 Pro (DeepSeek, peak hours)$1.32$3.96$0.044$0.0033

The gap is wide even within one lab. An answer from GPT-6 Astra costs 100 times as much as one from GPT-6 Luna, and Claude Fable 5.1 costs 10 times as much as Claude Haiku 4.5. Gemini 3.1 Pro Preview also gets dearer once a prompt runs past 200,000 tokens, at $4 in and $18 out, as Google's Gemini API pricing page shows.

Whizi's AI Model Cost Index runs the same sum for each of the 126 models that Whizi tracks. On its 20 August 2026 snapshot, the cheapest answer was from Ling-3.0-flash at $0.000053, and the priciest was from Claude Opus 4.7 in its Fast mode at $0.105. That's a spread of 2,000x. The model in the middle, Kimi K2 Thinking, cost $0.00185 an answer. If you want two vendors side by side with their quality scores, read ChatGPT API vs Claude API pricing.

Are LLM API prices going down?

Yes, for the same level of capability, and they're falling fast. Epoch AI found in March 2025 that the price of reaching a given benchmark score fell by between 9x and 900x a year, depending on which task it was. For the score GPT-4 got on PhD level science questions, the price fell 40x a year, according to Epoch AI's analysis.

a16z's Guido Appenzeller put it in one line in November 2024: "For an LLM of equivalent performance, the cost is decreasing by 10x every year." In his example, GPT-3 cost $60 per million tokens in November 2021, and Llama 3.2 3B matched its MMLU score for $0.06. Stanford's 2025 AI Index found that the cost of GPT-3.5 level performance fell by more than 280 times between November 2022 and October 2024.

That doesn't mean new models are cheap. The 84 text models that OpenRouter listed in the 90 days up to 6 October 2026 have a median output price of $2.50 per million tokens. The 132 it listed more than a year before that sit at $1.01. Old capability gets cheaper. New capability starts high. Google has even put a date on a rise. Its pricing page lists Gemini 3.6, 3.7 and 3.8 Flash at $0.75 in and $3.75 out through 31 December 2026, and then at $1.50 and $7.50 from 1 January 2027 on.

Listed prices also move after a model is out, though not always because a lab moved them. Of the 100 models in Whizi's 20 August 2026 snapshot, 9 had left OpenRouter by 6 October, and 20 of the rest showed a new price: 11 lower and 9 higher at our 05:49 UTC pull. NVIDIA's Nemotron 3 Ultra fell from $0.60 in and $3.60 out to $0.50 and $2.20. Some moves are about how OpenRouter reports a price: its rate for DeepSeek V4 Pro follows DeepSeek's clock, showing the off-peak $0.66 in and $1.98 out at our 05:49 UTC pull and the peak $1.32 and $3.96 after 06:00. Gemini 3.7 Flash went the other way, from $0.375 and $1.875 (Google's batch rate) to the standard $0.75 and $3.75. For open models, the price also follows the hosts OpenRouter sends traffic to.

Why do output tokens cost more than input tokens?

Output costs more because the model writes it one token at a time, while it reads the whole prompt in one parallel pass. On 6 October 2026, 95.8% of the 335 paid text models on OpenRouter charged more for output than for input. The median charged 4x. The middle half ran from 3x to 5x.

The flagships all follow the pattern. Claude Opus 5.5 and GPT-6 Astra charge 5x as much for output, GPT-5.5 and Gemini 3.1 Pro Preview charge 6x, and DeepSeek V4 Pro charges 3x. So an app that writes long answers costs more than one that reads long files, token for token. If you're weighing that against a plan, ChatGPT API vs subscription works out where the API passes $20 a month, and AI subscription statistics covers what people pay for the plans themselves.

What is the cheapest LLM API in October 2026?

The cheapest paid LLM API on OpenRouter on 6 October 2026 is Mistral Nemo, at $0.019 per million tokens in and $0.03 per million out. That's $0.000034 for a standard answer, or 3.4 cents for a thousand of them. Of the three biggest labs, Google and OpenAI sell the cheapest models. Gemini 2.5 Flash-Lite costs $0.10 in and $0.40 out, and GPT-6 Luna costs $0.10 in and $0.50 out, so both come to about $0.0003 an answer. The cheapest current model from Anthropic is Claude Haiku 4.5, at $1 in and $5 out.

There are free models, but not many of them. Only 17 of the 352 text endpoints on OpenRouter were free on 6 October 2026, which is 4.8%, and 16 of those were free variants of a model. Google's Gemini API also has a free tier on most of its models.

At the top, OpenAI's o1-pro still lists at $150 in and $600 out, which is $0.45 an answer and 13,235 times the price of Mistral Nemo. GPT-5.5 Pro and GPT-5.4 Pro charge $30 in and $180 out. Across all of the 335 paid models, 37% charge less than $1 per million output tokens, and 19% charge $10 or more.

How much do batch and caching cut an LLM API bill?

Batch processing cuts the price in half at each of the big labs. Anthropic, OpenAI and Google all charge 50% less for jobs that can wait. On 6 October 2026, 66 of the 73 batch endpoints on OpenRouter were at exactly half of the normal rate.

Caching cuts the input side even more. Across the 214 models that list a price for cache reads, the median cached token costs 12% of what a normal input token does. The flagships go further than that. Claude Opus 5.5 and GPT-6.1 Sol take 95% off cached input, and Claude Fable 5.1 takes off 97.5%. GPT-6 Astra and Gemini 3.1 Pro Preview take off 90%.

DeepSeek adds a clock. Its API charges half price outside peak hours, which DeepSeek sets at 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, per DeepSeek's pricing page. V4 Pro drops from $1.32 in and $3.96 out to $0.66 and $1.98.

How we compiled these numbers

Whizi Research pulled the public model list from OpenRouter (api/v1/models, which needs no key) at 05:49 UTC on 6 October 2026. We kept the 449 entries that output only text. Then we took out 24 routers and meta endpoints, as well as each free or batch variant of a model. That left 336 models from 52 providers, and 335 of them were paid. For the percentiles we used linear interpolation.

For the 47 day comparison we matched the 100 snapshot rows to the 6 October list by OpenRouter id. Three of the 9 that are gone have successors under new ids: Nex N2 Mini and Pro became Nex N2.5 Mini and Pro, and Qwen3.8 Max became qwen3.8-max-0902. At least 6 of the 20 changes are listing artifacts, such as a batch, off-peak or cached rate shown as the main price, rather than a lab repricing a model.

A standard answer is 1,000 tokens in and 500 out at list price. It's the same sum that's behind the Whizi AI Model Cost Index, and all 126 of its rows are free to download from the cost index page. Neither of these figures counts caching, batch or any negotiated discount. We took the flagship prices from the pricing pages of OpenAI, Anthropic, Google and DeepSeek on the same day.

We left out any figure that we could only trace to a stats aggregator. The price OpenRouter shows for an open model depends on the hosts behind it, and the date it lists a model stands in for the day that model came out. Trackers that weight their sample toward flagship models report higher medians than ours. So when you quote a median, quote the sample with it.

How to cite these LLM API pricing statistics

You're free to use every figure on this page, as long as you link back to it. A plain credit line like this one works:

```text

Source: Whizi Research, "LLM API pricing statistics 2026", whizi.io/resources/llm-api-pricing-statistics, updated October 2026.

```

There are also other pages from the same research batch, which went live on the same day: AI chatbot market share, ChatGPT statistics and AI hallucination statistics.

Workflow checklist
  • Check the date: 29 of 100 indexed models showed a new price or left OpenRouter in 47 days
  • Say whether a figure is input, output or a blended answer
  • Name the sample, because a median over 335 models is lower than one over flagships
  • Note whether batch, caching or off-peak discounts are in the figure
  • For a trend claim, say whether capability was held fixed
Common questions

Frequently asked questions

How much does an LLM API cost per million tokens?

The median paid LLM API charged $0.40 per million input tokens and $1.60 per million output tokens on 6 October 2026, across the 335 text models on OpenRouter. The flagships cost more than that. Claude Opus 5.5 lists at $4 in and $20 out, and GPT-6 Astra at $10 in and $50 out.

Are LLM API prices dropping in 2026?

Yes, for the same level of capability: Epoch AI measured falls of 9x to 900x a year. But not at the very top. OpenAI's priciest standard GPT went from $10 to $50 per million output tokens between August 2025 and September 2026, and Anthropic's Fable tier also charges $50, even though Anthropic cut Claude Opus from $75 to $20.

Why is output more expensive than input in LLM APIs?

Output costs more because the model generates it one token at a time, while it reads a prompt in one parallel pass. On 6 October 2026 the median paid model on OpenRouter charged 4x as much for output, and 95.8% charged more for output than input.

Is there a free LLM API?

Yes, a few. On 6 October 2026, 17 of 352 text endpoints on OpenRouter were free, about 4.8%. Google's Gemini API also offers a free tier on most of its models, with paid rates once you move past it.

Still have a question?

Type it here. After you sign up, Whizi answers it first thing.