AI model comparison: price, context and what each is good at
Every few weeks a new model claims the top spot. This table stays current so you do not have to follow the news. Prices are per million tokens on the API, in US dollars.
Last updated: 26.07.2026| Model | Released | Context | Max output | Price / 1M tokens | What it is good at |
|---|---|---|---|---|---|
| Claude Opus 5NEW Anthropic | 24.07.2026 | 1M | 128K |
5$ in 25$ out |
Code, multi step agent work and long document analysis. The effort setting lets you scale cost to the weight of the task. Anthropic still points to Fable 5 for the hardest fully autonomous projects. |
| Claude Fable 5 Anthropic | 09.06.2026 | 1M | 128K |
10$ in 50$ out |
The hardest reasoning and long running autonomous agent tasks. No long context surcharge. Twice the price of Opus 5, which most work does not justify. |
| Claude Sonnet 5VALUE Anthropic | 09.06.2026 | 1M | 64K |
2$ in 10$ out |
The balanced model for daily work. Good value for writing, analysis and moderate coding. Introductory pricing runs to 31 August 2026, after which it moves to 3 and 15 dollars per million tokens. |
| GPT-5.6 Sol OpenAI | 09.07.2026 | 1M | 128K |
5$ in 30$ out |
OpenAI's flagship. Broad general reasoning and the widest tool ecosystem. Output costs more than Opus 5, and the gap grows on work that produces long answers. |
| GPT-5.6 Terra OpenAI | 09.07.2026 | 1M | 128K |
2,50$ in 15$ out |
The balanced tier. Enough for most daily work without paying for Sol. Falls behind the flagship on the hardest reasoning tasks. |
| GPT-5.6 LunaVALUE OpenAI | 09.07.2026 | 1M | 128K |
1$ in 6$ out |
Speed and low cost. Built for high volume, simple, repetitive work. Struggles with deep analysis and multi step tasks. |
| Gemini 3.1 Pro Google | 2026 | 2M | 64K |
2$ in 12$ out |
The largest context window available. Huge document sets and multimodal input such as video, audio and PDF. Requests above 200 thousand tokens move to 4 and 18 dollars per million tokens. |
| Gemini 3.6 FlashNEW Google | 21.07.2026 | 1M | 64K |
1,50$ in 7,50$ out |
Unusually strong for the price. It scores above Google's own Pro tier on the independent intelligence index and spends fewer output tokens than 3.5 Flash. Behind flagship models on top tier reasoning. |
| Kimi K3OPEN Moonshot AI | 16.07.2026 | 1M | - | Self hosted | The first open model at 2.8 trillion parameters. It took first place on an independent frontend code arena. The weights are downloadable, so you can run it on your own servers. Running an open model needs servers and a technical team. Independent tests show it leading in specific areas, not across the board. |
A token is roughly three quarters of a word. These are API prices: if you use the chat apps instead, you pay a monthly subscription and these numbers do not apply to you. Model providers change pricing often, so confirm on their own page before committing to one.
Which one for which job?
Benchmark scores rarely answer the only question that matters: which model should I open for the work in front of me right now.
Claude Sonnet 5 or GPT-5.6 Terra
Both handle daily work well and neither makes you pay flagship prices for it.
Gemini 3.1 Pro or Claude Opus 5
Gemini has the widest context at 2M. Opus 5 reads 1M and reasons harder on what it finds.
Claude Opus 5
Strongest results on coding benchmarks and it verifies its own work, which means fewer rounds of fixing.
GPT-5.6 Luna or Gemini 3.6 Flash
When you run thousands of small jobs, per token price decides the bill, not benchmark scores.
Kimi K3
The only model here with downloadable weights. It needs a technical team but your data never leaves your infrastructure.
Frequently asked questions
Which AI model is the best right now?
There is no single best model, only a best fit for a given job. On coding and multi step agent work Claude Opus 5 leads. For the widest context window Gemini 3.1 Pro leads. For high volume cheap work GPT-5.6 Luna and Gemini 3.6 Flash lead. Pick by the work you actually do.
What do the prices in this table mean?
They are API prices per million tokens, in US dollars, input and output separately. A token is roughly three quarters of a word. If you use the chat apps rather than the API you pay a monthly subscription instead and these numbers do not apply to you.
Why is output more expensive than input?
Producing text costs the provider far more compute than reading it. That is why work that generates long answers costs much more than work that reads long documents and replies briefly.
How often is this table updated?
Whenever a model launches or a price changes. The last update date is shown at the top of the page. These models move fast, so always confirm pricing on the provider's own page before committing.
Do I need to pay for any of these to get started?
No. ChatGPT, Claude and Gemini all have free tiers that are enough for most everyday work. Paid plans start to make sense once you use AI daily and hit the usage limits.
The model matters less than the workflow
Switching to a better model gains you a few percent. Building a workflow around it changes how your week runs. Start with the free prompt library.
Free prompt library See the training