AI model comparison: price, context and what each is good at
Every few weeks a new model claims the top spot. This table stays current so you do not have to follow the news. Prices are per million tokens on the API, in US dollars.
Last updated: 10.09.2026| Model | Released | Context | Max output | Price / 1M tokens | What it is good at |
|---|---|---|---|---|---|
| GPT-6 AstraNEW OpenAI | 03.09.2026 | 1,05M | 128K |
10$ in 50$ out |
The first widely available model that can drive a computer on its own. It takes a goal, researches on the web, uses the screen, writes code and finishes multi step tasks without a human in the loop. It scored 0.86 on the OSWorld 2.0 computer use benchmark, and tasks that took 75 minutes on the previous model drop to around 40 minutes. Access rolls out in stages: Daybreak enterprise customers first, then Pro, Plus, Business and Enterprise, then the API, Azure and AWS Bedrock. It costs twice Opus 5, and fast mode doubles that again. As OpenAI's first model past the critical cybersecurity threshold it refuses work such as vulnerability discovery on standard access. |
| Claude Opus 5 Anthropic | 24.07.2026 | 1M | 128K |
5$ in 25$ out |
Code, multi step agent work and long document analysis. The effort setting lets you scale cost to the weight of the task. Anthropic still points to Fable 5 for the hardest fully autonomous projects. |
| Claude Fable 5.1NEW Anthropic | 01.09.2026 | 1M | 128K |
10$ in 50$ out |
The hardest reasoning and long running autonomous agent tasks. The 5.1 release on 1 September 2026 kept the price but cut cache reads from 1 dollar to 0.25 dollars per million tokens, which Anthropic puts at roughly a quarter off typical workloads and close to half off cache heavy agent work. List price is twice Opus 5. Most work does not justify the gap; outside long cache heavy agent flows, Opus 5 is the better call. |
| Claude Sonnet 5VALUE Anthropic | 09.06.2026 | 1M | 64K |
2$ in 10$ out |
The balanced model for daily work. Good value for writing, analysis and moderate coding. It was scheduled to move to 3 and 15 dollars on 1 September 2026; Anthropic cancelled that increase in August 2026 and the 2 and 10 dollar pricing became permanent. |
| GPT-5.6 Sol OpenAI | 09.07.2026 | 1M | 128K |
5$ in 30$ out |
Broad general reasoning and the most mature tool ecosystem. While Astra rolls out in stages, most corporate work still runs on this model. No longer OpenAI's flagship after GPT-6 Astra, but with broad access and a mature tool ecosystem it is still the first pick for most day to day corporate work. Output costs more than Opus 5. |
| GPT-5.6 Terra OpenAI | 09.07.2026 | 1M | 128K |
2,50$ in 15$ out |
The balanced tier. Enough for most daily work without paying for Sol. Falls behind the flagship on the hardest reasoning tasks. |
| GPT-5.6 LunaVALUE OpenAI | 09.07.2026 | 1M | 128K |
1$ in 6$ out |
Speed and low cost. Built for high volume, simple, repetitive work. Struggles with deep analysis and multi step tasks. |
| Gemini 3.1 Pro Google | 2026 | 2M | 64K |
2$ in 12$ out |
The largest context window available. Huge document sets and multimodal input such as video, audio and PDF. Requests above 200 thousand tokens move to 4 and 18 dollars per million tokens. |
| Gemini 3.6 FlashNEW Google | 21.07.2026 | 1M | 64K |
1,50$ in 7,50$ out |
Unusually strong for the price. It scores above Google's own Pro tier on the independent intelligence index and spends fewer output tokens than 3.5 Flash. Behind flagship models on top tier reasoning. |
| Kimi K3OPEN Moonshot AI | 16.07.2026 | 1M | - | Self hosted | The first open model at 2.8 trillion parameters. It took first place on an independent frontend code arena. The weights are downloadable, so you can run it on your own servers. Running an open model needs servers and a technical team. Independent tests show it leading in specific areas, not across the board. |
A token is roughly three quarters of a word. These are API prices: if you use the chat apps instead, you pay a monthly subscription and these numbers do not apply to you. Model providers change pricing often, so confirm on their own page before committing to one.
Which one for which job?
Benchmark scores rarely answer the only question that matters: which model should I open for the work in front of me right now.
GPT-6 Astra
The first model that drives a browser and a computer end to end. Worth it when the job is "go do this and come back", not when you want a quick answer.
Claude Sonnet 5 or GPT-5.6 Terra
Both handle daily work well and neither makes you pay flagship prices for it.
Gemini 3.1 Pro or Claude Opus 5
Gemini has the widest context at 2M. Opus 5 reads 1M and reasons harder on what it finds.
Claude Opus 5
Strongest results on coding benchmarks and it verifies its own work, which means fewer rounds of fixing.
GPT-5.6 Luna or Gemini 3.6 Flash
When you run thousands of small jobs, per token price decides the bill, not benchmark scores.
Kimi K3
The only model here with downloadable weights. It needs a technical team but your data never leaves your infrastructure.
Frequently asked questions
Which AI model is the best right now?
There is no single best model, only a best fit for a given job. For handing over a whole multi step task to the computer, GPT-6 Astra (3 September 2026) leads. On coding Claude Opus 5 leads. For the widest context window Gemini 3.1 Pro leads. For high volume cheap work GPT-5.6 Luna and Gemini 3.6 Flash lead. Pick by the work you actually do.
What do the prices in this table mean?
They are API prices per million tokens, in US dollars, input and output separately. A token is roughly three quarters of a word. If you use the chat apps rather than the API you pay a monthly subscription instead and these numbers do not apply to you.
Why is output more expensive than input?
Producing text costs the provider far more compute than reading it. That is why work that generates long answers costs much more than work that reads long documents and replies briefly.
How often is this table updated?
Whenever a model launches or a price changes. The last update date is shown at the top of the page. These models move fast, so always confirm pricing on the provider's own page before committing.
Do I need to pay for any of these to get started?
No. ChatGPT, Claude and Gemini all have free tiers that are enough for most everyday work. Paid plans start to make sense once you use AI daily and hit the usage limits.
The model matters less than the workflow
Switching to a better model gains you a few percent. Building a workflow around it changes how your week runs. Start with the free prompt library.
Free prompt library See consulting