ChatGPTGeminiClaudeGrok40,000+users on the Chrome Web Store4.7/ 5

AI Model Comparison: GPT, Claude, Gemini and Grok Side by Side

Context window, maximum output and API price per million tokens for every current model from OpenAI, Anthropic, Google and xAI, on one table with a verification date. Filter by provider, sort by the column that decides your choice, then price the shortlist on your real token counts.

Comparison free forever | Context gauge free inside ChatGPT, Gemini, Claude and Grok | Premium $9.99/mo or $99 once

Provider

  • 26 models26 with a published price
  • 2MLargest context: Grok 4.20
  • 128KLargest single reply: GPT-6 Astra
  • $0.40Cheapest output per 1M: GPT-5 nano
26 models, provider order. Prices per 1M tokens on the standard rate card.
ModelProvider
GPT-6 AstraOpenAI1.05M tokens128K tokens$10.00$50.00
GPT-5.6 SolOpenAI1.05M tokens128K tokens$4.00$20.00
GPT-5.6 TerraOpenAI1.05M tokens128K tokens$2.00$12.00
GPT-5.6 LunaOpenAI1.05M tokens128K tokens$0.20$1.20
GPT-5.6 CyberRequires approval via OpenAI's Daybreak program; Responses API onlyOpenAI400K tokens128K tokens$12.50$75.00
GPT-5.5OpenAI1.05M tokens128K tokens$5.00$30.00
GPT-5.5 ProOpenAI1.05M tokens128K tokens$30.00$180.00
GPT-5.4OpenAI1.05M tokens128K tokens$2.50$15.00
GPT-5.4 ProOpenAI1.05M tokens128K tokens$30.00$180.00
GPT-5.4 miniOpenAI400K tokens128K tokens$0.75$4.50
GPT-5.4 nanoOpenAI400K tokens128K tokens$0.20$1.25
GPT-5 miniOpenAI400K tokens128K tokens$0.25$2.00
GPT-5 nanoOpenAI400K tokens128K tokens$0.05$0.40
GPT-5OpenAI400K tokens128K tokens$1.25$10.00
GPT-4.1OpenAI1,047,576 tokens32,768 tokens$2.00$8.00
Claude Fable 5.1Anthropic1M tokens128K tokens$10.00$50.00
Claude Opus 5Anthropic1M tokens128K tokens$5.00$25.00
Claude Sonnet 5Anthropic1M tokens128K tokens$2.00$10.00
Claude Haiku 4.5Anthropic200K tokens64K tokens$1.00$5.00
Gemini 3.1 ProGoogle1M tokens64,000 tokens$2.00$12.00
Gemini 3 FlashGoogle1M tokens65,535 tokens$0.50$3.00
Gemini 3.1 Flash-LiteGoogle1M tokens65,536 tokens$0.25$1.50
Grok 4.5xAI500K tokens (API)Bounded by context$2.00$6.00
Grok 4.3xAI1M tokensBounded by context$1.25$2.50
Grok 4.20xAI2M tokens (API); 128K Grok web UIBounded by context$1.25$2.50
Grok Build 0.1xAI256K tokensBounded by context$1.00$2.00

Figures from official provider pricing pages: OpenAI verified September 9, 2026 from developers.openai.com, Google as of March 2026, xAI Grok verified July 2026 from docs.x.ai, and Anthropic Claude verified September 2026 from platform.claude.com. GPT-6 Astra, released September 4, 2026, is OpenAI's current flagship; GPT-5.6 Sol, Terra and Luna, GPT-5.5, GPT-5.4 and their Pro, mini and nano tiers remain listed, and GPT-5, GPT-5 mini, GPT-5 nano and GPT-4.1 stay on the price list although OpenAI now points new work to newer models. OpenAI bills prompts above 272K input tokens at 2x input and 1.5x output on its 1.05M-context models; the table shows the standard rate. Claude Fable 5.1 (released September 1, 2026), Opus 5, Sonnet 5 and Haiku 4.5 are current; Fable 5, Opus 4.8 and Sonnet 4.6 are now legacy. Claude Sonnet 5's $2 / $10 per 1M tokens launched as introductory pricing through August 31, 2026 and is now the standard price: Anthropic cancelled the scheduled rise to $3 / $15. GPT-5.6 Cyber was verified September 2026 from OpenAI's model page; it is an approval-gated API model (Daybreak program), so its row carries an access note. Always confirm current rates at openai.com/pricing, platform.claude.com, ai.google.dev/pricing, and docs.x.ai/pricing. Nothing is uploaded; filtering and sorting run in your browser.

Have a shortlist? Price it on your real token counts, or paste a text into the token counter to see how much of each window it takes.

Reading the table

What each column decides

Three numbers settle most model choices: how much fits in one call, how long one reply can be, and what a million tokens costs each way. They are independent, and the largest window is not the largest reply.

Every row is a current model from the provider's own documentation, with its API figure first. Where the chat product exposes a smaller window than the API, as Grok does, both are stated, because the number that applies to you depends on where you use the model.

The table never carries an estimate. A model whose vendor has not published a list price on this page keeps its limits and shows a pointer to the vendor's rate card instead, and the provenance note under the table states when each published figure was last checked.

The four columns

How to read them:

  • Context window: the total tokens the model can hold at once, your prompt, the conversation so far, any attached files and the reply. Where the API and the chat product differ, both are stated.
  • Max output: the cap on a single reply. A 1M context model can still stop at 32K tokens in one answer, so long generations need a model with a large output limit, not just a large window.
  • Input and output price: per million tokens on the provider's standard rate card. Output is the expensive side on every listed provider, 2 to 6 times the input rate.
  • Long-context surcharge: OpenAI bills prompts above 272K input tokens at 2x input and 1.5x output on its 1.05M-context models. The table carries the standard rate, so price a very long prompt at the higher tier.

Sort by any of them from the column heading; the quick facts above the table follow the current filter.

Step by step

How to compare AI models on this page

Four steps, no account, and nothing leaves your browser.

  • Filter by provider

    Show all four vendors, or narrow to OpenAI, Anthropic, Google or xAI to compare one lineup. The counts and quick facts above the table update with the filter.

  • Sort by what you need

    Largest context first, largest output first, or cheapest input or output. Click a column heading or use the sort control; models without a published price sort last on price.

  • Read the quick facts

    The tiles state the largest window, the largest single reply and the cheapest output rate in the current view, so the answer is visible before you read every row.

  • Price your own workload

    Once you have a shortlist, put your real token counts into the cost calculator to see which of those models is cheapest for that exact shape. Nothing is uploaded; both tools run in your browser.

By provider

What stands out in each lineup

The same table, read one vendor at a time, with a link to the full model guide for each.

  • OpenAI

    GPT-6 Astra on top, and every current tier writes 128K

    GPT-6 Astra (September 2026) and the GPT-5.6 trio (Sol, Terra, Luna) share a 1.05M-token context and a 128K output cap, from $10 down to $0.20 per million input tokens. GPT-5.5 and GPT-5.4 keep the same 1.05M window and 128K cap; the mini, nano and GPT-5 tiers run a 400K window.

    ChatGPT models guide
  • Anthropic

    Every Claude 5 model writes up to 128K tokens

    Fable 5.1, Opus 5 and Sonnet 5 all pair a 1M context with a 128K output limit; they differ on price, from $10 to $2 per million input tokens. Haiku 4.5 is the budget tier at 200K context and 64K output.

    Claude models guide
  • Google

    Low rates on a full 1M window

    Gemini 3.1 Flash-Lite lists at $0.25 input and $1.50 output per million and Gemini 3 Flash at $0.50 and $3.00, both on a 1M window. Gemini 3.1 Pro is the flagship at $2.00 and $12.00.

    Gemini models guide
  • xAI

    The largest window, with a smaller one in the chat UI

    Grok 4.20 accepts 2M tokens through the API, the largest listed, while the grok.com chat product caps near 128K. Grok 4.5 is the newest model at 500K, and every Grok row prices output at 2 to 3 times input, the narrowest gap here.

    Grok models guide

In the chat itself

The window is a number here. In the chat, it is a gauge.

This page compares the models. AI Toolbox works inside ChatGPT, Gemini, Claude and Grok, whichever model is selected.

Inside the chat

See how much of the window a conversation has used

The table tells you the size of the window; the Context Meter tells you how much of it is left in the chat you are in. It reads the running token count on ChatGPT, Gemini, Claude and Grok, warns when the window is nearly full, and can summarize the thread into a fresh chat before the oldest turns drop out. The gauge is free on all four platforms.

Plan: Free: the gauge on all four platforms. Premium: hand the context to another platform.

Context Meter panel reading 479k of 500k tokens with a context window is nearly full warning, about 9 messages left, and a Continue in a fresh chat handoff

Switching models

Keep the same prompts whichever model you move to

Comparing models usually means running the same prompt on each. AI Toolbox keeps one prompt library across ChatGPT, Gemini, Claude and Grok: save a prompt once, insert it by typing two slashes on any of the four, and chain steps with variables so the test is identical everywhere. The free plan holds two saved prompts and five from the catalog.

Plan: Free: 2 saved prompts, 5 catalog prompts. Premium: unlimited prompts and chains.

AI Toolbox prompt library inside ChatGPT with saved prompts inserted by typing two slashes and a catalog of ready-made prompts

Use cases

Who compares models before picking one

  • Developers

    • Pick a model by output cap before a long generation
    • Check whether a 1M window is API-only
    • Shortlist by price, then price the workload
    • Compare provider lineups tier by tier
  • Product and finance

    • Read the cheapest published rate at a glance
    • See which vendors publish list prices
    • Compare flagship tiers across providers
    • Set a model policy per use case
  • Researchers

    • Find models that hold a whole corpus
    • Match the window to a document set
    • Sort by output for long syntheses
    • Cite verified rates with dates
  • Writers and analysts

    • Choose a window that fits a book-length draft
    • Avoid models that stop at 32K per reply
    • Keep one prompt across every model
    • Know when a long thread is about to truncate
  • Anyone choosing a plan

    • Compare the models behind each subscription
    • See what the chat UI caps versus the API
    • Understand context versus output
    • Decide whether one provider covers every job
  • Use them all from one place

    Install AI Toolbox for a free context gauge on all four platforms and one prompt library that follows you from model to model.

    Add to Chrome, free

Pricing

This comparison is free. So is a lot of the extension.

The context gauge and two saved prompts are on the free plan. Premium adds unlimited prompts and chains, folders, full-text search and export. Not an AI subscription: you keep using your own accounts and pay the providers directly.

Free

$0forever

  • This comparison, no account
  • Context window gauge on all four platforms
  • 2 saved prompts inserted with two slashes
  • 5 search results per query
Add to Chrome

Premium

$9.99/month

  • Unlimited prompts and prompt chains
  • Cross-platform context handoff
  • Unlimited folders and search
  • Bulk export in every format
Get Premium

All Access LifetimeBest value

$199one-time

  • ChatGPT, Gemini, Claude and Grok included
  • Every future AI platform we add
  • All Premium features unlocked
  • Save $197 vs single Lifetimes
Get All Access

Teams$15/seat/month

All Premium features across the four modules ยท Admin dashboard ยท Team analytics ยท $12/seat/mo billed annually ยท 14-day free trial

See Teams
  • GDPR compliant
  • Secure checkout by Polar
  • 14-day money-back on a first purchase
  • Nothing on this page leaves the browser

Refunds follow the refund policy; renewals are not refundable. Premium unlocks for the email you enter at checkout, so use the same address you sign in to ChatGPT with.

FAQ

Model comparison questions people actually ask

Which AI model has the largest context window?

Grok 4.20 from xAI, at 2 million tokens through the API; its grok.com chat product caps near 128K. GPT-6 Astra, the GPT-5.6 models (Sol, Terra, Luna), GPT-5.5 and GPT-5.4 accept 1.05 million tokens, and GPT-4.1, Claude Fable 5.1, Opus 5 and Sonnet 5, the Gemini 3 family and Grok 4.3 all accept 1 million.

Which AI model is cheapest?

Among models with a published list price, GPT-5 nano at $0.05 input and $0.40 output per million tokens, then GPT-5.6 Luna at $0.20 and $1.20 and GPT-5.4 nano at $0.20 and $1.25, then Gemini 3.1 Flash-Lite at $0.25 and $1.50. Which is cheapest for you depends on the ratio of input to output tokens in your workload, which the cost calculator prices exactly.

What is the difference between context window and max output?

The context window is everything the model can consider at once: your prompt, the conversation history, attached files and the reply it is writing. Max output is the ceiling on one reply. GPT-4.1 has a 1M-token window but stops at 32,768 tokens per answer, while GPT-6 Astra, the GPT-5.6 trio, GPT-5.5, GPT-5.4 and the Claude 5 models can write up to 128K in one reply.

Which model can write the longest single reply?

GPT-6 Astra, GPT-5.6 Sol, Terra and Luna, GPT-5.5, GPT-5.4 and their Pro, mini and nano tiers, and Claude Fable 5.1, Opus 5 and Sonnet 5, all list a 128K-token max output. Gemini 3.1 Pro, Gemini 3 Flash and Gemini 3.1 Flash-Lite are in the 64K to 65K range, Claude Haiku 4.5 at 64K, and GPT-4.1 at 32,768. xAI does not publish a separate output cap for Grok; the reply is bounded by the remaining context.

Do OpenAI prices change with prompt length?

Yes. On its 1.05M-context models (GPT-6 Astra, the GPT-5.6 trio, GPT-5.5 and GPT-5.4), OpenAI bills any request whose input exceeds 272K tokens at 2x the input rate and 1.5x the output rate for the whole request. The table and the cost calculator carry the standard rate, so price a very long prompt at the higher tier.

Does a large context window make a model better?

It makes more fit in one call, which matters for long documents, large codebases and long-running threads. It does not by itself improve reasoning, and a long prompt costs more per call because every input token is billed. Match the window to the job and check the max output separately if you need long replies.

How current is this comparison?

Every figure carries a verification date against the provider's own pricing page: OpenAI verified September 9, 2026 from developers.openai.com, five days after GPT-6 Astra shipped; Google as of March 2026; xAI Grok verified July 2026 from docs.x.ai; and Anthropic verified September 2026 from platform.claude.com, the day after Claude Fable 5.1 shipped. Claude Sonnet 5 stays at $2 and $10 per million tokens: the rise to $3 and $15 that Anthropic had scheduled for September 1, 2026 was cancelled. Confirm on openai.com/pricing, platform.claude.com, ai.google.dev/pricing and docs.x.ai/pricing before budgeting.

Does AI Toolbox work with all of these models?

Yes, whichever model you pick inside ChatGPT, Gemini, Claude or Grok. AI Toolbox is a browser extension that adds folders, full-text search, export, a context gauge and a shared prompt library on top of the accounts you already have; it is not an AI subscription and does not sell model access. The gauge and two saved prompts are free, and Premium is $9.99/month or $99 once per module.

40,000+ users ยท 4.7/5 on the Chrome Web Store

Pick the model here. Keep your work in order there.

Install AI Toolbox for a free context gauge on ChatGPT, Gemini, Claude and Grok, and one prompt library across all four.