LatestClaude Opus 5.5See versions →Updated
betterllmsBETA
¢Credits→

Model cards

Compare at a glance.

Each card names the exact version (Sonnet 4.5 vs 4.6 vs 5), what it’s best at, and cost in dollars for a normal coding chat — plus the old “$/M in · $/M out” line explained in English.

API pricing snapshots, not Copilot availability. Previous models retain historical rates. Gemini 3.8 Flash promotional rates end December 31, 2026; DeepSeek rates shown are peak, uncached prices.

Claude Opus 5.5

Current

Anthropic · version 5.5

claude-opus-5-5

Complex writing, coding agents, and knowledge work

Strength
Current Opus with adaptive reasoning and long context
Weakness
More expensive than Sonnet for routine work
Context and pricing notes
1M-token context; up to 128K output tokens.
How versions differ
$4 / $20 per million tokens, down from Opus 5 at $5 / $25.

Typical coding chat: $0.08

Reading what you send: about $4 each time it reads a million tokens.

Writing the reply: about $20 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost HighSpeed MediumContext Very long

GPT-6 Sol

Current

OpenAI · version Sol

gpt-6-sol

Daily coding, tool use, and general reasoning

Strength
Balances capability and cost in the GPT-6 lineup
Weakness
Astra is the stronger option for the hardest tasks
Context and pricing notes
1.05M-token context; base rates shown, extra-long prompts may cost more.
How versions differ
$2 / $10 per million tokens; one fifth of Astra's base rates.

Typical coding chat: $0.04

Reading what you send: about $2 each time it reads a million tokens.

Writing the reply: about $10 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed FastContext Very long

GPT-6 Luna

Current

OpenAI · version Luna

gpt-6-luna

Focused tasks, simple edits, and high-volume processing

Strength
Lowest base rates in the GPT-6 lineup
Weakness
Validate quality before using it for complex reasoning
Context and pricing notes
1.05M-token context; up to 128K output tokens.
How versions differ
$0.10 / $0.50 per million tokens, below GPT-4o mini's base rates.

Typical coding chat: less than a cent

Reading what you send: about $0.1 each time it reads a million tokens.

Writing the reply: about $0.5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Very long

Gemini 3.8 Flash

Current

Google · version 3.8 Flash

gemini-3.8-flash

Long documents, coding, and multi-step tool use

Strength
Current stable Flash for agentic and everyday work
Weakness
Promotional rates expire December 31, 2026
Context and pricing notes
Long-context model. Output pricing includes thinking tokens.
How versions differ
$0.75 / $3.75 through 2026; $1.50 / $7.50 from January 1, 2027.

Typical coding chat: $0.02

Reading what you send: about $0.75 each time it reads a million tokens.

Writing the reply: about $3.75 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Very long

Gemini 3.5 Flash-Lite

Current

Google · version 3.5 Flash-Lite

gemini-3.5-flash-lite

High-volume extraction, translation, and simple processing

Strength
Low-cost stable Gemini for new projects
Weakness
Use Flash or Pro for harder reasoning
Context and pricing notes
Long-context processing; output pricing includes thinking tokens.
How versions differ
$0.30 / $2.50 per million tokens; a current alternative to 2.5 Flash.

Typical coding chat: < $0.01

Reading what you send: about $0.3 each time it reads a million tokens.

Writing the reply: about $2.5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Very long

DeepSeek V4.1 Flash

Current

DeepSeek · version V4.1 Flash

deepseek-flash

Budget coding, tool use, and visual inputs

Strength
Thinking and non-thinking modes with a 1M-token context
Weakness
Peak rates depend on time of day; test reliability on your tasks
Context and pricing notes
1M-token context; use the deepseek-flash API alias.
How versions differ
Peak uncached rates: $0.30 / $1.20. Off-peak rates are half.

Typical coding chat: < $0.01

Reading what you send: about $0.3 each time it reads a million tokens.

Writing the reply: about $1.2 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Very long

DeepSeek V4 Pro

Current

DeepSeek · version V4 Pro 0813

deepseek-v4-pro

More demanding reasoning and coding in the DeepSeek lineup

Strength
Long context and thinking mode at moderate rates
Weakness
No vision support; more expensive than Flash
Context and pricing notes
1M-token context; current version is DeepSeek-V4-Pro-0813.
How versions differ
Peak uncached rates: $1.32 / $3.96. Off-peak rates are half.

Typical coding chat: $0.02

Reading what you send: about $1.32 each time it reads a million tokens.

Writing the reply: about $3.96 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed MediumContext Very long

Qwen3.8 Max

Current

Alibaba Cloud · version 3.8 Max

qwen3.8-max

Complex reasoning, coding, and long-document work

Strength
Current Qwen flagship with thinking, tools, and 1M context
Weakness
Higher cost than Plus or Flash; deployment region matters
Context and pricing notes
1M-token context. qwen3.8-max-0902 is also listed as a snapshot.
How versions differ
International rates: $2 / $6 per million tokens up to 1M input.

Typical coding chat: $0.03

Reading what you send: about $2 each time it reads a million tokens.

Writing the reply: about $6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed MediumContext Very long

Qwen3.7 Plus

Current

Alibaba Cloud · version 3.7 Plus

qwen3.7-plus

Balanced coding, office work, and tool-calling agents

Strength
Provider-recommended balance of capability and cost
Weakness
Higher price tier above 256K input; promotional discounts vary
Context and pricing notes
1M-token context; thinking and non-thinking output have the same listed rate.
How versions differ
International list rates: $0.40 / $1.60 up to 256K; $1.20 / $4.80 above. Promotions excluded.

Typical coding chat: < $0.01

Reading what you send: about $0.4 each time it reads a million tokens.

Writing the reply: about $1.6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Very long

Qwen3.8 Flash

Current

Alibaba Cloud · version 3.8 Flash

qwen3.8-flash

High-volume processing, summaries, and everyday questions

Strength
Low international rates with 1M context and tool calling
Weakness
Compare quality with Plus or Max on difficult reasoning
Context and pricing notes
1M-token context. Confirm availability and pricing in your deployment region.
How versions differ
International rates: $0.15 / $0.47 per million tokens up to 1M input.

Typical coding chat: less than a cent

Reading what you send: about $0.15 each time it reads a million tokens.

Writing the reply: about $0.47 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Very long

Grok 4.7

Current

xAI · version 4.7

grok-4.7

General chat, coding, and tool-assisted research

Strength
Current xAI model for text and code with 500K context
Weakness
Live information requires search tools; long prompts cost more
Context and pricing notes
500K-token context. Higher-tier rates apply to the whole request.
How versions differ
$2 / $6 per million tokens below 200K input; $4 / $12 at or above 200K.

Typical coding chat: $0.03

Reading what you send: about $2 each time it reads a million tokens.

Writing the reply: about $6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed MediumContext Very long

Kimi K3

Current

Moonshot AI · version K3

kimi-k3

Long-context reasoning and multi-step coding work

Strength
Current Kimi with a 1,048,576-token context window
Weakness
Cache writes are billed separately from base input and output
Context and pricing notes
About 1M context. API cache writes: $3/M for 5min or $6/M for 1h.
How versions differ
$3 / $15 per million tokens. Cached input is $0.30; cache-write rates depend on TTL.

Typical coding chat: $0.06

Reading what you send: about $3 each time it reads a million tokens.

Writing the reply: about $15 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed MediumContext Very long

Mistral Medium 3.5

Current

Mistral AI · version 3.5

mistral-medium-3-5

Agentic coding, document questions, and multimodal work

Strength
Tool calling and structured outputs; open weights under Modified MIT
Weakness
Smaller context than 1M-token alternatives; review license terms
Context and pricing notes
256K-token context. Hosted API pricing is separate from self-hosting costs.
How versions differ
$1.50 / $7.50 per million tokens; Small 4 is the lower-cost Mistral option.

Typical coding chat: $0.03

Reading what you send: about $1.5 each time it reads a million tokens.

Writing the reply: about $7.5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed MediumContext Very long

Mistral Small 4

Current

Mistral AI · version 4

mistral-small-2603

Budget instruction following, reasoning, and coding

Strength
Hybrid model with tool calling and Apache 2.0 open weights
Weakness
Test hard tasks against Medium before choosing solely on price
Context and pricing notes
256K-token context; dated API snapshot mistral-small-2603.
How versions differ
$0.15 / $0.60 per million tokens; much lower base rates than Medium 3.5.

Typical coding chat: less than a cent

Reading what you send: about $0.15 each time it reads a million tokens.

Writing the reply: about $0.6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Very long

Claude Fable 5.1

Current

Anthropic · version 5.1

claude-fable-5-1

Long unattended coding and knowledge-work agents

Strength
Built for long-horizon autonomous tasks in Copilot
Weakness
Expensive. Admins often must enable it. Data-retention rules differ from other Claude models.
Context and pricing notes
Meant for long agent runs. Don’t dump trivia into it.
How versions differ
Costs like a flagship ($10 / $50). Use for multi-step agents, not daily chat. Sonnet 5 is cheaper for normal coding.

Typical coding chat: $0.21

Reading what you send: about $10 each time it reads a million tokens.

Writing the reply: about $50 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost HighSpeed SlowContext Very long

Claude Sonnet 5

Current

Anthropic · version 5

claude-sonnet-5

Daily coding, writing, and planning at a lower Sonnet price

Strength
Newest Sonnet, cheaper than 4.5 / 4.6
Weakness
If a company still requires Sonnet 4.6, you cannot swap silently
Context and pricing notes
Excellent memory for most projects.
How versions differ
$2 to read / $10 to write per million tokens — less than Sonnet 4.6’s $3 / $15.

Typical coding chat: $0.04

Reading what you send: about $2 each time it reads a million tokens.

Writing the reply: about $10 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed FastContext Very long

Claude Haiku 4.5

Current

Anthropic · version 4.5

claude-haiku-4-5

Quick replies, light edits, cheap Claude volume

Strength
Fast and inexpensive, still a Claude
Weakness
Weaker on hard reasoning than Sonnet
Context and pricing notes
Fine for short threads. Don’t dump huge files.
How versions differ
Newer than retired Haiku 3.5. Slightly higher list price, stronger quality.

Typical coding chat: $0.02

Reading what you send: about $1 each time it reads a million tokens.

Writing the reply: about $5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost LowSpeed FastContext Good

GPT-6 Astra

Current

OpenAI · version Astra

gpt-6-astra

Long coding agents and hard multi-step work

Strength
Newest OpenAI flagship. Plans, checks, and finishes agent tasks with fewer loops
Weakness
Very expensive in Copilot credits. Not on Copilot Free/Pro in all rollouts
Context and pricing notes
Huge context, but long prompts cost extra after ~272k tokens.
How versions differ
$10 / $50 per million tokens at base rates. GPT-6 Sol costs $2 / $10; choose Astra when its extra capability justifies the cost.

Typical coding chat: $0.21

Reading what you send: about $10 each time it reads a million tokens.

Writing the reply: about $50 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost HighSpeed MediumContext Very long

Gemini 3.1 Pro Preview

Current

Google · version 3.1 Pro

gemini-3.1-pro-preview

Long documents that still need careful answers

Strength
Newer Pro with huge context
Weakness
Overkill for short chats
Context and pricing notes
Base rates up to 200K input tokens; above that, $4 / $18 per million.
How versions differ
Use 3.8 Flash for simpler questions. Preview availability and limits can change.

Typical coding chat: $0.05

Reading what you send: about $2 each time it reads a million tokens.

Writing the reply: about $12 for a million tokens it types back. Replies cost more than reading, because generating text is harder.

Cost MediumSpeed MediumContext Very long