API pricing snapshots, not Copilot availability. Previous models retain historical rates. Gemini 3.8 Flash promotional rates end December 31, 2026; DeepSeek rates shown are peak, uncached prices.
Claude Opus 5.5
CurrentAnthropic · version 5.5
claude-opus-5-5
Complex writing, coding agents, and knowledge work
- Strength
- Current Opus with adaptive reasoning and long context
- Weakness
- More expensive than Sonnet for routine work
- Context and pricing notes
- 1M-token context; up to 128K output tokens.
- How versions differ
- $4 / $20 per million tokens, down from Opus 5 at $5 / $25.
Typical coding chat: $0.08
Reading what you send: about $4 each time it reads a million tokens.
Writing the reply: about $20 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost HighSpeed MediumContext Very long
GPT-6 Sol
CurrentOpenAI · version Sol
gpt-6-sol
Daily coding, tool use, and general reasoning
- Strength
- Balances capability and cost in the GPT-6 lineup
- Weakness
- Astra is the stronger option for the hardest tasks
- Context and pricing notes
- 1.05M-token context; base rates shown, extra-long prompts may cost more.
- How versions differ
- $2 / $10 per million tokens; one fifth of Astra's base rates.
Typical coding chat: $0.04
Reading what you send: about $2 each time it reads a million tokens.
Writing the reply: about $10 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed FastContext Very long
GPT-6 Luna
CurrentOpenAI · version Luna
gpt-6-luna
Focused tasks, simple edits, and high-volume processing
- Strength
- Lowest base rates in the GPT-6 lineup
- Weakness
- Validate quality before using it for complex reasoning
- Context and pricing notes
- 1.05M-token context; up to 128K output tokens.
- How versions differ
- $0.10 / $0.50 per million tokens, below GPT-4o mini's base rates.
Typical coding chat: less than a cent
Reading what you send: about $0.1 each time it reads a million tokens.
Writing the reply: about $0.5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Very long
Gemini 3.8 Flash
CurrentGoogle · version 3.8 Flash
gemini-3.8-flash
Long documents, coding, and multi-step tool use
- Strength
- Current stable Flash for agentic and everyday work
- Weakness
- Promotional rates expire December 31, 2026
- Context and pricing notes
- Long-context model. Output pricing includes thinking tokens.
- How versions differ
- $0.75 / $3.75 through 2026; $1.50 / $7.50 from January 1, 2027.
Typical coding chat: $0.02
Reading what you send: about $0.75 each time it reads a million tokens.
Writing the reply: about $3.75 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Very long
Gemini 3.5 Flash-Lite
CurrentGoogle · version 3.5 Flash-Lite
gemini-3.5-flash-lite
High-volume extraction, translation, and simple processing
- Strength
- Low-cost stable Gemini for new projects
- Weakness
- Use Flash or Pro for harder reasoning
- Context and pricing notes
- Long-context processing; output pricing includes thinking tokens.
- How versions differ
- $0.30 / $2.50 per million tokens; a current alternative to 2.5 Flash.
Typical coding chat: < $0.01
Reading what you send: about $0.3 each time it reads a million tokens.
Writing the reply: about $2.5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Very long
DeepSeek V4.1 Flash
CurrentDeepSeek · version V4.1 Flash
deepseek-flash
Budget coding, tool use, and visual inputs
- Strength
- Thinking and non-thinking modes with a 1M-token context
- Weakness
- Peak rates depend on time of day; test reliability on your tasks
- Context and pricing notes
- 1M-token context; use the deepseek-flash API alias.
- How versions differ
- Peak uncached rates: $0.30 / $1.20. Off-peak rates are half.
Typical coding chat: < $0.01
Reading what you send: about $0.3 each time it reads a million tokens.
Writing the reply: about $1.2 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Very long
DeepSeek V4 Pro
CurrentDeepSeek · version V4 Pro 0813
deepseek-v4-pro
More demanding reasoning and coding in the DeepSeek lineup
- Strength
- Long context and thinking mode at moderate rates
- Weakness
- No vision support; more expensive than Flash
- Context and pricing notes
- 1M-token context; current version is DeepSeek-V4-Pro-0813.
- How versions differ
- Peak uncached rates: $1.32 / $3.96. Off-peak rates are half.
Typical coding chat: $0.02
Reading what you send: about $1.32 each time it reads a million tokens.
Writing the reply: about $3.96 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed MediumContext Very long
Qwen3.8 Max
CurrentAlibaba Cloud · version 3.8 Max
qwen3.8-max
Complex reasoning, coding, and long-document work
- Strength
- Current Qwen flagship with thinking, tools, and 1M context
- Weakness
- Higher cost than Plus or Flash; deployment region matters
- Context and pricing notes
- 1M-token context. qwen3.8-max-0902 is also listed as a snapshot.
- How versions differ
- International rates: $2 / $6 per million tokens up to 1M input.
Typical coding chat: $0.03
Reading what you send: about $2 each time it reads a million tokens.
Writing the reply: about $6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed MediumContext Very long
Qwen3.7 Plus
CurrentAlibaba Cloud · version 3.7 Plus
qwen3.7-plus
Balanced coding, office work, and tool-calling agents
- Strength
- Provider-recommended balance of capability and cost
- Weakness
- Higher price tier above 256K input; promotional discounts vary
- Context and pricing notes
- 1M-token context; thinking and non-thinking output have the same listed rate.
- How versions differ
- International list rates: $0.40 / $1.60 up to 256K; $1.20 / $4.80 above. Promotions excluded.
Typical coding chat: < $0.01
Reading what you send: about $0.4 each time it reads a million tokens.
Writing the reply: about $1.6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Very long
Qwen3.8 Flash
CurrentAlibaba Cloud · version 3.8 Flash
qwen3.8-flash
High-volume processing, summaries, and everyday questions
- Strength
- Low international rates with 1M context and tool calling
- Weakness
- Compare quality with Plus or Max on difficult reasoning
- Context and pricing notes
- 1M-token context. Confirm availability and pricing in your deployment region.
- How versions differ
- International rates: $0.15 / $0.47 per million tokens up to 1M input.
Typical coding chat: less than a cent
Reading what you send: about $0.15 each time it reads a million tokens.
Writing the reply: about $0.47 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Very long
Grok 4.7
CurrentxAI · version 4.7
grok-4.7
General chat, coding, and tool-assisted research
- Strength
- Current xAI model for text and code with 500K context
- Weakness
- Live information requires search tools; long prompts cost more
- Context and pricing notes
- 500K-token context. Higher-tier rates apply to the whole request.
- How versions differ
- $2 / $6 per million tokens below 200K input; $4 / $12 at or above 200K.
Typical coding chat: $0.03
Reading what you send: about $2 each time it reads a million tokens.
Writing the reply: about $6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed MediumContext Very long
Kimi K3
CurrentMoonshot AI · version K3
kimi-k3
Long-context reasoning and multi-step coding work
- Strength
- Current Kimi with a 1,048,576-token context window
- Weakness
- Cache writes are billed separately from base input and output
- Context and pricing notes
- About 1M context. API cache writes: $3/M for 5min or $6/M for 1h.
- How versions differ
- $3 / $15 per million tokens. Cached input is $0.30; cache-write rates depend on TTL.
Typical coding chat: $0.06
Reading what you send: about $3 each time it reads a million tokens.
Writing the reply: about $15 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed MediumContext Very long
Mistral Medium 3.5
CurrentMistral AI · version 3.5
mistral-medium-3-5
Agentic coding, document questions, and multimodal work
- Strength
- Tool calling and structured outputs; open weights under Modified MIT
- Weakness
- Smaller context than 1M-token alternatives; review license terms
- Context and pricing notes
- 256K-token context. Hosted API pricing is separate from self-hosting costs.
- How versions differ
- $1.50 / $7.50 per million tokens; Small 4 is the lower-cost Mistral option.
Typical coding chat: $0.03
Reading what you send: about $1.5 each time it reads a million tokens.
Writing the reply: about $7.5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed MediumContext Very long
Mistral Small 4
CurrentMistral AI · version 4
mistral-small-2603
Budget instruction following, reasoning, and coding
- Strength
- Hybrid model with tool calling and Apache 2.0 open weights
- Weakness
- Test hard tasks against Medium before choosing solely on price
- Context and pricing notes
- 256K-token context; dated API snapshot mistral-small-2603.
- How versions differ
- $0.15 / $0.60 per million tokens; much lower base rates than Medium 3.5.
Typical coding chat: less than a cent
Reading what you send: about $0.15 each time it reads a million tokens.
Writing the reply: about $0.6 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Very long
Claude Fable 5.1
CurrentAnthropic · version 5.1
claude-fable-5-1
Long unattended coding and knowledge-work agents
- Strength
- Built for long-horizon autonomous tasks in Copilot
- Weakness
- Expensive. Admins often must enable it. Data-retention rules differ from other Claude models.
- Context and pricing notes
- Meant for long agent runs. Don’t dump trivia into it.
- How versions differ
- Costs like a flagship ($10 / $50). Use for multi-step agents, not daily chat. Sonnet 5 is cheaper for normal coding.
Typical coding chat: $0.21
Reading what you send: about $10 each time it reads a million tokens.
Writing the reply: about $50 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost HighSpeed SlowContext Very long
Claude Sonnet 5
CurrentAnthropic · version 5
claude-sonnet-5
Daily coding, writing, and planning at a lower Sonnet price
- Strength
- Newest Sonnet, cheaper than 4.5 / 4.6
- Weakness
- If a company still requires Sonnet 4.6, you cannot swap silently
- Context and pricing notes
- Excellent memory for most projects.
- How versions differ
- $2 to read / $10 to write per million tokens — less than Sonnet 4.6’s $3 / $15.
Typical coding chat: $0.04
Reading what you send: about $2 each time it reads a million tokens.
Writing the reply: about $10 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed FastContext Very long
Claude Haiku 4.5
CurrentAnthropic · version 4.5
claude-haiku-4-5
Quick replies, light edits, cheap Claude volume
- Strength
- Fast and inexpensive, still a Claude
- Weakness
- Weaker on hard reasoning than Sonnet
- Context and pricing notes
- Fine for short threads. Don’t dump huge files.
- How versions differ
- Newer than retired Haiku 3.5. Slightly higher list price, stronger quality.
Typical coding chat: $0.02
Reading what you send: about $1 each time it reads a million tokens.
Writing the reply: about $5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost LowSpeed FastContext Good
GPT-6 Astra
CurrentOpenAI · version Astra
gpt-6-astra
Long coding agents and hard multi-step work
- Strength
- Newest OpenAI flagship. Plans, checks, and finishes agent tasks with fewer loops
- Weakness
- Very expensive in Copilot credits. Not on Copilot Free/Pro in all rollouts
- Context and pricing notes
- Huge context, but long prompts cost extra after ~272k tokens.
- How versions differ
- $10 / $50 per million tokens at base rates. GPT-6 Sol costs $2 / $10; choose Astra when its extra capability justifies the cost.
Typical coding chat: $0.21
Reading what you send: about $10 each time it reads a million tokens.
Writing the reply: about $50 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost HighSpeed MediumContext Very long
Gemini 3.1 Pro Preview
CurrentGoogle · version 3.1 Pro
gemini-3.1-pro-preview
Long documents that still need careful answers
- Strength
- Newer Pro with huge context
- Weakness
- Overkill for short chats
- Context and pricing notes
- Base rates up to 200K input tokens; above that, $4 / $18 per million.
- How versions differ
- Use 3.8 Flash for simpler questions. Preview availability and limits can change.
Typical coding chat: $0.05
Reading what you send: about $2 each time it reads a million tokens.
Writing the reply: about $12 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cost MediumSpeed MediumContext Very long