Reading note
Don’t pay for a genius to summarize a spreadsheet.
Match the model to the job. Save the expensive one for hard thinking.
The right model for your work. More capable answers, less wasted spend.
Task first
Write, debug, or review code
You need a model that follows instructions closely and still stays affordable for many back-and-forth edits.
Best overall
Anthropic · version 5
claude-sonnet-5
Daily coding, writing, and planning at a lower Sonnet price
Typical coding chat: $0.04
Reading what you send: about $2 each time it reads a million tokens.
Writing the reply: about $10 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Best value
DeepSeek · version V4.1 Flash
deepseek-flash
Budget coding, tool use, and visual inputs
Typical coding chat: < $0.01
Reading what you send: about $0.3 each time it reads a million tokens.
Writing the reply: about $1.2 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Cheapest usable
OpenAI · version Luna
gpt-6-luna
Focused tasks, simple edits, and high-volume processing
Typical coding chat: less than a cent
Reading what you send: about $0.1 each time it reads a million tokens.
Writing the reply: about $0.5 for a million tokens it types back. Replies cost more than reading, because generating text is harder.
Rule of thumb: use the cheapest usable model until the answer quality starts to cost you time. Then step up one level.
Context switch
Switching mid-chat can drop details and change the bill. Pick two models.
Extra cost
Lower bill: a typical coding session goes from about $0.04 to less than a cent.
Risk of losing context
Low
Recommendation
Switch
You keep similar memory and pay less. Paste a short recap of the goal when you switch.