LLM Pricing: The True Cost of Intelligence
Beyond the sticker price — understanding token economics, hidden costs, and how to actually budget for AI at scale. Updated with October 2025 pricing.
Pricing Models
Per-Token Pricing
The standard model: cost per 1M tokens processed.
Important distinction:
- Input tokens: Your prompt + context (cheaper)
- Output tokens: Model’s response (more expensive)
- Cached input: Repeated context (cheapest, ~90% discount)
Current Major Provider Pricing (per 1M tokens, October 2025)
GPT-6 Astra (OpenAI)
- Input: $10.00
- Cached Input: $1.00
- Output: $50.00
GPT-5.6 Sol (OpenAI)
- Input: $2.00
- Cached Input: $0.20
- Output: $10.00
GPT-5.6 Luna (OpenAI)
- Input: $0.20
- Cached Input: $0.02
- Output: $1.20
Claude Fable 5.1 (Anthropic)
- Input: $10.00
- Cached Input: $0.10
- Output: $50.00
Claude 3 Haiku (Anthropic)
- Input: $0.25
- Cached Input: $0.03
- Output: $1.25
Gemini 3.8 Flash (Google)
- Input: $0.75
- Cached Input: $0.075
- Output: $3.75
DeepSeek V4.1 Flash (DeepSeek)
- Input: $0.15-$0.30 (time-based)
- Cached Input: $0.006
- Output: $0.60-$1.20
Ling 3.0 Flash VL (InclusionAI)
- Input: $0.06
- Cached Input: $0.012
- Output: $0.18
Meta Muse Spark 1.3 (Meta)
- Input: $0.10
- Cached Input: $0.025
- Output: $0.20
Mercury 2.5 (Inception)
- Input: $0.04
- Output: $0.15
Hidden Costs
Context Window Tax
Every token in your context window costs money. A 1M context window with 800K of system prompt = 800K tokens × input cost per request.
At scale:
- 1M requests/day × 800K input tokens × $10/1M = $8,000/day just for context
- That’s $2.9M/year in context alone
Token Counting Gotchas
- Whitespace counts as tokens
- Special characters count as tokens
- Images count as tokens (base64 encoded, very expensive)
- Tool calls count as tokens (both directions)
- System prompts are repeated every request
The “Cheap Model” Trap
A cheaper model that requires 3x more tokens to achieve the same result isn’t actually cheaper.
Example:
- GPT-6 Astra: $10/1M input, generates 500 tokens response
- DeepSeek V4.1: $0.15/1M input, generates 1500 tokens response
For a task requiring 1000 input tokens:
- GPT-6: ($10 × 1K) + ($50 × 500) = $10 + $25 = $35
- DeepSeek: ($0.15 × 1K) + ($0.60 × 1500) = $0.15 + $0.90 = $1.05
DeepSeek is 33x cheaper for this task.
Cost Optimization Strategies
1. Aggressive Caching
Cache repeated context aggressively:
- Anthropic’s prompt caching: 90% discount on cached tokens
- OpenAI’s automatic caching
- Redis/Memcached for deterministic responses
2. Model Tiering
Not every request needs the most powerful model:
- Tier 1: GPT-6 Astra for complex reasoning, final outputs
- Tier 2: Claude Fable 5.1 for general tasks
- Tier 3: Gemini 3.8 Flash for simple queries
- Tier 4: DeepSeek V4.1 Flash for batch processing
- Tier 5: Ling 3.0 Flash VL / Mercury 2.5 for high-volume
Result: 60-70% cost reduction vs. using GPT-6 for everything.
3. Batching
Process multiple requests in parallel:
- Batch API for offline tasks (50% discount)
- Queue and batch similar requests
- Combine multiple queries into single prompts
4. Distillation
Train smaller models on outputs from larger models:
- GPT-6 generates training data
- Fine-tune DeepSeek on that data
- Deploy DeepSeek at fraction of cost
- 80-90% cost reduction with minimal quality loss
5. Token Budgeting
Hard limits per request:
- Maximum context length
- Maximum output length
- Maximum tokens per conversation
- Per-user daily limits
Total Cost of Ownership
Beyond per-token costs, factor in:
- Engineering time: Integration, maintenance, monitoring
- Latency costs: User patience, abandoned sessions
- Quality costs: Bad responses, retries, manual correction
- Compliance: Data residency, audit trails
- Opportunity cost: What else could that engineering time build?
The Open-Source Alternative
Self-hosting open-source models changes the economics:
Pros:
- No per-token costs
- Full control over data
- No rate limits
- Customizable
Cons:
- GPU infrastructure ($5K-$50K+ per server)
- Engineering time for deployment
- Model management overhead
- Usually lower quality than frontier models
Break-even analysis:
- Self-hosting Muse Spark 1.3: ~$15K upfront + $3K/month
- Equivalent API spend: breaks even at ~$3M tokens/month
- For most teams, API pricing is cheaper until massive scale
The Future
Pricing trends to watch:
- Hardware efficiency: New chips reduce inference costs
- Competition: Prices falling 2-3x per year
- Specialized models: Cheaper models for specific tasks
- On-device inference: Zero API costs for local models
- Commoditization: Inference becoming a commodity market
The Bottom Line
Don’t optimize for cheapest tokens. Optimize for:
- Cost per successful task completion
- Latency requirements
- Quality requirements
- Engineering time
A $0.01 request that fails and requires retry costs more than a $0.50 request that succeeds first time.