INTEL DOSSIER|CLASSIFIED
DECRYPTED
File #09 • October 10, 2025

LLM Pricing: The True Cost of Intelligence

llmpricingeconomics
Authorcoderunner
Categoryllm
StatusPUBLISHED
ClearancePUBLIC

Beyond the sticker price — understanding token economics, hidden costs, and how to actually budget for AI at scale. Updated with October 2025 pricing.

Pricing Models

Per-Token Pricing

The standard model: cost per 1M tokens processed.

Important distinction:

  • Input tokens: Your prompt + context (cheaper)
  • Output tokens: Model’s response (more expensive)
  • Cached input: Repeated context (cheapest, ~90% discount)

Current Major Provider Pricing (per 1M tokens, October 2025)

GPT-6 Astra (OpenAI)

  • Input: $10.00
  • Cached Input: $1.00
  • Output: $50.00

GPT-5.6 Sol (OpenAI)

  • Input: $2.00
  • Cached Input: $0.20
  • Output: $10.00

GPT-5.6 Luna (OpenAI)

  • Input: $0.20
  • Cached Input: $0.02
  • Output: $1.20

Claude Fable 5.1 (Anthropic)

  • Input: $10.00
  • Cached Input: $0.10
  • Output: $50.00

Claude 3 Haiku (Anthropic)

  • Input: $0.25
  • Cached Input: $0.03
  • Output: $1.25

Gemini 3.8 Flash (Google)

  • Input: $0.75
  • Cached Input: $0.075
  • Output: $3.75

DeepSeek V4.1 Flash (DeepSeek)

  • Input: $0.15-$0.30 (time-based)
  • Cached Input: $0.006
  • Output: $0.60-$1.20

Ling 3.0 Flash VL (InclusionAI)

  • Input: $0.06
  • Cached Input: $0.012
  • Output: $0.18

Meta Muse Spark 1.3 (Meta)

  • Input: $0.10
  • Cached Input: $0.025
  • Output: $0.20

Mercury 2.5 (Inception)

  • Input: $0.04
  • Output: $0.15

Hidden Costs

Context Window Tax

Every token in your context window costs money. A 1M context window with 800K of system prompt = 800K tokens × input cost per request.

At scale:

  • 1M requests/day × 800K input tokens × $10/1M = $8,000/day just for context
  • That’s $2.9M/year in context alone

Token Counting Gotchas

  • Whitespace counts as tokens
  • Special characters count as tokens
  • Images count as tokens (base64 encoded, very expensive)
  • Tool calls count as tokens (both directions)
  • System prompts are repeated every request

The “Cheap Model” Trap

A cheaper model that requires 3x more tokens to achieve the same result isn’t actually cheaper.

Example:

  • GPT-6 Astra: $10/1M input, generates 500 tokens response
  • DeepSeek V4.1: $0.15/1M input, generates 1500 tokens response

For a task requiring 1000 input tokens:

  • GPT-6: ($10 × 1K) + ($50 × 500) = $10 + $25 = $35
  • DeepSeek: ($0.15 × 1K) + ($0.60 × 1500) = $0.15 + $0.90 = $1.05

DeepSeek is 33x cheaper for this task.

Cost Optimization Strategies

1. Aggressive Caching

Cache repeated context aggressively:

  • Anthropic’s prompt caching: 90% discount on cached tokens
  • OpenAI’s automatic caching
  • Redis/Memcached for deterministic responses

2. Model Tiering

Not every request needs the most powerful model:

  • Tier 1: GPT-6 Astra for complex reasoning, final outputs
  • Tier 2: Claude Fable 5.1 for general tasks
  • Tier 3: Gemini 3.8 Flash for simple queries
  • Tier 4: DeepSeek V4.1 Flash for batch processing
  • Tier 5: Ling 3.0 Flash VL / Mercury 2.5 for high-volume

Result: 60-70% cost reduction vs. using GPT-6 for everything.

3. Batching

Process multiple requests in parallel:

  • Batch API for offline tasks (50% discount)
  • Queue and batch similar requests
  • Combine multiple queries into single prompts

4. Distillation

Train smaller models on outputs from larger models:

  • GPT-6 generates training data
  • Fine-tune DeepSeek on that data
  • Deploy DeepSeek at fraction of cost
  • 80-90% cost reduction with minimal quality loss

5. Token Budgeting

Hard limits per request:

  • Maximum context length
  • Maximum output length
  • Maximum tokens per conversation
  • Per-user daily limits

Total Cost of Ownership

Beyond per-token costs, factor in:

  • Engineering time: Integration, maintenance, monitoring
  • Latency costs: User patience, abandoned sessions
  • Quality costs: Bad responses, retries, manual correction
  • Compliance: Data residency, audit trails
  • Opportunity cost: What else could that engineering time build?

The Open-Source Alternative

Self-hosting open-source models changes the economics:

Pros:

  • No per-token costs
  • Full control over data
  • No rate limits
  • Customizable

Cons:

  • GPU infrastructure ($5K-$50K+ per server)
  • Engineering time for deployment
  • Model management overhead
  • Usually lower quality than frontier models

Break-even analysis:

  • Self-hosting Muse Spark 1.3: ~$15K upfront + $3K/month
  • Equivalent API spend: breaks even at ~$3M tokens/month
  • For most teams, API pricing is cheaper until massive scale

The Future

Pricing trends to watch:

  1. Hardware efficiency: New chips reduce inference costs
  2. Competition: Prices falling 2-3x per year
  3. Specialized models: Cheaper models for specific tasks
  4. On-device inference: Zero API costs for local models
  5. Commoditization: Inference becoming a commodity market

The Bottom Line

Don’t optimize for cheapest tokens. Optimize for:

  • Cost per successful task completion
  • Latency requirements
  • Quality requirements
  • Engineering time

A $0.01 request that fails and requires retry costs more than a $0.50 request that succeeds first time.

END OF FILE|DISTRIBUTION: UNLIMITED
← Back to Archive