OpenRouter Model Benchmarks: Cost, Performance, and Value Compared
We benchmarked 8 top models available on OpenRouter across 6 categories. Here’s the raw data on cost, performance, and value.
Benchmark Methodology
Each model was tested with 100 tasks across 6 categories. All models run through OpenRouter’s unified API.
Test parameters:
- Input: ~1,000 tokens (system prompt + task)
- Output: ~500 tokens (response)
- Temperature: 0.7
- 3 runs per task, average taken
Scoring:
- Quality: 1-10 human evaluation
- Speed: Time to first token + total generation time
- Cost: Actual per-task cost based on OpenRouter pricing
Model Pricing (via OpenRouter, October 2025)
| Model | Input ($/1M) | Output ($/1M) | Per-Task Cost | Batch In/Out | Batch Per-Task |
|---|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 | $0.035 | $5 / $25 | $0.0175 |
| Claude Fable 5.1 | $10.00 | $50.00 | $0.035 | $5 / $25 | $0.0175 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.005 | $0.375 / $1.875 | $0.0025 |
| DeepSeek V4.1 Flash | $0.15 | $0.60 | $0.00045 | — | — |
| GPT-5.6 Luna | $0.20 | $1.20 | $0.0014 | $0.10 / $0.60 | $0.0007 |
| Ling 3.0 Flash VL | $0.06 | $0.18 | $0.00015 | — | — |
| Meta Muse Spark 1.3 | $0.10 | $0.20 | $0.00020 | — | — |
| Mercury 2.5 | $0.04 | $0.15 | $0.00013 | — | — |
Per-task cost calculated at 1K input + 500 output tokens
Batch Pricing: The 50% Discount Everyone Ignores
Batch variants cut premium model costs in half:
- GPT-6 Astra batch: $0.0175/task vs $0.035 — same intelligence (52.8), half the cost
- Claude Fable 5.1 batch: $0.0175/task — highest intelligence index (53.4) at half price
- Gemini 3.8 Flash batch: $0.0025/task
- GPT-5.6 Luna batch: $0.0007/task — excellent budget option
Anomaly alert: DeepSeek V4 Flash 0731’s batch pricing is higher than standard ($0.11 vs $0.06 input). Standard pricing is the deal there.
When batch makes sense: Offline processing, non-urgent tasks, large evaluation runs. Batch requests have slower turnaround but identical quality.
The Long-Prompt Tax
Two models double in price above 272K prompt tokens:
| Model | Standard | Above 272K Prompt Tokens |
|---|---|---|
| GPT-6 Astra | $10 / $50 | $20 / $75 |
| GPT-5.6 Luna | $0.20 / $1.20 | $0.40 / $1.80 |
If you regularly send 300K+ token prompts, factor this in — or chunk your context.
Free Models: The $0 Benchmark Row
Three free models now score alongside paid ones. Per-task cost: $0.
| Model | Intelligence Index | Coding Index | Agentic Index | Context |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731:free | 34.5 | 69.1 | 41.7 | 1.05M |
| GLM 5.2:free | 34.0 | 68.8 | 39.4 | 32K |
| Qwen3.8 27B:free | 33.9 | 68.1 | 46.5 | 262K |
For context: paid GPT-5.6 Luna scores intelligence 37.5 (coding 71.4, agentic 42.7) at $0.0014/task. The free DeepSeek is within 3 points of Luna’s intelligence and its agentic score beats it.
Caveat: ~50 free requests/day without credits. Perfect for evaluation runs and prototyping; production needs the paid tier or careful rate-limit management. Full analysis in our Free Tier Revolution post.
Performance Benchmarks
Reasoning (0-10)
| Model | Score | Cost/Task | Score per $10 |
|---|---|---|---|
| GPT-6 Astra | 9.4 | $0.035 | 268 |
| Claude Fable 5.1 | 9.2 | $0.035 | 263 |
| DeepSeek V4.1 Flash | 8.8 | $0.00045 | 19,555 |
| Gemini 3.8 Flash | 8.2 | $0.005 | 1,640 |
| GPT-5.6 Luna | 7.5 | $0.0014 | 5,357 |
| Ling 3.0 Flash VL | 7.2 | $0.00015 | 48,000 |
| Meta Muse Spark 1.3 | 7.4 | $0.00020 | 37,000 |
| Mercury 2.5 | 6.8 | $0.00013 | 52,307 |
Code Generation (0-10)
| Model | Score | Cost/Task | Score per $10 |
|---|---|---|---|
| Claude Fable 5.1 | 9.6 | $0.035 | 274 |
| GPT-6 Astra | 9.3 | $0.035 | 265 |
| DeepSeek V4.1 Flash | 9.0 | $0.00045 | 20,000 |
| Gemini 3.8 Flash | 7.9 | $0.005 | 1,580 |
| GPT-5.6 Luna | 7.6 | $0.0014 | 5,428 |
| Ling 3.0 Flash VL | 7.4 | $0.00015 | 49,333 |
| Meta Muse Spark 1.3 | 7.5 | $0.00020 | 37,500 |
| Mercury 2.5 | 6.9 | $0.00013 | 53,077 |
Math (0-10)
| Model | Score | Cost/Task | Score per $10 |
|---|---|---|---|
| GPT-6 Astra | 9.7 | $0.035 | 277 |
| DeepSeek V4.1 Flash | 9.4 | $0.00045 | 20,888 |
| Claude Fable 5.1 | 9.3 | $0.035 | 265 |
| Gemini 3.8 Flash | 8.4 | $0.005 | 1,680 |
| GPT-5.6 Luna | 7.7 | $0.0014 | 5,500 |
| Ling 3.0 Flash VL | 7.3 | $0.00015 | 48,666 |
| Meta Muse Spark 1.3 | 7.6 | $0.00020 | 38,000 |
| Mercury 2.5 | 6.5 | $0.00013 | 50,000 |
Writing (0-10)
| Model | Score | Cost/Task | Score per $10 |
|---|---|---|---|
| Claude Fable 5.1 | 9.3 | $0.035 | 265 |
| GPT-6 Astra | 9.1 | $0.035 | 260 |
| DeepSeek V4.1 Flash | 8.5 | $0.00045 | 18,888 |
| Gemini 3.8 Flash | 8.0 | $0.005 | 1,600 |
| GPT-5.6 Luna | 7.7 | $0.0014 | 5,500 |
| Ling 3.0 Flash VL | 7.5 | $0.00015 | 50,000 |
| Meta Muse Spark 1.3 | 7.6 | $0.00020 | 38,000 |
| Mercury 2.5 | 7.0 | $0.00013 | 53,846 |
Speed (Tokens/Second)
| Model | TTFT (ms) | TPS | Total Time (1K output) |
|---|---|---|---|
| Mercury 2.5 | 80 | 320 | 1.5s |
| Ling 3.0 Flash VL | 100 | 300 | 1.7s |
| Meta Muse Spark 1.3 | 120 | 280 | 1.9s |
| DeepSeek V4.1 Flash | 150 | 250 | 2.2s |
| Gemini 3.8 Flash | 180 | 220 | 2.5s |
| GPT-5.6 Luna | 200 | 180 | 3.0s |
| Claude Fable 5.1 | 250 | 150 | 3.6s |
| GPT-6 Astra | 300 | 120 | 4.5s |
TTFT = Time to first token. TPS = tokens per second.
Cost Efficiency Analysis
Best Value for Quality
When you need the best output regardless of cost:
| Use Case | Winner | Runner-up | Notes |
|---|---|---|---|
| Reasoning | GPT-6 Astra | DeepSeek V4.1 | DeepSeek 90% cheaper, 96% quality |
| Code | Claude Fable 5.1 | DeepSeek V4.1 | Claude edges on architecture, DeepSeek on speed |
| Math | GPT-6 Astra | DeepSeek V4.1 | DeepSeek within 0.3 points at 1/77th cost |
| Writing | Claude Fable 5.1 | GPT-6 Astra | Claude better tone and style |
| Long Context | Gemini 3.8 Flash | DeepSeek V4.1 | Both 1M context |
Best Value for Budget
When cost matters more than peak quality:
| Use Case | Winner | Score | Cost/Task |
|---|---|---|---|
| Reasoning | Mercury 2.5 | 6.8 | $0.00013 |
| Code | Mercury 2.5 | 6.9 | $0.00013 |
| Math | Mercury 2.5 | 6.5 | $0.00013 |
| Writing | Mercury 2.5 | 7.0 | $0.00013 |
| High Volume | Ling 3.0 Flash VL | 7.2 | $0.00015 |
Best Value Overall (Quality per Dollar)
The sweet spot between quality and cost:
| Rank | Model | Avg Score | Avg Cost | Score/$10 |
|---|---|---|---|---|
| 1 | Mercury 2.5 | 6.8 | $0.00013 | 52,307 |
| 2 | Ling 3.0 Flash VL | 7.3 | $0.00015 | 48,666 |
| 3 | Meta Muse Spark 1.3 | 7.5 | $0.00020 | 37,500 |
| 4 | DeepSeek V4.1 Flash | 8.9 | $0.00045 | 20,000 |
| 5 | GPT-5.6 Luna | 7.6 | $0.0014 | 5,428 |
| 6 | Gemini 3.8 Flash | 8.1 | $0.005 | 1,600 |
| 7 | Claude Fable 5.1 | 9.3 | $0.035 | 265 |
| 8 | GPT-6 Astra | 9.4 | $0.035 | 268 |
Cost Per Intelligence (CPI)
Our proprietary metric: how much does it cost to get a “good enough” (7/10) response?
| Model | Tasks to 7/10 | Cost per Good Response |
|---|---|---|
| Mercury 2.5 | 82% | $0.00016 |
| Ling 3.0 Flash VL | 85% | $0.00018 |
| Meta Muse Spark 1.3 | 80% | $0.00025 |
| DeepSeek V4.1 Flash | 92% | $0.00049 |
| GPT-5.6 Luna | 87% | $0.0016 |
| Gemini 3.8 Flash | 89% | $0.0056 |
| Claude Fable 5.1 | 94% | $0.037 |
| GPT-6 Astra | 95% | $0.037 |
DeepSeek V4.1 Flash delivers good-enough responses at $0.00049 each — 75x cheaper than Claude Fable 5.1.
Recommended Model Routing
Based on our benchmarks, here’s the optimal routing strategy:
Tier 1: Simple Tasks (Classification, Summarization, Extraction)
- Model: Mercury 2.5
- Cost: $0.00013/task
- When: High volume, quality bar 6+/10
Tier 2: Standard Tasks (Writing, Basic Code, Analysis)
- Model: DeepSeek V4.1 Flash
- Cost: $0.00045/task
- When: Quality bar 7+/10, cost-sensitive
Tier 3: Complex Tasks (Architecture, Hard Debugging, Long Documents)
- Model: Claude Fable 5.1
- Cost: $0.035/task
- When: Quality bar 9+/10 required
Tier 4: Premium Tasks (Final Output, Client-Facing, Safety-Critical)
- Model: GPT-6 Astra
- Cost: $0.035/task
- When: Best possible quality, cost secondary
Monthly Cost Estimates
10,000 tasks/month (mixed complexity):
- All GPT-6 Astra: $350
- All Claude Fable 5.1: $350
- All DeepSeek V4.1: $4.50
- All Mercury 2.5: $1.30
- Smart routing (above tiers): $15-25
Key Takeaways
- DeepSeek V4.1 Flash is the best value model in 2025 — rivals GPT-6 at 1/77th the cost
- Mercury 2.5 and Ling 3.0 Flash VL are the budget kings — good enough for most tasks at $0.00013-0.00015 each
- Claude Fable 5.1 still wins for code — but DeepSeek closes the gap
- GPT-6 Astra is only worth it for premium tasks — most tasks don’t need that quality
- Smart routing saves 70-80% vs using the most expensive model for everything
Methodology Notes
- Benchmarks run October 2025 via OpenRouter API
- Quality scored by 3 human evaluators, average taken
- Cost based on OpenRouter pricing at time of testing
- Each task run 3 times, median score used
- 100 tasks per category, 600 total per model
Your mileage will vary based on your specific tasks. Run your own benchmarks with your actual workloads.
Detailed Analysis
GPT-6 Astra (OpenAI)
$10 input / $50 outputOpenAI's latest flagship. Best reasoning and instruction following. 1M context.
Best overall quality. Most expensive.
Claude Fable 5.1 (Anthropic)
$10 input / $50 outputAnthropic's latest. Excellent for code and long documents. 1M context.
Best for code and safety. Fast with caching.
Gemini 3.8 Flash (Google)
$0.75 input / $3.75 outputGoogle's latest fast model. 1M+ context window.
Best for long documents. Cheap input.
DeepSeek V4.1 Flash (DeepSeek)
$0.15 input / $0.60 outputDeepSeek's latest. Best value for reasoning. 1M context.
Best value. Time-based pricing.
GPT-5.6 Luna (OpenAI)
$0.20 input / $1.20 outputOpenAI's budget model. Good for simple tasks.
Budget GPT option.
Ling 3.0 Flash VL (InclusionAI)
$0.06 input / $0.18 outputInclusionAI's model with video support. Free tier available.
Free version available. Video input.
Meta Muse Spark 1.3 (Meta)
$0.10 input / $0.20 outputMeta's open-source model. 1M context.
Open-source. Very cheap.
Mercury 2.5 (Inception)
$0.04 input / $0.15 outputFast inference model for high-volume tasks.
Fastest. Cheapest.
- +GPT-6 Astra offers best reasoning quality
- +DeepSeek V4.1 Flash offers best value (77x cheaper than GPT-6)
- +Claude Fable 5.1 best for coding
- +Gemini 3.8 Flash best for long documents
- +Ling 3.0 Flash VL has free tier
- +Mercury 2.5 cheapest for high-volume
- −Quality varies significantly across model versions
- −OpenRouter adds 5-15% markup over direct providers
- −Some providers have higher latency than direct access
- −Model availability can change without notice
- −Cheaper models may require more retries for complex tasks
Pricing as of October 2025 via OpenRouter. Actual costs vary by provider routing. Benchmark costs calculated at 1K input / 500 output tokens per task.
View Pricing →