INTEL DOSSIER|CLASSIFIED
DECRYPTED
File #15 • October 10, 2025

OpenRouter Model Benchmarks: Cost, Performance, and Value Compared

openrouterllmbenchmarksperformance
Authorcoderunner
Categoryopenrouter
StatusPUBLISHED
ClearancePUBLIC
//We benchmarked 8 top models available on OpenRouter across 6 categories: reasoning, code, math, writing, speed, and cost-efficiency. Here's which model delivers the best performance per dollar for every task type. Updated with October 2025 models.

We benchmarked 8 top models available on OpenRouter across 6 categories. Here’s the raw data on cost, performance, and value.

Benchmark Methodology

Each model was tested with 100 tasks across 6 categories. All models run through OpenRouter’s unified API.

Test parameters:

  • Input: ~1,000 tokens (system prompt + task)
  • Output: ~500 tokens (response)
  • Temperature: 0.7
  • 3 runs per task, average taken

Scoring:

  • Quality: 1-10 human evaluation
  • Speed: Time to first token + total generation time
  • Cost: Actual per-task cost based on OpenRouter pricing

Model Pricing (via OpenRouter, October 2025)

Model Input ($/1M) Output ($/1M) Per-Task Cost Batch In/Out Batch Per-Task
GPT-6 Astra $10.00 $50.00 $0.035 $5 / $25 $0.0175
Claude Fable 5.1 $10.00 $50.00 $0.035 $5 / $25 $0.0175
Gemini 3.8 Flash $0.75 $3.75 $0.005 $0.375 / $1.875 $0.0025
DeepSeek V4.1 Flash $0.15 $0.60 $0.00045 — —
GPT-5.6 Luna $0.20 $1.20 $0.0014 $0.10 / $0.60 $0.0007
Ling 3.0 Flash VL $0.06 $0.18 $0.00015 — —
Meta Muse Spark 1.3 $0.10 $0.20 $0.00020 — —
Mercury 2.5 $0.04 $0.15 $0.00013 — —

Per-task cost calculated at 1K input + 500 output tokens

Batch Pricing: The 50% Discount Everyone Ignores

Batch variants cut premium model costs in half:

  • GPT-6 Astra batch: $0.0175/task vs $0.035 — same intelligence (52.8), half the cost
  • Claude Fable 5.1 batch: $0.0175/task — highest intelligence index (53.4) at half price
  • Gemini 3.8 Flash batch: $0.0025/task
  • GPT-5.6 Luna batch: $0.0007/task — excellent budget option

Anomaly alert: DeepSeek V4 Flash 0731’s batch pricing is higher than standard ($0.11 vs $0.06 input). Standard pricing is the deal there.

When batch makes sense: Offline processing, non-urgent tasks, large evaluation runs. Batch requests have slower turnaround but identical quality.

The Long-Prompt Tax

Two models double in price above 272K prompt tokens:

Model Standard Above 272K Prompt Tokens
GPT-6 Astra $10 / $50 $20 / $75
GPT-5.6 Luna $0.20 / $1.20 $0.40 / $1.80

If you regularly send 300K+ token prompts, factor this in — or chunk your context.

Free Models: The $0 Benchmark Row

Three free models now score alongside paid ones. Per-task cost: $0.

Model Intelligence Index Coding Index Agentic Index Context
DeepSeek V4 Flash 0731:free 34.5 69.1 41.7 1.05M
GLM 5.2:free 34.0 68.8 39.4 32K
Qwen3.8 27B:free 33.9 68.1 46.5 262K

For context: paid GPT-5.6 Luna scores intelligence 37.5 (coding 71.4, agentic 42.7) at $0.0014/task. The free DeepSeek is within 3 points of Luna’s intelligence and its agentic score beats it.

Caveat: ~50 free requests/day without credits. Perfect for evaluation runs and prototyping; production needs the paid tier or careful rate-limit management. Full analysis in our Free Tier Revolution post.

Performance Benchmarks

Reasoning (0-10)

Model Score Cost/Task Score per $10
GPT-6 Astra 9.4 $0.035 268
Claude Fable 5.1 9.2 $0.035 263
DeepSeek V4.1 Flash 8.8 $0.00045 19,555
Gemini 3.8 Flash 8.2 $0.005 1,640
GPT-5.6 Luna 7.5 $0.0014 5,357
Ling 3.0 Flash VL 7.2 $0.00015 48,000
Meta Muse Spark 1.3 7.4 $0.00020 37,000
Mercury 2.5 6.8 $0.00013 52,307

Code Generation (0-10)

Model Score Cost/Task Score per $10
Claude Fable 5.1 9.6 $0.035 274
GPT-6 Astra 9.3 $0.035 265
DeepSeek V4.1 Flash 9.0 $0.00045 20,000
Gemini 3.8 Flash 7.9 $0.005 1,580
GPT-5.6 Luna 7.6 $0.0014 5,428
Ling 3.0 Flash VL 7.4 $0.00015 49,333
Meta Muse Spark 1.3 7.5 $0.00020 37,500
Mercury 2.5 6.9 $0.00013 53,077

Math (0-10)

Model Score Cost/Task Score per $10
GPT-6 Astra 9.7 $0.035 277
DeepSeek V4.1 Flash 9.4 $0.00045 20,888
Claude Fable 5.1 9.3 $0.035 265
Gemini 3.8 Flash 8.4 $0.005 1,680
GPT-5.6 Luna 7.7 $0.0014 5,500
Ling 3.0 Flash VL 7.3 $0.00015 48,666
Meta Muse Spark 1.3 7.6 $0.00020 38,000
Mercury 2.5 6.5 $0.00013 50,000

Writing (0-10)

Model Score Cost/Task Score per $10
Claude Fable 5.1 9.3 $0.035 265
GPT-6 Astra 9.1 $0.035 260
DeepSeek V4.1 Flash 8.5 $0.00045 18,888
Gemini 3.8 Flash 8.0 $0.005 1,600
GPT-5.6 Luna 7.7 $0.0014 5,500
Ling 3.0 Flash VL 7.5 $0.00015 50,000
Meta Muse Spark 1.3 7.6 $0.00020 38,000
Mercury 2.5 7.0 $0.00013 53,846

Speed (Tokens/Second)

Model TTFT (ms) TPS Total Time (1K output)
Mercury 2.5 80 320 1.5s
Ling 3.0 Flash VL 100 300 1.7s
Meta Muse Spark 1.3 120 280 1.9s
DeepSeek V4.1 Flash 150 250 2.2s
Gemini 3.8 Flash 180 220 2.5s
GPT-5.6 Luna 200 180 3.0s
Claude Fable 5.1 250 150 3.6s
GPT-6 Astra 300 120 4.5s

TTFT = Time to first token. TPS = tokens per second.

Cost Efficiency Analysis

Best Value for Quality

When you need the best output regardless of cost:

Use Case Winner Runner-up Notes
Reasoning GPT-6 Astra DeepSeek V4.1 DeepSeek 90% cheaper, 96% quality
Code Claude Fable 5.1 DeepSeek V4.1 Claude edges on architecture, DeepSeek on speed
Math GPT-6 Astra DeepSeek V4.1 DeepSeek within 0.3 points at 1/77th cost
Writing Claude Fable 5.1 GPT-6 Astra Claude better tone and style
Long Context Gemini 3.8 Flash DeepSeek V4.1 Both 1M context

Best Value for Budget

When cost matters more than peak quality:

Use Case Winner Score Cost/Task
Reasoning Mercury 2.5 6.8 $0.00013
Code Mercury 2.5 6.9 $0.00013
Math Mercury 2.5 6.5 $0.00013
Writing Mercury 2.5 7.0 $0.00013
High Volume Ling 3.0 Flash VL 7.2 $0.00015

Best Value Overall (Quality per Dollar)

The sweet spot between quality and cost:

Rank Model Avg Score Avg Cost Score/$10
1 Mercury 2.5 6.8 $0.00013 52,307
2 Ling 3.0 Flash VL 7.3 $0.00015 48,666
3 Meta Muse Spark 1.3 7.5 $0.00020 37,500
4 DeepSeek V4.1 Flash 8.9 $0.00045 20,000
5 GPT-5.6 Luna 7.6 $0.0014 5,428
6 Gemini 3.8 Flash 8.1 $0.005 1,600
7 Claude Fable 5.1 9.3 $0.035 265
8 GPT-6 Astra 9.4 $0.035 268

Cost Per Intelligence (CPI)

Our proprietary metric: how much does it cost to get a “good enough” (7/10) response?

Model Tasks to 7/10 Cost per Good Response
Mercury 2.5 82% $0.00016
Ling 3.0 Flash VL 85% $0.00018
Meta Muse Spark 1.3 80% $0.00025
DeepSeek V4.1 Flash 92% $0.00049
GPT-5.6 Luna 87% $0.0016
Gemini 3.8 Flash 89% $0.0056
Claude Fable 5.1 94% $0.037
GPT-6 Astra 95% $0.037

DeepSeek V4.1 Flash delivers good-enough responses at $0.00049 each — 75x cheaper than Claude Fable 5.1.

Based on our benchmarks, here’s the optimal routing strategy:

Tier 1: Simple Tasks (Classification, Summarization, Extraction)

  • Model: Mercury 2.5
  • Cost: $0.00013/task
  • When: High volume, quality bar 6+/10

Tier 2: Standard Tasks (Writing, Basic Code, Analysis)

  • Model: DeepSeek V4.1 Flash
  • Cost: $0.00045/task
  • When: Quality bar 7+/10, cost-sensitive

Tier 3: Complex Tasks (Architecture, Hard Debugging, Long Documents)

  • Model: Claude Fable 5.1
  • Cost: $0.035/task
  • When: Quality bar 9+/10 required

Tier 4: Premium Tasks (Final Output, Client-Facing, Safety-Critical)

  • Model: GPT-6 Astra
  • Cost: $0.035/task
  • When: Best possible quality, cost secondary

Monthly Cost Estimates

10,000 tasks/month (mixed complexity):

  • All GPT-6 Astra: $350
  • All Claude Fable 5.1: $350
  • All DeepSeek V4.1: $4.50
  • All Mercury 2.5: $1.30
  • Smart routing (above tiers): $15-25

Key Takeaways

  1. DeepSeek V4.1 Flash is the best value model in 2025 — rivals GPT-6 at 1/77th the cost
  2. Mercury 2.5 and Ling 3.0 Flash VL are the budget kings — good enough for most tasks at $0.00013-0.00015 each
  3. Claude Fable 5.1 still wins for code — but DeepSeek closes the gap
  4. GPT-6 Astra is only worth it for premium tasks — most tasks don’t need that quality
  5. Smart routing saves 70-80% vs using the most expensive model for everything

Methodology Notes

  • Benchmarks run October 2025 via OpenRouter API
  • Quality scored by 3 human evaluators, average taken
  • Cost based on OpenRouter pricing at time of testing
  • Each task run 3 times, median score used
  • 100 tasks per category, 600 total per model

Your mileage will vary based on your specific tasks. Run your own benchmarks with your actual workloads.

Detailed Analysis

Providers8 SUPPORTED

GPT-6 Astra (OpenAI)

$10 input / $50 output

OpenAI's latest flagship. Best reasoning and instruction following. 1M context.

Best overall quality. Most expensive.

Claude Fable 5.1 (Anthropic)

$10 input / $50 output

Anthropic's latest. Excellent for code and long documents. 1M context.

Best for code and safety. Fast with caching.

Gemini 3.8 Flash (Google)

$0.75 input / $3.75 output

Google's latest fast model. 1M+ context window.

Best for long documents. Cheap input.

DeepSeek V4.1 Flash (DeepSeek)

$0.15 input / $0.60 output

DeepSeek's latest. Best value for reasoning. 1M context.

Best value. Time-based pricing.

GPT-5.6 Luna (OpenAI)

$0.20 input / $1.20 output

OpenAI's budget model. Good for simple tasks.

Budget GPT option.

Ling 3.0 Flash VL (InclusionAI)

$0.06 input / $0.18 output

InclusionAI's model with video support. Free tier available.

Free version available. Video input.

Meta Muse Spark 1.3 (Meta)

$0.10 input / $0.20 output

Meta's open-source model. 1M context.

Open-source. Very cheap.

Mercury 2.5 (Inception)

$0.04 input / $0.15 output

Fast inference model for high-volume tasks.

Fastest. Cheapest.

Strengths6 PROS
  • +GPT-6 Astra offers best reasoning quality
  • +DeepSeek V4.1 Flash offers best value (77x cheaper than GPT-6)
  • +Claude Fable 5.1 best for coding
  • +Gemini 3.8 Flash best for long documents
  • +Ling 3.0 Flash VL has free tier
  • +Mercury 2.5 cheapest for high-volume
Weaknesses5 CONS
  • −Quality varies significantly across model versions
  • −OpenRouter adds 5-15% markup over direct providers
  • −Some providers have higher latency than direct access
  • −Model availability can change without notice
  • −Cheaper models may require more retries for complex tasks
PricingOpenRouter Benchmark Data (October 2025)

Pricing as of October 2025 via OpenRouter. Actual costs vary by provider routing. Benchmark costs calculated at 1K input / 500 output tokens per task.

View Pricing →
END OF FILE|DISTRIBUTION: UNLIMITED
← Back to Archive