INTEL DOSSIER|CLASSIFIED
DECRYPTED
File #10 • October 11, 2025

Top 10 LLM Models by Weekly Token Consumption (Full Data)

openrouterllmbenchmarksusage
Authorcoderunner
Categoryopenrouter
StatusPUBLISHED
ClearancePUBLIC
//Complete breakdown of the top 10 LLM models by weekly token consumption on OpenRouter, with full data — real token counts, pricing, context length, and capabilities for each model.

Real data from OpenRouter — October 2025.

The top 10 LLM models by weekly token consumption, with complete data from the OpenRouter API.


Usage ≠ Intelligence

The biggest insight from this data: the most-used models are NOT the most intelligent. Intelligence Index (Artificial Analysis) vs weekly usage:

Model Weekly Usage Rank Intelligence Index The Gap
GPT-5.6 Luna #1 37.5 Most used, mid-tier intelligence
Claude Fable 5.1 #10* 53.4 Highest intelligence, least used
GPT-6 Astra #10 52.8 Frontier intelligence, premium price
GLM 5.3 Flash #4 41.9 Agentic index 51.2 — nearly matches GPT-6 Astra (51.5)
DeepSeek V4.1 Flash #2 39.5 Best value-to-intelligence ratio

*Claude Fable 5.1 sits outside the weekly top 10 by usage but tops the intelligence chart.

The market optimizes for cost, not capability. Users pick models that are “good enough” at 1/50th the price.


#1: GPT-5.6 Luna (OpenAI)

Weekly Tokens: 16.1T | Rank: #1

Spec Value
Context 1,050,000 tokens
Input $0.20/M tokens
Output $1.20/M tokens
Modality Text + Image + File
Created July 9, 2026

Description: GPT-5.6 Luna is a fast, cost-efficient model in OpenAI’s GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

Why It’s #1: 50x cheaper than GPT-6 Astra ($0.20 vs $10 input). For most tasks, good enough quality at a fraction of the cost.


#2: DeepSeek V4.1 Flash (DeepSeek)

Weekly Tokens: 12.5T | Rank: #2

Spec Value
Context 1,048,576 tokens
Input $0.15/M tokens
Output $0.60/M tokens
Modality Text + Image
Created September 10, 2026

Description: DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company’s Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone. Suited for coding, terminal, and computer-use agents, along with long-horizon tasks. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation.

Why It’s #2: 67x cheaper than GPT-6 Astra. Exceeds V4 Pro on performance, speed, and task completion time.


#3: Tencent Hy4 Preview (Tencent)

Weekly Tokens: 12.4T | Rank: #3

Spec Value
Context 1,048,576 tokens
Input $0.834/M tokens
Output $2.501/M tokens
Modality Text
Created August 28, 2026

Description: Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

Why It’s #3: 12x cheaper than GPT-6 Astra. Strong on coding agents and complex tool-use workflows.


#4: Z.ai GLM 5.3 Flash (Z.ai)

Weekly Tokens: 11.9T | Rank: #4

Spec Value
Context 1,310,720 tokens
Input $0.075/M tokens
Output $0.25/M tokens
Modality Text
Created August 26, 2026

Description: GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead. Currently 50% off on OpenRouter.

Why It’s #4: 133x cheaper than GPT-6 Astra. 1.31M context window — largest in top 10. 50% discount makes it even cheaper.


#5: DeepSeek V4 Flash 0731 (DeepSeek)

Weekly Tokens: 10.9T | Rank: #5

Spec Value
Context 1,310,720 tokens
Input $0.056/M tokens
Output $0.177/M tokens
Modality Text
Created July 31, 2026

Description: DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

Why It’s #5: 178x cheaper than GPT-6 Astra. 1.31M context window. Ranked #1 in Academia, Finance, Health, and Marketing categories.


#6: Xiaomi MiMo-V2.5 (Xiaomi)

Weekly Tokens: 7.77T | Rank: #6

Spec Value
Context 1,048,576 tokens
Input $0.119/M tokens
Output $0.238/M tokens
Modality Text + Image + Audio
Created April 22, 2026

Description: MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass. Currently 15% off on OpenRouter.

Why It’s #6: 84x cheaper than GPT-6 Astra. Omnimodal (text, image, audio, video). 15% discount.


#7: Tencent Hy3 (Tencent)

Weekly Tokens: 5.05T | Rank: #7

Spec Value
Context 256,000 tokens
Input 25% off
Output 25% off
Modality Text
Created August 28, 2026

Description: Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems.

Why It’s #7: 25% discount on OpenRouter. Strong on reasoning and agentic workflows.


#8: DeepSeek V4 Flash Vision (DeepSeek)

Weekly Tokens: 4.8T | Rank: #8

Spec Value
Context 1,048,576 tokens
Input $0.05/M tokens
Output $0.15/M tokens
Modality Text + Image
Created July 31, 2026

Description: Experimental version of DeepSeek V4 Flash with vision capabilities. Built for coding, reasoning, and agent workflows with image understanding.

Why It’s #8: 200x cheaper than GPT-6 Astra. Vision + 1M context. DeepSeek’s experimental vision model.


#9: Gemini 3.8 Flash (Google)

Weekly Tokens: 4.5T | Rank: #9

Spec Value
Context 1,048,576 tokens
Input $0.75/M tokens
Output $3.75/M tokens
Modality Text + Image + Audio
Created August 30, 2026

Description: Google’s latest fast model. 1M context standard, native multimodal (text, image, audio), built-in web search.

Why It’s #9: 13x cheaper than GPT-6 Astra. Native multimodal. Google’s fastest model.


#10: GPT-6 Astra (OpenAI)

Weekly Tokens: 4.2T | Rank: #10

Spec Value
Context 1,050,000 tokens
Input $10.00/M tokens
Output $50.00/M tokens
Modality Text + Image + File
Created September 4, 2026

Description: GPT-6 Astra is OpenAI’s flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and complex reasoning. Mandatory reasoning with 5 effort levels (max, xhigh, high, medium, low).

Why It’s #10: Most expensive model in top 10. Best reasoning quality. For tasks that need the absolute best, regardless of cost.


Summary Table

Rank Model Weekly Tokens Input Output Context
1 GPT-5.6 Luna 16.1T $0.20 $1.20 1.05M
2 DeepSeek V4.1 Flash 12.5T $0.15 $0.60 1.05M
3 Tencent Hy4 12.4T $0.834 $2.501 1.05M
4 Z.ai GLM 5.3 Flash 11.9T $0.075 $0.25 1.31M
5 DeepSeek V4 Flash 0731 10.9T $0.056 $0.177 1.31M
6 Xiaomi MiMo-V2.5 7.77T $0.119 $0.238 1.05M
7 Tencent Hy3 5.05T 25% off 25% off 256K
8 DeepSeek V4 Flash Vision 4.8T $0.05 $0.15 1.05M
9 Gemini 3.8 Flash 4.5T $0.75 $3.75 1.05M
10 GPT-6 Astra 4.2T $10.00 $50.00 1.05M

Key Takeaways

  1. Price = Usage: Luna ($0.20) gets 4x more tokens than Astra ($10)
  2. DeepSeek dominates: 3 models in top 5, 4 models in top 8
  3. Chinese models rising: Tencent, Z.ai, Xiaomi all in top 7
  4. 1M context standard: 9 of 10 models have 1M+ context
  5. GPT-6 is niche: Only #10 despite being “the best” — users choose value
  6. Usage ≠ Intelligence: Claude Fable 5.1 has the highest intelligence index (53.4) but sits outside the top 10 by usage
  7. Batch pricing halves costs: GPT-6 Astra batch is $5/$25 vs $10/$50 standard — check batch before paying full price
  8. Watch the long-prompt tax: GPT-6 Astra and Luna cost 2x above 272K prompt tokens
  9. Free tiers are real: DeepSeek V4 Flash 0731:free offers intelligence 34.5 at $0

Data: OpenRouter API + Top Weekly rankings, October 2025. Intelligence indices via Artificial Analysis.

Detailed Analysis

Strengths3 PROS
  • +Complete data: token counts, pricing, context, capabilities
  • +Based on real OpenRouter API data
  • +Shows what models are actually being used in production
Weaknesses2 CONS
  • −Data is from October 2025 and changes weekly
  • −Only includes models on OpenRouter (not direct API usage)
PricingVarious

Complete data for the top 10 LLM models by weekly token consumption on OpenRouter.

View Pricing →
END OF FILE|DISTRIBUTION: UNLIMITED
← Back to Archive