Top 10 LLM Models by Weekly Token Consumption (Full Data)
Real data from OpenRouter — October 2025.
The top 10 LLM models by weekly token consumption, with complete data from the OpenRouter API.
Usage ≠ Intelligence
The biggest insight from this data: the most-used models are NOT the most intelligent. Intelligence Index (Artificial Analysis) vs weekly usage:
| Model | Weekly Usage Rank | Intelligence Index | The Gap |
|---|---|---|---|
| GPT-5.6 Luna | #1 | 37.5 | Most used, mid-tier intelligence |
| Claude Fable 5.1 | #10* | 53.4 | Highest intelligence, least used |
| GPT-6 Astra | #10 | 52.8 | Frontier intelligence, premium price |
| GLM 5.3 Flash | #4 | 41.9 | Agentic index 51.2 — nearly matches GPT-6 Astra (51.5) |
| DeepSeek V4.1 Flash | #2 | 39.5 | Best value-to-intelligence ratio |
*Claude Fable 5.1 sits outside the weekly top 10 by usage but tops the intelligence chart.
The market optimizes for cost, not capability. Users pick models that are “good enough” at 1/50th the price.
#1: GPT-5.6 Luna (OpenAI)
Weekly Tokens: 16.1T | Rank: #1
| Spec | Value |
|---|---|
| Context | 1,050,000 tokens |
| Input | $0.20/M tokens |
| Output | $1.20/M tokens |
| Modality | Text + Image + File |
| Created | July 9, 2026 |
Description: GPT-5.6 Luna is a fast, cost-efficient model in OpenAI’s GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.
Why It’s #1: 50x cheaper than GPT-6 Astra ($0.20 vs $10 input). For most tasks, good enough quality at a fraction of the cost.
#2: DeepSeek V4.1 Flash (DeepSeek)
Weekly Tokens: 12.5T | Rank: #2
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.15/M tokens |
| Output | $0.60/M tokens |
| Modality | Text + Image |
| Created | September 10, 2026 |
Description: DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company’s Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone. Suited for coding, terminal, and computer-use agents, along with long-horizon tasks. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation.
Why It’s #2: 67x cheaper than GPT-6 Astra. Exceeds V4 Pro on performance, speed, and task completion time.
#3: Tencent Hy4 Preview (Tencent)
Weekly Tokens: 12.4T | Rank: #3
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.834/M tokens |
| Output | $2.501/M tokens |
| Modality | Text |
| Created | August 28, 2026 |
Description: Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.
Why It’s #3: 12x cheaper than GPT-6 Astra. Strong on coding agents and complex tool-use workflows.
#4: Z.ai GLM 5.3 Flash (Z.ai)
Weekly Tokens: 11.9T | Rank: #4
| Spec | Value |
|---|---|
| Context | 1,310,720 tokens |
| Input | $0.075/M tokens |
| Output | $0.25/M tokens |
| Modality | Text |
| Created | August 26, 2026 |
Description: GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead. Currently 50% off on OpenRouter.
Why It’s #4: 133x cheaper than GPT-6 Astra. 1.31M context window — largest in top 10. 50% discount makes it even cheaper.
#5: DeepSeek V4 Flash 0731 (DeepSeek)
Weekly Tokens: 10.9T | Rank: #5
| Spec | Value |
|---|---|
| Context | 1,310,720 tokens |
| Input | $0.056/M tokens |
| Output | $0.177/M tokens |
| Modality | Text |
| Created | July 31, 2026 |
Description: DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.
Why It’s #5: 178x cheaper than GPT-6 Astra. 1.31M context window. Ranked #1 in Academia, Finance, Health, and Marketing categories.
#6: Xiaomi MiMo-V2.5 (Xiaomi)
Weekly Tokens: 7.77T | Rank: #6
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.119/M tokens |
| Output | $0.238/M tokens |
| Modality | Text + Image + Audio |
| Created | April 22, 2026 |
Description: MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass. Currently 15% off on OpenRouter.
Why It’s #6: 84x cheaper than GPT-6 Astra. Omnimodal (text, image, audio, video). 15% discount.
#7: Tencent Hy3 (Tencent)
Weekly Tokens: 5.05T | Rank: #7
| Spec | Value |
|---|---|
| Context | 256,000 tokens |
| Input | 25% off |
| Output | 25% off |
| Modality | Text |
| Created | August 28, 2026 |
Description: Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems.
Why It’s #7: 25% discount on OpenRouter. Strong on reasoning and agentic workflows.
#8: DeepSeek V4 Flash Vision (DeepSeek)
Weekly Tokens: 4.8T | Rank: #8
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.05/M tokens |
| Output | $0.15/M tokens |
| Modality | Text + Image |
| Created | July 31, 2026 |
Description: Experimental version of DeepSeek V4 Flash with vision capabilities. Built for coding, reasoning, and agent workflows with image understanding.
Why It’s #8: 200x cheaper than GPT-6 Astra. Vision + 1M context. DeepSeek’s experimental vision model.
#9: Gemini 3.8 Flash (Google)
Weekly Tokens: 4.5T | Rank: #9
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.75/M tokens |
| Output | $3.75/M tokens |
| Modality | Text + Image + Audio |
| Created | August 30, 2026 |
Description: Google’s latest fast model. 1M context standard, native multimodal (text, image, audio), built-in web search.
Why It’s #9: 13x cheaper than GPT-6 Astra. Native multimodal. Google’s fastest model.
#10: GPT-6 Astra (OpenAI)
Weekly Tokens: 4.2T | Rank: #10
| Spec | Value |
|---|---|
| Context | 1,050,000 tokens |
| Input | $10.00/M tokens |
| Output | $50.00/M tokens |
| Modality | Text + Image + File |
| Created | September 4, 2026 |
Description: GPT-6 Astra is OpenAI’s flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and complex reasoning. Mandatory reasoning with 5 effort levels (max, xhigh, high, medium, low).
Why It’s #10: Most expensive model in top 10. Best reasoning quality. For tasks that need the absolute best, regardless of cost.
Summary Table
| Rank | Model | Weekly Tokens | Input | Output | Context |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna | 16.1T | $0.20 | $1.20 | 1.05M |
| 2 | DeepSeek V4.1 Flash | 12.5T | $0.15 | $0.60 | 1.05M |
| 3 | Tencent Hy4 | 12.4T | $0.834 | $2.501 | 1.05M |
| 4 | Z.ai GLM 5.3 Flash | 11.9T | $0.075 | $0.25 | 1.31M |
| 5 | DeepSeek V4 Flash 0731 | 10.9T | $0.056 | $0.177 | 1.31M |
| 6 | Xiaomi MiMo-V2.5 | 7.77T | $0.119 | $0.238 | 1.05M |
| 7 | Tencent Hy3 | 5.05T | 25% off | 25% off | 256K |
| 8 | DeepSeek V4 Flash Vision | 4.8T | $0.05 | $0.15 | 1.05M |
| 9 | Gemini 3.8 Flash | 4.5T | $0.75 | $3.75 | 1.05M |
| 10 | GPT-6 Astra | 4.2T | $10.00 | $50.00 | 1.05M |
Key Takeaways
- Price = Usage: Luna ($0.20) gets 4x more tokens than Astra ($10)
- DeepSeek dominates: 3 models in top 5, 4 models in top 8
- Chinese models rising: Tencent, Z.ai, Xiaomi all in top 7
- 1M context standard: 9 of 10 models have 1M+ context
- GPT-6 is niche: Only #10 despite being “the best” — users choose value
- Usage ≠ Intelligence: Claude Fable 5.1 has the highest intelligence index (53.4) but sits outside the top 10 by usage
- Batch pricing halves costs: GPT-6 Astra batch is $5/$25 vs $10/$50 standard — check batch before paying full price
- Watch the long-prompt tax: GPT-6 Astra and Luna cost 2x above 272K prompt tokens
- Free tiers are real: DeepSeek V4 Flash 0731:free offers intelligence 34.5 at $0
Data: OpenRouter API + Top Weekly rankings, October 2025. Intelligence indices via Artificial Analysis.
Detailed Analysis
- +Complete data: token counts, pricing, context, capabilities
- +Based on real OpenRouter API data
- +Shows what models are actually being used in production
- −Data is from October 2025 and changes weekly
- −Only includes models on OpenRouter (not direct API usage)
Complete data for the top 10 LLM models by weekly token consumption on OpenRouter.
View Pricing →