Top 10 Weekly LLM Models: Complete Reviews
Complete reviews of the top 10 LLM models by weekly token consumption. Based on OpenRouter data, community feedback, and pricing analysis.
#1: GPT-5.6 Luna (OpenAI)
Weekly Tokens: 16.1T | Rank: #1
| Spec | Value |
|---|---|
| Context | 1,050,000 tokens |
| Input | $0.20/M tokens |
| Output | $1.20/M tokens |
| Modality | Text + Image + File |
| Created | July 9, 2026 |
Strengths
- 1M context window — handle entire codebases in one prompt
- Fast inference — optimized for low-latency applications
- Multimodal — text, image, and file inputs
- Cheapest in GPT-6 family — 50x cheaper than GPT-6 Astra
- Good enough quality — capable reasoning for most tasks
Weaknesses
- Not the best at reasoning — GPT-6 Astra significantly better for complex logic
- Limited agentic capabilities — not designed for complex multi-step workflows
- Newer model — less community testing and tooling support
Best Use Cases
- High-volume chat applications
- Classification and extraction tasks
- Lightweight agentic workflows
- Content generation at scale
- Cost-sensitive applications
Verdict
The volume king. Luna proves that for most tasks, good enough quality at a fraction of the cost wins. It’s #1 because 90% of AI tasks don’t need frontier reasoning — they need fast, cheap, reliable responses.
#2: DeepSeek V4.1 Flash (DeepSeek)
Weekly Tokens: 12.5T | Rank: #2
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.15/M tokens |
| Output | $0.60/M tokens |
| Modality | Text + Image |
| Created | September 10, 2026 |
Strengths
- Best value for coding — rivals GPT-6 on many coding tasks
- CED architecture — first Causal Encoder-Decoder from DeepSeek
- Asymmetric MoE — 8B active on input, 16B on output (552B total)
- Compressed KV caching — 75% less cache memory than V3
- Time-based pricing — cheaper off-peak, free weekends
Weaknesses
- Limited modality — text + image only (no file input)
- Newer model — less community tooling
- Time-based pricing confusion — costs vary by time of day
Best Use Cases
- Coding agents and IDE integrations
- Terminal and computer-use agents
- Long-horizon multi-step tasks
- Cost-sensitive coding workflows
- High-volume batch processing
Verdict
The value champion. DeepSeek V4.1 Flash proves you don’t need to spend OpenAI prices for great coding results. The CED architecture and compressed KV caching make it ideal for agentic workloads.
#3: Tencent Hy4 Preview (Tencent)
Weekly Tokens: 12.4T | Rank: #3
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.834/M tokens |
| Output | $2.501/M tokens |
| Modality | Text |
| Created | August 28, 2026 |
Strengths
- Strong on agentic workflows — designed for complex tool-use
- 49B active / 770B total MoE — large model, efficient inference
- 1M context — suitable for long documents
- Multi-step execution — built for sustained workflows
Weaknesses
- Preview status — may have bugs and instability
- Text only — no image or file input
- Limited availability — newer model, less widespread adoption
Best Use Cases
- Complex coding agents
- Multi-step tool-use workflows
- Productivity applications
- Automation tasks
Verdict
The agentic contender. Tencent’s Hy4 is purpose-built for agentic workflows. It’s impressive that a Chinese model is competing directly with OpenAI and DeepSeek on complex tasks.
#4: Z.ai GLM 5.3 Flash (Z.ai)
Weekly Tokens: 11.9T | Rank: #4
| Spec | Value |
|---|---|
| Context | 1,310,720 tokens |
| Input | $0.075/M tokens |
| Output | $0.25/M tokens |
| Modality | Text |
| Created | August 26, 2026 |
Strengths
- Largest context in top 10 — 1.31M tokens
- Hybrid attention — sparse + linear for efficient long-context
- 133x cheaper than GPT-6 — $0.075 vs $10 input
- 50% discount on OpenRouter — even cheaper currently
- Strong in Academia — ranked #1
Weaknesses
- Text only — no image or file input
- Less known — smaller community and tooling ecosystem
- Newer model — less battle-tested
Best Use Cases
- Long document analysis
- Academic research
- Legal document review
- Any task requiring maximum context
Verdict
The context king. GLM 5.3 Flash has the largest context window in the top 10 and is the cheapest per token. For tasks that need to process entire documents, it’s the clear choice.
#5: DeepSeek V4 Flash 0731 (DeepSeek)
Weekly Tokens: 10.9T | Rank: #5
| Spec | Value |
|---|---|
| Context | 1,310,720 tokens |
| Input | $0.056/M tokens |
| Output | $0.177/M tokens |
| Modality | Text |
| Created | July 31, 2026 |
Strengths
- Category leader — #1 in Academia, Finance, Health, Marketing
- 1.31M context — largest context in top 10
- Cheapest in top 5 — $0.056 input
- GA release — stable, production-ready
- 13B active / 284B total MoE — efficient architecture
Weaknesses
- Text only — no image or file input
- Older than V4.1 — superseded by newer model
Best Use Cases
- Academic research
- Financial analysis
- Health/medical text processing
- Marketing content generation
- Category-specific applications
Verdict
The specialist. DeepSeek V4 Flash 0731 dominates specific categories. If you’re in academia, finance, health, or marketing, this is the model to use.
#6: Xiaomi MiMo-V2.5 (Xiaomi)
Weekly Tokens: 7.77T | Rank: #6
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.119/M tokens |
| Output | $0.238/M tokens |
| Modality | Text + Image + Audio |
| Created | April 22, 2026 |
Strengths
- Omnimodal — text, image, AND audio input
- Pro-level agentic performance at half the cost
- 1M context — suitable for long documents
- 15% discount on OpenRouter — even cheaper currently
Weaknesses
- Less known — smaller community
- Newer model — less battle-tested for production
Best Use Cases
- Multimodal applications (audio + image + text)
- Agentic coding workflows
- Video understanding tasks
- Cost-sensitive multimodal projects
Verdict
The multimodal dark horse. Xiaomi’s MiMo-V2.5 is the only omnimodal model in the top 10. For applications that need audio understanding, it’s the clear choice.
#7: Tencent Hy3 (Tencent)
Weekly Tokens: 5.05T | Rank: #7
| Spec | Value |
|---|---|
| Context | 256,000 tokens |
| Input | 25% off on OpenRouter |
| Output | 25% off on OpenRouter |
| Modality | Text |
| Created | August 28, 2026 |
Strengths
- Configurable reasoning — no-think, low-thought, high-thought modes
- 295B MoE with 21B active — large but efficient
- Agentic workflows — built for real-world production
- 25% discount on OpenRouter — good value
Weaknesses
- Smaller context — only 256K (others have 1M+)
- Newer model — less community support
Best Use Cases
- Reasoning-heavy tasks
- Agentic coding workflows
- Math and logic problems
- Configurable reasoning needs
Verdict
The reasoning specialist. Hy3’s configurable reasoning modes let you trade speed for quality as needed. Good for tasks where you sometimes need deep thinking and sometimes need fast responses.
#8: DeepSeek V4 Flash Vision (DeepSeek)
Weekly Tokens: 4.8T | Rank: #8
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.05/M tokens |
| Output | $0.15/M tokens |
| Modality | Text + Image |
| Created | July 31, 2026 |
Strengths
- 200x cheaper than GPT-6 — $0.05 vs $10 input
- Vision capable — text + image input
- 1M context — suitable for long documents
- Vision + coding — good for screenshot-to-code workflows
Weaknesses
- Experimental — may have bugs and limitations
- Vision is experimental — not as capable as dedicated vision models
Best Use Cases
- Screenshot-to-code workflows
- Image understanding in coding contexts
- Cost-sensitive vision tasks
- Document analysis with images
Verdict
The budget vision model. If you need image understanding but can’t afford GPT-6 prices, this is your best bet.
#9: Gemini 3.8 Flash (Google)
Weekly Tokens: 4.5T | Rank: #9
| Spec | Value |
|---|---|
| Context | 1,048,576 tokens |
| Input | $0.75/M tokens |
| Output | $3.75/M tokens |
| Modality | Text + Image + Audio |
| Created | August 30, 2026 |
Strengths
- Native multimodal — text, image, AND audio
- Google ecosystem — integrates with Google services
- 1M context — suitable for long documents
- Built-in web search — no extra tool needed
Weaknesses
- More expensive — $0.75 input vs $0.05-0.15 for cheaper models
- Google dependency — tied to Google infrastructure
Best Use Cases
- Multimodal applications
- Google ecosystem integration
- Tasks requiring web search
- Long document analysis
Verdict
The Google play. Gemini 3.8 Flash is Google’s fastest model and a solid choice if you’re already in the Google ecosystem.
#10: GPT-6 Astra (OpenAI)
Weekly Tokens: 4.2T | Rank: #10
| Spec | Value |
|---|---|
| Context | 1,050,000 tokens |
| Input | $10.00/M tokens |
| Output | $50.00/M tokens |
| Modality | Text + Image + File |
| Created | September 4, 2026 |
Strengths
- Best reasoning quality — OpenAI’s flagship
- Mandatory reasoning — always thinks before responding
- 5 effort levels — max, xhigh, high, medium, low
- 1M context — handle large inputs
- Multimodal — text, image, file inputs
Weaknesses
- Most expensive — $10 input, $50 output
- Slowest — mandatory reasoning adds latency
- Overkill for simple tasks — most tasks don’t need this quality
Best Use Cases
- Final output generation
- Client-facing content
- Complex reasoning tasks
- Safety-critical applications
- When cost is no object
Verdict
The premium choice. GPT-6 Astra is the best model available — but it’s 50-200x more expensive than alternatives. Use it only when you absolutely need the best.
Summary: Which Model for Which Task?
| Task | Best Model | Why |
|---|---|---|
| High-volume chat | GPT-5.6 Luna | Cheapest, fast, 1M context |
| Coding agent | DeepSeek V4.1 Flash | Best value for coding |
| Long documents | Z.ai GLM 5.3 Flash | 1.31M context, cheapest |
| Academia/Finance | DeepSeek V4 Flash 0731 | Category leader |
| Multimodal | Xiaomi MiMo-V2.5 | Only omnimodal in top 10 |
| Reasoning | Tencent Hy3 | Configurable reasoning |
| Vision | DeepSeek V4 Flash Vision | Cheapest with vision |
| Google ecosystem | Gemini 3.8 Flash | Native Google integration |
| Best quality | Claude Fable 5.1 | Highest intelligence index (53.4) |
| Agentic workflows | GLM 5.3 Flash | Agentic 51.2 at 1/111th GPT-6 price |
Intelligence Index Rankings (Artificial Analysis)
How the top 10 stack up on measured intelligence:
| Rank | Model | Intelligence | Coding | Agentic |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 53.4 | — | — |
| 2 | GPT-6 Astra | 52.8 | 76.9 | 51.5 |
| 3 | GLM 5.3 Flash | 41.9 | 71.5 | 51.2 |
| 4 | Gemini 3.8 Flash | 41.2 | 76.3 | 41.1 |
| 5 | DeepSeek V4.1 Flash | 39.5 | — | — |
| 6 | GPT-5.6 Luna | 37.5 | 71.4 | 42.7 |
| 7 | DeepSeek V4 Flash 0731 | 34.5 | 69.1 | 41.7 |
| 8 | Xiaomi MiMo-V2.5 | 22.3 | 56.8 | 17.4 |
Note: Tencent Hy4, Tencent Hy3, DeepSeek V4 Flash Vision, Gemini and others without indices haven’t been scored by Artificial Analysis yet. Absence of data ≠ absence of capability.
Reasoning Effort Defaults
Each model ships with different default thinking behavior — this affects both latency and quality:
| Model | Reasoning | Default Effort | Supported Efforts |
|---|---|---|---|
| GLM 5.3 Flash | Mandatory | MAX | max, high, low |
| GPT-6 Astra | Mandatory | medium | max, xhigh, high, medium, low |
| Gemini 3.8 Flash | Mandatory | medium | high, medium, low |
| Muse Spark 1.3 | Mandatory | medium | max, xhigh, high, medium, low, minimal |
| Sakana Fugu Ultra | Mandatory | xhigh | max, xhigh, high |
| DeepSeek V4.1 Flash | Optional | high | max, high, low |
| GPT-5.6 Luna | Optional | medium | max, xhigh, high, medium, low, none |
| Tencent Hy4 | Optional | high | high, low, none |
| Mercury 2.5 | Optional | medium | high, medium, low, none |
Why this matters: GLM 5.3 Flash defaults to MAX effort — it thinks hardest by default, explaining its outsized agentic index. GPT-6 Astra defaults to medium. If you’re comparing latency, remember you’re not comparing equal thinking time.
Cost lever: On models with optional reasoning (Luna, DeepSeek, Hy4), setting effort to none trades quality for maximum speed and minimum cost.
Data: OpenRouter API, October 2025.
Detailed Analysis
- +In-depth analysis of each top 10 model
- +Real-world strengths and weaknesses
- +Best use cases for each model
- +Pricing analysis and value comparison
- −Reviews based on limited public data
- −Performance varies by use case
- −Models update frequently
Reviews based on OpenRouter API data, community feedback, and pricing analysis.
View Pricing →