The Latest LLM Models: October 2025 Landscape
This post is based on live OpenRouter API data from October 2025.
The Current State
| Metric | Value |
|---|---|
| Total Models on OpenRouter | 445 |
| Max Context | 1,050,000 tokens |
| Cheapest Input | Free (Ling 3.0 Flash VL) |
| Most Expensive | $50/1M output (GPT-6 Astra) |
September 2025 Releases
The last week of September brought two major releases that shift the coding landscape:
Claude Sonnet 4.5 (Anthropic)
Released September 29, 2025. Anthropic’s best coding model — “the strongest model for building complex agents” and “best at using computers.”
| Spec | Value |
|---|---|
| Context | 200K |
| Input | $3/1M tokens |
| Output | $15/1M tokens |
| SWE-bench Verified | State-of-the-art |
| OSWorld (computer use) | 61.4% (up from 42.2%) |
Key features:
- 30+ hours of autonomous coding without performance degradation
- Checkpoints in Claude Code (save/roll back progress)
- Native VS Code extension
- Claude Agent SDK released — build your own agents with the same infrastructure
- Context editing + memory tool for longer sessions
- Code execution and file creation in conversations
Available via: claude-sonnet-4-5 on Claude API, Claude Code, and Claude apps.
GPT-5-Codex (OpenAI)
Announced September 15, 2025. GPT-5 optimized for agentic coding — equally proficient at quick interactive sessions and long, complex independent tasks.
| Spec | Value |
|---|---|
| Context | 1M+ (via Codex) |
| Input | Same as GPT-5 |
| Output | Same as GPT-5 |
| Code review | Built-in, catches critical bugs |
Key features:
- Available in Codex CLI, IDE extension, web, and API
- Trained for real-world software engineering
- Code review capability matches intent to diff, reasons over entire codebase
- Sandboxed execution (network disabled by default)
- MCP support for external tool integration
- Image support for design specs and UI bugs
- Three approval modes: read-only, auto, full access
Available via: Codex CLI, VS Code extension, ChatGPT subscription, or API key.
Current Frontier Models
GPT-6 Series (OpenAI)
| Model | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
| GPT-6 Astra | 1.05M | $10 | $50 | Best reasoning, mandatory reasoning |
| GPT-6 Astra Pro | 1.05M | $10 | $50 | Production variant with reasoning.mode=pro |
| GPT-5.6 Sol | 1.05M | $2 | $10 | Moderated, knowledge cutoff Feb 2026 |
| GPT-5.6 Terra | 1.05M | $2 | $12 | Not moderated |
| GPT-5.6 Luna | 1.05M | $0.20 | $1.20 | Budget option, 16.1T weekly tokens |
Key Features:
- 1M context window standard
- Web search available ($0.01/request)
- Reasoning: 5 effort levels (max, xhigh, high, medium, low)
- Knowledge cutoff: February 2026
Claude Fable 5.1 (Anthropic)
| Model | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
| Claude Fable 5.1 | 1M | $10 | $50 | Latest flagship |
| Claude Fable 5.1 Batch | 1M | $5 | $25 | Batch processing |
Key Features:
- Latest Claude model (replaces Claude 3.5/4 series)
- 1M context
- Best for coding and safety
- Computer use built-in
Gemini 3.8 Series (Google)
| Model | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
| Gemini 3.8 Flash | 1.05M | $0.75 | $3.75 | Fast, cheap |
Key Features:
- Latest Google model (replaces Gemini 1.5/2.5)
- 1M context standard
- Native multimodal
- Best for long documents
DeepSeek V4.1 (DeepSeek)
| Model | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 1.05M | $0.15-$0.30 | $0.60-$1.20 | Time-based pricing |
| DeepSeek V4 Pro | 1.05M | $0.70 | $2.96 | Higher quality |
Key Features:
- Causal Encoder-Decoder (CED) architecture
- 8B params on input, 16B on output
- Time-of-day pricing (cheaper off-peak, free weekends)
- Best value for reasoning
Free Tier Models ($0)
Zero-cost models on OpenRouter with real capability. Full ranking in our Free Tier Revolution post — 25 free models total. The standouts:
| Model | Context | Intelligence Index | Notes |
|---|---|---|---|
| DeepSeek V4 Flash 0731:free | 1.05M | 34.5 | Best free model — coding 69.1, rivals paid Luna |
| GLM 5.2:free | 32K | 34.0 | Deep reasoning, project-level engineering |
| Qwen3.8 27B:free | 262K | 33.9 | Agentic 46.5 beats paid GPT-5.6 Luna — video input |
| Inkling Small:free | 1.05M | 26.1 | Thinking Machines, text+image+audio |
| Nemotron 3 Ultra:free | 1M | 23.4 | NVIDIA 550B MoE, Transformer-Mamba hybrid |
| Ling 3.0 Flash VL:free | 262K | 25 | Vision + video — free has 2x context of paid |
| Ling 3.0 Flash Fin:free | 262K | 23 | Finance-tuned |
| Ling 3.0 Flash Sante:free | 262K | — | Healthcare-tuned |
| Nex-N2.5 Mini:free | 262K | — | Agentic coding, visual feedback loop |
| Nex-N2.5 Pro:free | 262K | — | Agentic coding, verified outcomes |
The headline: DeepSeek V4 Flash 0731:free delivers intelligence 34.5 — better than paid tiers of several models in our benchmarks — at zero cost with 1M context. Qwen3.8 27B:free’s agentic score beats paid models. The catch is rate limits (~50 requests/day without credits), not capability.
Plus openrouter/free — a router that randomly selects from free models, supporting tools and structured outputs.
Other Notable Models
| Model | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
| Qwen 3.8 Max | 1M | $2 | $6 | Alibaba, multilingual |
| Sakana Fugu Ultra | 1M | $5 | $30 | Multi-agent orchestration |
| Ling 3.0 Flash VL | 128K | $0.06 | $0.18 | Free version available |
| Meta Muse Spark 1.3 | 1.05M | $0.10 | $0.20 | Open-source |
| Mercury 2.5 | 260K | $0.04 | $0.15 | Fastest inference |
What Changed (Old → New)
| Category | Old Model | Current Model | Improvement |
|---|---|---|---|
| OpenAI Flagship | GPT-4o | GPT-6 Astra | 1M context, reasoning |
| Anthropic Flagship | Claude 3.5 Sonnet | Claude Fable 5.1 | Better coding, 1M context |
| Google Flagship | Gemini 1.5 Pro | Gemini 3.8 Flash | Faster, 1M context |
| Best Value | DeepSeek V3 | DeepSeek V4.1 | 2x better, still 1/50th cost |
| Open Source | Llama 3.1 | Meta Muse Spark | Newer, cheaper |
| Budget | GPT-4o-mini | GPT-5.6 Luna | 10x cheaper |
What’s Actually Worth Using (October 2025)
Best Overall
- GPT-6 Astra — Best reasoning, 1M context, tools
- Claude Fable 5.1 — Best for coding and safety
- Gemini 3.8 Flash — Best for long documents
Best Value
- DeepSeek V4.1 Flash — 90% of GPT-6 quality at 1/77th cost
- Ling 3.0 Flash VL — Free tier available
- GPT-5.6 Luna — $0.20/1M input
Best for Specific Tasks
- Coding: Claude Fable 5.1, GPT-6 Astra
- Math/Reasoning: GPT-6 Astra, DeepSeek V4.1
- Long Documents: Gemini 3.8 Flash, GPT-6 Astra
- Multimodal: GPT-6 Astra, Ling 3.0 Flash VL (video!)
- High Volume: DeepSeek V4.1, Mercury 2.5
- Open Source: Meta Muse Spark, Ling 3.0
Pricing Comparison (Per 1K Input + 500 Output Tokens)
| Model | Cost/Task | Context |
|---|---|---|
| GPT-6 Astra | $0.035 | 1.05M |
| Claude Fable 5.1 | $0.035 | 1M |
| GPT-5.6 Sol | $0.009 | 1.05M |
| DeepSeek V4.1 Flash | $0.00045 | 1.05M |
| Ling 3.0 Flash VL | $0.00015 | 128K |
| Meta Muse Spark 1.3 | $0.00020 | 1.05M |
| Mercury 2.5 | $0.00013 | 260K |
OpenRouter Auto-Router
Use openrouter/auto to automatically get the best model for your prompt. It routes based on what the OpenRouter community spends on similar tasks.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer sk-or-..." \
-d '{"model": "openrouter/auto", "messages": [{"role": "user", "content": "Hello"}]}'
How to Stay Updated
- OpenRouter Models Page: https://openrouter.ai/models
- OpenRouter API:
curl https://openrouter.ai/api/v1/models - This Blog: We update posts when models change significantly
The LLM landscape moves fast. What’s best today may not be best in 3 months. Always verify current pricing and capabilities.
Detailed Analysis
Visual Presentations

Major model releases over the past year
- +Based on live OpenRouter API data
- +Covers current frontier models from all major providers
- +Includes both open-source and closed-source models
- −LLM landscape changes weekly — verify with OpenRouter
Pricing varies significantly. See individual models for details. Data verified via OpenRouter API.
View Pricing →