INTEL DOSSIER|CLASSIFIED
DECRYPTED
File #23 • October 11, 2025

The Latest LLM Models: October 2025 Landscape

llmai-modelslatest
Authorcoderunner
Categoryllm
StatusPUBLISHED
ClearancePUBLIC
//The LLM landscape as of October 2025, based on live OpenRouter API data. 445 models available, with GPT-6 Astra, Claude Fable 5.1, and DeepSeek V4.1 Flash leading their respective categories.

This post is based on live OpenRouter API data from October 2025.

The Current State

Metric Value
Total Models on OpenRouter 445
Max Context 1,050,000 tokens
Cheapest Input Free (Ling 3.0 Flash VL)
Most Expensive $50/1M output (GPT-6 Astra)

September 2025 Releases

The last week of September brought two major releases that shift the coding landscape:

Claude Sonnet 4.5 (Anthropic)

Released September 29, 2025. Anthropic’s best coding model — “the strongest model for building complex agents” and “best at using computers.”

Spec Value
Context 200K
Input $3/1M tokens
Output $15/1M tokens
SWE-bench Verified State-of-the-art
OSWorld (computer use) 61.4% (up from 42.2%)

Key features:

  • 30+ hours of autonomous coding without performance degradation
  • Checkpoints in Claude Code (save/roll back progress)
  • Native VS Code extension
  • Claude Agent SDK released — build your own agents with the same infrastructure
  • Context editing + memory tool for longer sessions
  • Code execution and file creation in conversations

Available via: claude-sonnet-4-5 on Claude API, Claude Code, and Claude apps.


GPT-5-Codex (OpenAI)

Announced September 15, 2025. GPT-5 optimized for agentic coding — equally proficient at quick interactive sessions and long, complex independent tasks.

Spec Value
Context 1M+ (via Codex)
Input Same as GPT-5
Output Same as GPT-5
Code review Built-in, catches critical bugs

Key features:

  • Available in Codex CLI, IDE extension, web, and API
  • Trained for real-world software engineering
  • Code review capability matches intent to diff, reasons over entire codebase
  • Sandboxed execution (network disabled by default)
  • MCP support for external tool integration
  • Image support for design specs and UI bugs
  • Three approval modes: read-only, auto, full access

Available via: Codex CLI, VS Code extension, ChatGPT subscription, or API key.


Current Frontier Models

GPT-6 Series (OpenAI)

Model Context Input $/1M Output $/1M Notes
GPT-6 Astra 1.05M $10 $50 Best reasoning, mandatory reasoning
GPT-6 Astra Pro 1.05M $10 $50 Production variant with reasoning.mode=pro
GPT-5.6 Sol 1.05M $2 $10 Moderated, knowledge cutoff Feb 2026
GPT-5.6 Terra 1.05M $2 $12 Not moderated
GPT-5.6 Luna 1.05M $0.20 $1.20 Budget option, 16.1T weekly tokens

Key Features:

  • 1M context window standard
  • Web search available ($0.01/request)
  • Reasoning: 5 effort levels (max, xhigh, high, medium, low)
  • Knowledge cutoff: February 2026

Claude Fable 5.1 (Anthropic)

Model Context Input $/1M Output $/1M Notes
Claude Fable 5.1 1M $10 $50 Latest flagship
Claude Fable 5.1 Batch 1M $5 $25 Batch processing

Key Features:

  • Latest Claude model (replaces Claude 3.5/4 series)
  • 1M context
  • Best for coding and safety
  • Computer use built-in

Gemini 3.8 Series (Google)

Model Context Input $/1M Output $/1M Notes
Gemini 3.8 Flash 1.05M $0.75 $3.75 Fast, cheap

Key Features:

  • Latest Google model (replaces Gemini 1.5/2.5)
  • 1M context standard
  • Native multimodal
  • Best for long documents

DeepSeek V4.1 (DeepSeek)

Model Context Input $/1M Output $/1M Notes
DeepSeek V4.1 Flash 1.05M $0.15-$0.30 $0.60-$1.20 Time-based pricing
DeepSeek V4 Pro 1.05M $0.70 $2.96 Higher quality

Key Features:

  • Causal Encoder-Decoder (CED) architecture
  • 8B params on input, 16B on output
  • Time-of-day pricing (cheaper off-peak, free weekends)
  • Best value for reasoning

Free Tier Models ($0)

Zero-cost models on OpenRouter with real capability. Full ranking in our Free Tier Revolution post — 25 free models total. The standouts:

Model Context Intelligence Index Notes
DeepSeek V4 Flash 0731:free 1.05M 34.5 Best free model — coding 69.1, rivals paid Luna
GLM 5.2:free 32K 34.0 Deep reasoning, project-level engineering
Qwen3.8 27B:free 262K 33.9 Agentic 46.5 beats paid GPT-5.6 Luna — video input
Inkling Small:free 1.05M 26.1 Thinking Machines, text+image+audio
Nemotron 3 Ultra:free 1M 23.4 NVIDIA 550B MoE, Transformer-Mamba hybrid
Ling 3.0 Flash VL:free 262K 25 Vision + video — free has 2x context of paid
Ling 3.0 Flash Fin:free 262K 23 Finance-tuned
Ling 3.0 Flash Sante:free 262K — Healthcare-tuned
Nex-N2.5 Mini:free 262K — Agentic coding, visual feedback loop
Nex-N2.5 Pro:free 262K — Agentic coding, verified outcomes

The headline: DeepSeek V4 Flash 0731:free delivers intelligence 34.5 — better than paid tiers of several models in our benchmarks — at zero cost with 1M context. Qwen3.8 27B:free’s agentic score beats paid models. The catch is rate limits (~50 requests/day without credits), not capability.

Plus openrouter/free — a router that randomly selects from free models, supporting tools and structured outputs.


Other Notable Models

Model Context Input $/1M Output $/1M Notes
Qwen 3.8 Max 1M $2 $6 Alibaba, multilingual
Sakana Fugu Ultra 1M $5 $30 Multi-agent orchestration
Ling 3.0 Flash VL 128K $0.06 $0.18 Free version available
Meta Muse Spark 1.3 1.05M $0.10 $0.20 Open-source
Mercury 2.5 260K $0.04 $0.15 Fastest inference

What Changed (Old → New)

Category Old Model Current Model Improvement
OpenAI Flagship GPT-4o GPT-6 Astra 1M context, reasoning
Anthropic Flagship Claude 3.5 Sonnet Claude Fable 5.1 Better coding, 1M context
Google Flagship Gemini 1.5 Pro Gemini 3.8 Flash Faster, 1M context
Best Value DeepSeek V3 DeepSeek V4.1 2x better, still 1/50th cost
Open Source Llama 3.1 Meta Muse Spark Newer, cheaper
Budget GPT-4o-mini GPT-5.6 Luna 10x cheaper

What’s Actually Worth Using (October 2025)

Best Overall

  1. GPT-6 Astra — Best reasoning, 1M context, tools
  2. Claude Fable 5.1 — Best for coding and safety
  3. Gemini 3.8 Flash — Best for long documents

Best Value

  1. DeepSeek V4.1 Flash — 90% of GPT-6 quality at 1/77th cost
  2. Ling 3.0 Flash VL — Free tier available
  3. GPT-5.6 Luna — $0.20/1M input

Best for Specific Tasks

  • Coding: Claude Fable 5.1, GPT-6 Astra
  • Math/Reasoning: GPT-6 Astra, DeepSeek V4.1
  • Long Documents: Gemini 3.8 Flash, GPT-6 Astra
  • Multimodal: GPT-6 Astra, Ling 3.0 Flash VL (video!)
  • High Volume: DeepSeek V4.1, Mercury 2.5
  • Open Source: Meta Muse Spark, Ling 3.0

Pricing Comparison (Per 1K Input + 500 Output Tokens)

Model Cost/Task Context
GPT-6 Astra $0.035 1.05M
Claude Fable 5.1 $0.035 1M
GPT-5.6 Sol $0.009 1.05M
DeepSeek V4.1 Flash $0.00045 1.05M
Ling 3.0 Flash VL $0.00015 128K
Meta Muse Spark 1.3 $0.00020 1.05M
Mercury 2.5 $0.00013 260K

OpenRouter Auto-Router

Use openrouter/auto to automatically get the best model for your prompt. It routes based on what the OpenRouter community spends on similar tasks.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer sk-or-..." \
  -d '{"model": "openrouter/auto", "messages": [{"role": "user", "content": "Hello"}]}'

How to Stay Updated

  1. OpenRouter Models Page: https://openrouter.ai/models
  2. OpenRouter API: curl https://openrouter.ai/api/v1/models
  3. This Blog: We update posts when models change significantly

The LLM landscape moves fast. What’s best today may not be best in 3 months. Always verify current pricing and capabilities.

Detailed Analysis

Visual Presentations

Timeline of LLM Releases 2025-2026
Timeline of LLM Releases 2025-2026

Major model releases over the past year

Strengths3 PROS
  • +Based on live OpenRouter API data
  • +Covers current frontier models from all major providers
  • +Includes both open-source and closed-source models
Weaknesses1 CONS
  • −LLM landscape changes weekly — verify with OpenRouter
PricingVarious

Pricing varies significantly. See individual models for details. Data verified via OpenRouter API.

View Pricing →
END OF FILE|DISTRIBUTION: UNLIMITED
← Back to Archive