OpenRouter: The Unified API for All LLMs
OpenRouter is a unified API gateway that provides access to 200+ AI models from every major provider through a single endpoint. It handles routing, fallovers, and billing automatically.
What Makes OpenRouter Different
Single Endpoint
Instead of managing multiple API keys and endpoints for different providers, you use one API key and one endpoint. OpenRouter routes your requests to the appropriate provider automatically.
# Instead of this:
client = Anthropic(api_key="sk-anthropic-...")
# You do this:
client = OpenRouter(api_key="sk-openrouter-...")
Automatic Routing
OpenRouter can automatically select the best model for your request based on:
- Price: Cheaper models when quality allows
- Availability: Route around provider outages
- Performance: Lowest latency for your region
- Capabilities: Match model strengths to task requirements
Fallback Chains
Define fallback chains for production reliability:
Primary: Claude 3.5 Sonnet
Fallback 1: GPT-4o
Fallback 2: Claude 3 Haiku
Fallback 3: Llama 3.1 405B
If your primary provider fails or is rate-limited, OpenRouter automatically tries the next provider in your chain.
Prompt Caching
OpenRouter supports Anthropic’s prompt caching and OpenAI’s equivalent:
- Repeated system prompts cache automatically
- Up to 90% cost reduction on cached tokens
- No configuration needed — works out of the box
Model Catalog
The model catalog shows real-time data for every model:
- Current pricing per token
- Context window size
- Provider availability
- Recent latency metrics
- Specialization tags (code, reasoning, multimodal)
Compliance Routing
Set provider preferences based on compliance needs:
- Block specific providers by name
- Require data residency in specific regions
- Only use providers with SOC 2 compliance
- Prefer open-source providers for data sensitivity
Getting Started
pip install openrouter
from openrouter import OpenRouter
client = OpenRouter(api_key="sk-or-...")
response = client.chat.completions.create(
model="anthropic/claude-3.5-sonnet",
messages=[{"role": "user", "content": "Hello"}]
)
Use Cases
1. Development Flexibility
Switch models in development without changing code. Test on GPT-4, deploy on Claude, fallback on Llama.
2. Cost Optimization
Automatically route simple tasks to cheaper models and complex tasks to premium models.
3. High Availability
Never go down because of provider outages. OpenRouter routes around failures automatically.
4. Model Evaluation
Compare models side-by-side on your actual workloads without managing multiple integrations.
The Verdict
OpenRouter is the clear choice when you want:
- Maximum model flexibility
- Production reliability
- Cost optimization
- Minimal vendor lock-in
The small markup is worth it for the routing, fallovers, and analytics alone. For teams serious about AI, it should be part of your infrastructure.
Detailed Analysis
Model Catalog
CoreBrowse 200+ models with real-time pricing, context windows, and performance data.
Auto-Routing
CoreAutomatically select the best model based on price, availability, and performance.
Fallbacks
CoreDefine fallback chains for high availability. If primary provider fails, automatically try alternatives.
Prompt Caching
OptimizationAutomatic prompt caching for Anthropic and OpenAI models. Reduces costs by up to 90%.
Usage Analytics
MonitoringDetailed cost tracking, token usage, and performance metrics per model and provider.
Direct Provider Access
AdvancedBypass OpenRouter and connect directly to providers for lower latency.
Model Preference
CustomizationSet preferred providers and block specific ones based on compliance needs.
Model Discovery
Find the right model for any task with filtering and comparison tools.
Cost Optimization
Switch models based on price/performance for each task type.
High Availability
Define fallback chains for production reliability.
Compliance Routing
Block providers based on data residency or compliance requirements.
Testing & Evaluation
Compare models side-by-side on your actual workloads.
OpenAI
Direct passthroughGPT-4, GPT-4o, o1, o3, and fine-tuned models.
Best for production quality and reliability.
Anthropic
Direct passthroughClaude 3.5 Sonnet, Opus, Haiku, and Claude 3 models.
Best for safety-critical applications.
Gemini 1.5 Pro and Flash models.
Best for multimodal and long-context.
Meta (Llama)
Varies by providerLlama 3.1 405B, 70B, 8B and Code Llama variants.
Best open-source model access.
Mistral
Varies by providerMixtral, Mistral Large, Mistral Small, and specialized models.
Best for EU compliance.
DeepSeek
$0.14-$0.28/1M tokensDeepSeek V3 and R1 reasoning model.
Best value for reasoning tasks.
- +One API for 200+ models — no provider lock-in
- +Automatic failover for high availability
- +Real-time pricing comparison across providers
- +Prompt caching for 90% cost reduction on repeated context
- +Model preferences for compliance routing
- +Detailed analytics and cost tracking
- +Free tier available for testing
- −Small markup on top of provider pricing
- −Can have slightly higher latency than direct provider
- −Not all providers support all features
- −Complex routing rules can be hard to debug
OpenRouter passes through provider pricing with a small markup (typically 5-15%). No additional fees for routing or fallbacks. Free credits available for new accounts.
View Pricing →