OpenRouter: Unified LLM API with Routing and Fallbacks

A unified proxy layer over multiple LLM providers with cost-optimised routing, fallback chains, and per-key security controls.

You are building an app that uses multiple AI features. Voice-to-text with OpenAI's Whisper. Text reasoning with Claude. Some experimental features with open-source models from Llama or Mistral. Each provider needs its own API key. Each has different API format, different error handling, different rate limits.

Your codebase starts having separate code paths for each provider. Your .env file has six keys. Your CI/CD pipeline passes secrets for OpenAI, Anthropic, Together AI separately. Your mobile app ships an API key. Someone extracts it. You get a $1200 bill for a service you did not use. You revoke the key, add a new one, ship an app update.

A single OpenRouter API key routing requests to OpenAI, Anthropic, and Meta model endpoints

Next month, a different key leaks from Docker. Same problem.

You realize you need a better approach. One place to manage API keys. One way to call any model. One dashboard to see all spending.

That is the real problem OpenRouter solves.


What is OpenRouter

OpenRouter is an API proxy service that sits between your application and multiple LLM providers (OpenAI, Anthropic, Meta, etc.). Instead of managing separate API keys and integrations for each provider, you use one OpenRouter API key to access 300+ models across different providers.

The service standardizes the API interface across all providers, meaning the same code can call Claude, GPT-4, or Llama without changing parameters or response handling.

What Problem Does It Solve

Multiple API keys and SDKs: When building with LLMs, you often need multiple providers. OpenAI for voice, Claude for reasoning, Llama for experiments. Each requires a separate SDK, separate credentials, separate error handling.

Key security: API keys in mobile apps, .env files, and Docker configs are vulnerable. A stolen OpenAI key can cost thousands in minutes. Managing multiple keys multiplies the attack surface.

Provider lock-in: Once you build on one provider's API, switching costs engineering time. OpenRouter lets you swap models without code changes.

Inconsistent interfaces: Different providers return different response formats, have different error codes, different rate limit behaviors. OpenRouter normalizes this.

No unified monitoring: When using multiple providers, spending and usage are scattered across different dashboards. Hard to see total LLM costs.

Competitors

Direct Provider APIs (OpenAI, Anthropic, Together AI, etc.)

  • You call the provider directly
  • Pros: Lowest latency, volume discounts, full control
  • Cons: Multiple integrations, no fallback, multiple keys to manage

Amazon Bedrock (AWS)

  • AWS-managed service offering multiple providers' models (Anthropic, Meta, Mistral, Amazon Nova) through one API
  • Pros: Enterprise features, AWS integration, IAM-based access control
  • Cons: AWS lock-in, added latency, pricing scales with AWS commitment level

LiteLLM

  • Open-source proxy that unifies LLM APIs
  • Pros: Self-hosted, no vendor dependency, free
  • Cons: Self-hosted burden, no managed fallover, no dashboard

Langchain

  • Framework layer above LLM APIs
  • Pros: Rich abstractions, agent support, community
  • Cons: Not a proxy (still need raw API keys), adds latency through framework, complex deployment

Together AI

  • Unified interface for multiple open-source models
  • Pros: Good for open-source only workloads
  • Cons: Limited to open models, no proprietary models like GPT-4 or Claude

Replicate

  • Inference platform for various models
  • Pros: Easy to use, good for production inference
  • Cons: Limited model selection, focus on image/video models

Vercel AI Gateway

  • Managed multi-provider proxy built for apps already deployed on Vercel
  • Pros: Native Vercel/Next.js integration, spend tracking, provider fallback
  • Cons: Newer product, most value tied to the Vercel platform

OpenRouter Advantages

One API key: Access 300+ models with single credentials. Real provider keys stay backend-only.

Provider independence: Switch models without code changes. Add fallback chains automatically.

Managed infrastructure: No self-hosting burden. OpenRouter handles scaling, uptime, rate limiting.

Real-time monitoring: Single dashboard shows spending by model, API usage patterns, cost alerts.

Built-in fallback: Automatic model fallback if primary provider is down or rate-limited.

Standardized responses: Same response format across all providers. No provider-specific error handling needed.

Easy integration: Works with existing code through simple endpoint and header changes.

No markup on token pricing: Provider rates pass through at list price. The 5.5% fee applies only to credit purchases, not per-request inference.

OpenRouter Disadvantages

Credit-purchase fee: Buying credits costs 5.5% (minimum $0.80 per transaction, 5% for crypto) — small top-ups carry a much higher effective rate. BYOK usage is free up to 1M requests/month, then 5% of the equivalent platform cost.

Latency: Adds a routing hop on top of the provider's own response time — noticeable if you are chasing sub-100ms first-token latency.

Single point of failure: If OpenRouter is down, all LLM calls fail.

No volume discounts: Cannot negotiate pricing with providers.

Limited rate limits: Rate limits are OpenRouter's, not the underlying provider's.


Implementation

Mobile App (iOS/Swift)

Direct OpenAI API call without OpenRouter:

let apiKey = Bundle.main.infoDictionary?["OPENAI_API_KEY"] as? String

var request = URLRequest(url: URL(string: "https://api.openai.com/v1/audio/transcriptions")!)
request.setValue("Bearer \(apiKey)", forHTTPHeaderField: "Authorization")

let task = URLSession.shared.dataTask(with: request) { data, response, error in
    if let data = data {
        let transcription = try JSONDecoder().decode(Transcription.self, from: data)
    }
}
task.resume()

With OpenRouter:

let openRouterKey = Bundle.main.infoDictionary?["OPENROUTER_KEY"] as? String

var request = URLRequest(url: URL(string: "https://openrouter.ai/api/v1/audio/transcriptions")!)
request.setValue("Bearer \(openRouterKey)", forHTTPHeaderField: "Authorization")
request.setValue("https://yourdomain.com", forHTTPHeaderField: "HTTP-Referer")

let task = URLSession.shared.dataTask(with: request) { data, response, error in
    if let data = data {
        let transcription = try JSONDecoder().decode(Transcription.self, from: data)
    }
}
task.resume()

Only change: endpoint URL and API key. Response format stays identical.

Backend Proxy Layer

Create one central endpoint for all LLM calls:

// services/llm-proxy.js

const router = require('express').Router();
const axios = require('axios');

const OPENROUTER_KEY = process.env.OPENROUTER_API_KEY;
const APP_DOMAIN = process.env.APP_DOMAIN;

// Transcription endpoint
router.post('/transcribe', async (req, res) => {
  const { audioUrl } = req.body;
  
  try {
    const response = await axios.post(
      'https://openrouter.ai/api/v1/audio/transcriptions',
      { url: audioUrl },
      {
        headers: {
          'Authorization': `Bearer ${OPENROUTER_KEY}`,
          'HTTP-Referer': APP_DOMAIN
        }
      }
    );
    res.json(response.data);
  } catch (error) {
    res.status(error.response?.status || 500).json({ error: error.message });
  }
});

// Chat completion endpoint
router.post('/completions', async (req, res) => {
  const { prompt, model = 'anthropic/claude-sonnet-5' } = req.body;
  
  try {
    const response = await axios.post(
      'https://openrouter.ai/api/v1/chat/completions',
      {
        model: model,
        messages: [{ role: 'user', content: prompt }],
        max_tokens: 1024
      },
      {
        headers: {
          'Authorization': `Bearer ${OPENROUTER_KEY}`,
          'HTTP-Referer': APP_DOMAIN
        }
      }
    );
    res.json(response.data);
  } catch (error) {
    res.status(error.response?.status || 500).json({ error: error.message });
  }
});

// Fallback chain with retries
router.post('/completions-with-fallback', async (req, res) => {
  const { prompt } = req.body;
  const models = [
    'anthropic/claude-sonnet-5',
    'anthropic/claude-opus-5',
    'openai/gpt-5'
  ];
  
  const maxRetries = 3;
  const baseDelayMs = 1000;
  
  for (const model of models) {
    for (let attempt = 0; attempt < maxRetries; attempt++) {
      try {
        const response = await axios.post(
          'https://openrouter.ai/api/v1/chat/completions',
          {
            model: model,
            messages: [{ role: 'user', content: prompt }],
            max_tokens: 1024
          },
          {
            headers: {
              'Authorization': `Bearer ${OPENROUTER_KEY}`,
              'HTTP-Referer': APP_DOMAIN
            }
          }
        );
        return res.json(response.data);
      } catch (error) {
        if (error.response?.status === 429 || error.response?.status === 503) {
          const delayMs = baseDelayMs * Math.pow(2, attempt);
          await new Promise(r => setTimeout(r, delayMs));
          continue;
        }
        break;
      }
    }
  }
  
  res.status(500).json({ error: 'All models failed' });
});

module.exports = router;

Environment Configuration

Create .env with one OpenRouter key:

OPENROUTER_API_KEY=your_key_here
APP_DOMAIN=https://yourdomain.com

Real provider keys stay in AWS Secrets Manager or similar, only accessible by backend services.

Model slugs shift as providers ship new versions — the ones above are current as of this post's publish date. Check openrouter.ai/models for the live catalog before deploying.

Request Examples

Chat completions:

curl -X POST http://localhost:3000/llm/completions \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Explain quantum computing",
    "model": "anthropic/claude-sonnet-5"
  }'

With fallback chain:

curl -X POST http://localhost:3000/llm/completions-with-fallback \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Explain quantum computing"
  }'

OpenRouter vs. Competitors

FeatureOpenRouterDirect APIsAmazon BedrockLiteLLMTogether AI
Multiple providersYesNoYesYesNo (open-source only)
Unified interfaceYesNoYesYesYes
Managed serviceYesNoYesNoYes
Cost overheadCredit fee only (~5.5%)*NoneAWS pricing + marginNoneNone
Added latencyRouting hopBaselineAWS routing hopDepends on your infraRouting hop
Real-time monitoringYesLimitedYesNoYes
Fallback chainsYesManualManualNoYes
Volume discountsNoYesYesN/AYes
Self-hosted optionNoN/ANoYesNo
Proprietary modelsYesYesYesYesNo

*OpenRouter's fee applies to credit purchases (minimum $0.80/transaction), not as a markup on top of provider token pricing. BYOK usage is free up to 1M requests/month, then 5% of the equivalent platform cost.

When to Use

Use OpenRouter when:

  • Managing multiple LLM providers in one codebase
  • Client apps need API access (mobile apps, browser)
  • Key security and centralized management matter
  • Fallback chains add value (avoiding provider outages)
  • Response format consistency is important
  • Volume is under 1M requests per month
  • Want managed infrastructure without self-hosting

Use direct APIs when:

  • Committed to single provider for long-term
  • Volume is >1M requests per month (cost matters)
  • Sub-100ms latency is critical
  • Qualify for volume discounts
  • Predictable, stable workload
  • Need absolute lowest cost

Use Amazon Bedrock when:

  • Deep AWS integration required
  • Enterprise compliance and features needed
  • Accept higher cost for managed enterprise experience

Use LiteLLM when:

  • Self-hosting is acceptable
  • Want open-source solution
  • Cost optimization is critical

The Real Talk

OpenRouter earns its place the moment you have more than one provider in production and a client — mobile app, browser — that cannot hold a real API key. The pitch is not that it is cheaper; once you count the credit fee, it usually is not. The pitch is that key management and fallback logic stop being your problem.

The number that actually matters: under 1M requests a month on BYOK, the fee is a rounding error next to the engineering time saved. Past that volume, run the math on direct APIs before assuming OpenRouter still wins.

Start on OpenRouter, watch the dashboard, and move the providers you have outgrown to direct APIs once volume justifies the switch. You do not have to pick one forever.