11/10/2026

Running Multiple LLM APIs in Production: Fallbacks and Rate Limits

Knowledge_seci_model

Most teams start with a single LLM provider and a single API key. It works — until the provider has an outage, a rate limit kicks in during a traffic spike, or the monthly bill makes someone ask why you're paying premium prices for requests that a cheaper model could have handled. That's the point where "call the API" stops being enough and you need an actual integration architecture.

This guide covers the three problems that show up as soon as you run more than one LLM API in production: what to do when a provider fails, how to stay under rate limits without dropping requests, and how to route traffic so you're not overpaying for simple calls.

Why a single LLM API call doesn't survive production

A direct call to one provider's endpoint is fine for a prototype. In production, three things break it:

  • Provider outages are not rare. OpenAI, Anthropic, and Google have all had multi-hour incidents. If your app calls one endpoint directly, an outage on their end is an outage on yours.
  • Rate limits are per-provider and per-model. A traffic spike that's routine for your app can hit a provider's requests-per-minute or tokens-per-minute ceiling, and the failure mode is usually a 429 with no useful retry guidance beyond "wait."
  • Not every request needs your most expensive model. Classifying a support ticket and drafting a legal summary have very different cost profiles, but a single hardcoded model call treats them the same.

The fix for all three is the same pattern: put a routing layer between your application code and the providers, so the application asks for "a completion" and the layer decides which provider, which model, and what to do if the first choice fails.

Fallbacks: what happens when a provider fails

A fallback chain means your code calls an abstraction, not a specific provider, and that abstraction tries providers in order until one succeeds.

const providers = [
  { name: "primary", call: callProviderA },
  { name: "secondary", call: callProviderB },
];

async function completeWithFallback(prompt, opts) {
  let lastError;
  for (const provider of providers) {
    try {
      return await withTimeout(provider.call(prompt, opts), opts.timeoutMs ?? 8000);
    } catch (err) {
      lastError = err;
      if (!isRetryable(err)) throw err;
    }
  }
  throw lastError;
}

A few details matter more than the happy-path code above:

  • Set a per-call timeout. A provider that's slow but not erroring will stall your fallback chain just as badly as one that's down. Don't rely on the provider SDK's default timeout — it's often too long for a user-facing request.
  • Distinguish retryable from non-retryable errors. A 500 or a timeout should trigger the next provider in the chain. A 400 (bad request, invalid prompt) will fail identically on every provider, so retrying it just adds latency for no benefit.
  • Log which provider actually served each response. When something goes wrong three days later, you need to know whether it was the primary or a fallback that answered, and how often you're falling back at all. This is the first thing worth tracing — see the dedicated guide on distributed tracing with OpenTelemetry for how to instrument calls like this one so a fallback shows up as a span, not a mystery in the logs.
  • Don't fall back silently forever. If your primary has been down for ten minutes, that's worth paging someone, not just quietly running on the backup provider indefinitely at a different cost and quality profile.

Keep the fallback chain short — two providers, three at most. Each additional hop adds latency to the worst case, and past three you're usually better off fixing the primary than adding more backups.

Rate limits: queue and back off instead of dropping requests

Rate limiting a single outbound API is a different problem from rate limiting inbound traffic to your own service. You don't control the limit, and the provider doesn't always tell you how close you are to it until you've already crossed it.

The practical approach is a token-bucket limiter in front of each provider, tuned slightly below their published ceiling:

class TokenBucket {
  constructor(ratePerSecond, burst) {
    this.rate = ratePerSecond;
    this.capacity = burst;
    this.tokens = burst;
    this.last = Date.now();
  }
  async take() {
    this.refill();
    while (this.tokens < 1) {
      await sleep(50);
      this.refill();
    }
    this.tokens -= 1;
  }
  refill() {
    const now = Date.now();
    const elapsed = (now - this.last) / 1000;
    this.tokens = Math.min(this.capacity, this.tokens + elapsed * this.rate);
    this.last = now;
  }
}

Three things make this reliable in practice rather than just in theory:

  1. Set the bucket's rate below the provider's documented limit, not at it. Token counting for LLM requests is approximate until the response comes back, so a limiter tuned exactly to the ceiling will still occasionally cross it.
  2. Respect Retry-After headers when a provider does send them. Treat the header as authoritative over your own backoff calculation — the provider knows its own state better than your client does.
  3. Queue requests instead of rejecting them, but bound the queue. An unbounded queue during a sustained spike just delays the failure and makes it harder to diagnose. A bounded queue with a clear "try again shortly" response to the caller is more honest about what's happening.

If you're running this on a cluster rather than a single process, the limiter state needs to be shared (Redis is the usual choice) — otherwise each instance thinks it has the full rate budget, and you're back to tripping the provider's actual limit with N times the traffic you intended.

Diagram of a rate limiter and fallback chain routing requests to two LLM providers with shared state across instances

Cost-aware routing: not every request needs the expensive model

Once fallbacks and rate limiting are in place, the next lever is routing requests to the right model for the task, not the same model for everything.

A simple version of this is a classifier step that runs before the main request:

function routeModel(request) {
  if (request.tokensEstimate < 500 && request.taskType === "classification") {
    return "small-model";
  }
  if (request.requiresReasoning || request.tokensEstimate > 4000) {
    return "large-model";
  }
  return "default-model";
}

Flowchart showing requests classified and routed to either a small fast model or a large reasoning model

This doesn't need to be sophisticated to pay off. Even a rough split — short classification and extraction tasks to a cheap, fast model, longer generation and reasoning tasks to a capable one — typically cuts the API bill meaningfully, because a large share of production LLM traffic is short, repetitive, and not actually reasoning-heavy.

The part teams skip is measuring it. Track cost and latency per model and per task type, not just in aggregate — aggregate numbers hide the fact that one endpoint is quietly routing everything to the expensive model regardless of the rule you wrote. Monitoring the service with Prometheus gives you a place to put those per-model, per-route metrics so cost routing decisions are based on what's actually happening, not what the code was supposed to do.

Where this runs

None of this — the fallback chain, the rate limiter, the routing layer — is specific to any one cloud or any one LLM provider. It's an ordinary stateful service: it needs to run continuously, scale with request volume, share limiter state across instances, and be observable when something upstream breaks. That's also why Kubo exists as an option worth knowing about — it's standard Kubernetes, not a proprietary platform, so a service like this one isn't locked into a single vendor's way of doing things.

That alignment between LLM infrastructure and Kubernetes isn't a coincidence specific to this one use case — it shows up across the industry. A recent survey found 66% of AI-driven companies already run on Kubernetes, largely for the same reasons: the workload needs horizontal scaling, rolling deploys when you swap a model or provider, and standard tooling for the observability this kind of routing layer depends on.

If you're weighing whether to build this routing layer on raw cloud VMs, a serverless platform with cold-start and timeout constraints, or a managed Kubernetes setup, the tradeoff usually comes down to how much of the "plumbing" — ingress, scaling rules, secrets for provider API keys, metrics collection — you want to assemble yourself versus have already wired up. When you're ready to compare, see the current plans — Kubo's entry-level Ashigaru plan starts at ¥8,800/month, and most small teams running a service like this land on the Ronin plan at ¥17,600/month.

Summary

  • A single direct call to one LLM provider doesn't survive production: outages, per-provider rate limits, and flat per-request cost all become real problems at scale.
  • A fallback chain with short timeouts and clear retryable/non-retryable error handling keeps a single provider's outage from becoming your outage.
  • A token-bucket limiter tuned below the provider's published ceiling, with shared state across instances, keeps you from tripping rate limits under load.
  • Routing requests by task complexity — not sending every request to your most capable (and most expensive) model — is the simplest lever for controlling cost, as long as you measure it per model and per route.
  • The service that does all of this is an ordinary piece of production infrastructure, which is why it runs well on standard Kubernetes rather than requiring anything LLM-specific.