Skip to content
EngineeringJuly 12, 2026

Building an LLM gateway on Cloudflare Workers

AnyRouter runs as a set of Cloudflare Workers, not a long-running server. That constrains the design in specific ways — a 128MB memory ceiling that rules out buffering responses, a 3MiB script cap that forced a multi-Worker split, and a hard rule that every upstream call goes through Cloudflare's AI Gateway. Here's what that architecture actually looks like.

Why Workers instead of a server

A Cloudflare Worker is an isolate, not a process — it starts on the request, handles it, and doesn't sit around holding memory or connections between calls. There's no server to keep warm, no fleet to scale, no region to pick; the same script runs at every Cloudflare edge location. For a gateway whose entire job is 'take a request, pick an upstream, stream the response back,' that request-scoped model is a good fit — but it also means every architectural decision has to respect isolate limits that a conventional server wouldn't have.

Two of those limits shape almost everything downstream: a 128MB memory ceiling per isolate, and a 3MiB compressed script-size cap per Worker.

Stream, never buffer

The 128MB ceiling makes one common server pattern actively dangerous: reading an entire upstream response into memory before forwarding it. An LLM completion can be arbitrarily long, and await response.text() on an unbounded stream has no ceiling of its own — it grows until the isolate runs out of memory.

Anti-patternWhy it breaksCorrect pattern
await response.text() on unbounded upstream dataMemory exhaustion against the 128MB isolate limitStream: return new Response(readableStream, { headers })
Module-level mutable state for request dataCross-request leakage between isolates handling different usersPer-request state on c.set(...) or closures
Floating promises for background workDropped results, swallowed errors after the response returnsc.executionCtx.waitUntil(promise)

In practice this means chat completions, Anthropic messages, and the Responses API all return a native Response wrapping the upstream's server-sent-event stream, piped through a format translator chunk by chunk rather than accumulated and replayed:

export default (async (c) => {
  const upstream = await fetch(upstreamUrl, { method: "POST", body, headers })

  // Never buffer: pipe the upstream stream straight through,
  // translating SSE dialect chunk-by-chunk as it passes.
  return new Response(upstream.body.pipeThrough(dialectTranslator), {
    status: upstream.status,
    headers: { "Content-Type": "text/event-stream", "Cache-Control": "no-cache" },
  })
}) satisfies HonoHandler
Streaming pattern used across the executor's chat/messages/responses paths.

The 3MiB cap forced a multi-Worker split

A single Worker outgrew Cloudflare's 3MiB compressed script-size cap once the routing engine, the model catalog, Durable Objects, MCP server, OAuth, and the full Hono API surface all lived in one script. The fix was a hub-and-spoke split: one web Worker owns the zone-wide route and either forwards a request to a dedicated spoke over a service binding, or serves the page itself.

WorkerHostOwns
anyrouter (web hub)anyrouter.dev (apex)SSR landing pages; routes and forwards everything else
anyrouter-apiservice binding onlyHono app, executor, catalog, Durable Objects
anyrouter-dashboarddash.anyrouter.devDashboard app
anyrouter-adminadmin.anyrouter.devAdmin app
anyrouter-mcpservice binding onlyMCP server (JSON-RPC + OAuth)
anyrouter-flowcron + service bindingWorkflows and scheduled jobs

Because each spoke proxies its own same-origin /api/* calls to anyrouter-api over a service binding, the browser never makes a cross-origin request — no CORS configuration needed, and no extra public surface exposed per Worker.

Every upstream goes through the AI Gateway

Dispatch to the actual model provider happens through exactly one of three transport shapes, all funneled through Cloudflare's AI Gateway rather than calling a provider's API directly:

TransportReachesHow
Workers AI bindingFirst-party and partner-hosted Workers AI catalog modelsc.env.AI.run(model, body, { gateway: { id }, returnRawResponse: true })
Unified AI Gateway RESTCF unified-billing catalog, third-party modelsOne API token, cf-aig-gateway-id header, usage bills via Unified Billing
Per-provider AI Gateway RESTBYOK and direct providers (OpenAI, Anthropic, xAI, Groq, DeepInfra, ...)cf-aig-authorization header plus the provider's own key in Authorization

That split is invisible to the end user — every backend shares the same 'Cloudflare AI Gateway' badge in the UI regardless of which transport actually served it. The rule behind it is stricter than it sounds: every upstream must route through gateway.ai.cloudflare.com, with exactly one documented exception for a first-party backend that has no HTTP origin to proxy in the first place (it reaches a user's own device over an outbound WebSocket instead).

Security primitives the edge gives you for free

Running on Workers also means the platform's crypto primitives are the default, not an add-on dependency:

  • crypto.getRandomValues() or crypto.randomUUID() — never Math.random() for tokens or ids
  • crypto.subtle for all cryptographic operations, including per-row encryption at rest
  • crypto.subtle.timingSafeEqual for secret comparison, so key checks aren't timing-attackable
  • Secrets live in wrangler secret put, never hardcoded in source or committed config

Explicit error handling matters more than it would in a framework with a global exception page, too — passThroughOnException() is avoided deliberately, because it hides real bugs behind a silent fallback instead of surfacing them to Hono's error handler where they can be logged and fixed.

What the edge trades away

None of this is free. Every limit above is a constraint the architecture works around, not one it eliminates — the 3MiB cap means new features have to think about which Worker they belong in, and the streaming-only rule means there's no shortcut for post-processing a full response before it reaches the client. What it buys back is a script that runs identically at every edge location with no server fleet behind it to operate.

AnyRouter live network stats dashboard showing aggregate request and token volume
The same Worker split serves this traffic — no dedicated servers behind it.

See the full request lifecycle, model catalog, and live routing in your own account.

Open the dashboard

Route your first request in 2 minutes

Start free with your own keys, or top up and pay per token. Get $4/mo in credits and free models on Go — $2/mo, or free when you donate a provider key.

Start free