The catalog is wide, not deep

AnyRouter's catalog covers 170+ models across 28+ providers, addressed by a single provider/model string — anthropic/claude-opus-4-8, z-ai/glm-5.2, google/gemini-2.5-pro. Models are organized by the organization that owns them, not by which upstream serves them, and that ownership is lopsided: a handful of labs account for most of the catalog, and a long tail of smaller labs each contribute a model or two.
| Owner | Models in catalog |
|---|---|
| openai | 32 |
| qwen | 16 |
| meta | 16 |
| 14 | |
| anthropic | 12 |
| x-ai | 11 |
| venice | 8 |
| wafer | 7 |
| deepseek | 7 |
| nvidia | 6 |
That split matters because a model can also come in two SKU classes on the Cloudflare side alone — hosted first-party (@cf/owner/model) or partner-hosted (owner/model, no @cf prefix) — before you even get to a third-party upstream like DeepInfra or a direct provider API. Breadth in the catalog doesn't mean breadth in what gets called.
What people actually call for free
Calling any model is one request against a single, OpenAI-compatible endpoint:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer sk-ar-v1-..." \
-H "Content-Type: application/json" \
-d '{"model": "z-ai/glm-5.2", "messages": [{"role": "user", "content": "Hello"}]}'But traffic through the free routes — no credit charged, not a BYOK call — doesn't spread evenly across the catalog. Two models account for the overwhelming majority of free requests, all-time through July 11, 2026:
Free-served requests by model
All-time through July 11, 2026
Two shapes of usage, not one
Request count alone hides what's actually happening. GLM-5.2 leads on requests and dwarfs everything else on tokens served — north of 200M free tokens — which is the signature of a model running real, sustained work, not one-off tests. gpt-4o-mini leads on requests too, but with a fraction of the tokens per call: it's the classifier-and-glue model, cheap enough to sit inside a loop where each call is a sentence or two.
Those are two different jobs wearing the same 'free model' label, and the request count on its own doesn't tell you which one a given model is doing — the tokens-per-request ratio does.
Coding agents show up as attribution, not as a model name
A meaningful share of production traffic doesn't come from a developer typing a prompt — it comes from a coding agent making the call on their behalf. Claude Code, Codex, and Cursor all speak different dialects against the same catalog:
| Dialect | Endpoint | Used by |
|---|---|---|
| OpenAI chat | /api/v1/chat/completions | Codex, Cursor, any OpenAI-compatible client |
| Anthropic messages | /api/v1/messages | Claude Code and other Claude-native clients |
| OpenAI Responses | /api/v1/responses | Newer OpenAI-compatible agent tooling |
Because every dialect terminates at the same catalog, a request from Claude Code and a request typed by hand can resolve to the same model — the only difference visible in usage analytics is the app attribution header the client sent, not a separate code path.
The free tier is funded, not subsidized from one balance sheet
The free capacity behind those request counts isn't purchased in bulk — it's a shared key pool. Members donate spare quota from a provider's own free tier, quotas sum into one pool, and donors earn a share back plus a free plan tier. The more capacity gets donated, the more the free-tier numbers above can grow without a paywall behind them.

Reading the shape correctly
- A wide catalog and a concentrated traffic pattern aren't a contradiction — reach and popularity are different things
- Request count and token count together identify what job a model is actually doing, not just how often it's called
- Model ownership in the catalog skews toward a handful of labs, but the smaller entries still get called — just less often
- Free-tier volume scales with donated capacity, not with a fixed budget
See live catalog, traffic, and per-model usage in your own account.
Open the dashboardRoute your first request in 2 minutes
Start free with your own keys, or top up and pay per token. Get $4/mo in credits and free models on Go — $2/mo, or free when you donate a provider key.
Start free