Skip to content
ProductJuly 11, 2026

Apple's on-device Foundation Model, now on the AnyRouter API

Run Apple Intelligence's on-device Foundation Model on your own Mac and call it through the same AnyRouter API you use for everything else. Your prompts stay on your hardware, you pay nothing for tokens, and there's no vendor account to manage.

A model that runs on your Mac

Apple ships an on-device Foundation Model as part of Apple Intelligence — a small language model that runs entirely on the local hardware, with no data leaving the device. It's now in the AnyRouter catalog as apple/foundation-model, reachable through the exact same OpenAI-compatible API you already use for every other model.

The shape is different from every other model we host, and it's worth being precise about it: AnyRouter does not run this model for you. Your own Mac does. AnyRouter is the connection between the public API and the model server running on your machine — nothing more. That's what makes it $0 and what keeps your prompts on your own hardware.

How the relay works

Apple's model has no public endpoint — it runs on a machine with no server address for us to call. So instead of AnyRouter reaching out to a provider, your Mac reaches out to AnyRouter. The anyr CLI opens an outbound WebSocket to AnyRouter and holds it open. When a request comes in for apple/foundation-model, AnyRouter pushes it down that connection to your Mac, your local model server answers, and the response streams back up the same connection to whoever made the call.

sequenceDiagram
participant C as API caller
participant GW as AnyRouter
participant M as Your Mac (anyr relay)
participant FM as On-device model
M->>GW: open outbound WebSocket, stay connected
C->>GW: POST /v1/chat/completions (apple/foundation-model)
GW->>M: push request down the socket
M->>FM: run on-device
FM-->>M: tokens
M-->>GW: stream response back up
GW-->>C: streamed response
Your Mac dials out and holds the connection; AnyRouter pushes each request down it and streams the answer back.

Because the origin is a device you own and control, this is the one place in AnyRouter where a request doesn't route through the usual gateway path — there's simply nothing on the public internet to route to. The model server on your Mac is the origin.

Set it up in two steps

You need a Mac that supports Apple Intelligence and a local server exposing the on-device model over an OpenAI-compatible endpoint (Apple's fm serve, listening on 127.0.0.1:1976). Once that's running, one command connects it to your AnyRouter account:

# 1. Start Apple's on-device model server on your Mac
fm serve

# 2. Connect it to AnyRouter (auto-detects fm serve on port 1976)
anyr relay start

relay start pairs the machine on first run — it uses an existing AnyRouter API key if it finds one, or walks you through a browser login otherwise — then holds the connection open and reconnects on its own if your network drops. Leave it running and your Mac is on call to serve requests for apple/foundation-model.

Call it like any other model

With the relay connected, apple/foundation-model behaves like every other id in the catalog. Point a standard chat-completions request at it:

curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer $ANYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "model": "apple/foundation-model",
  "messages": [{ "role": "user", "content": "Summarize this in one sentence." }]
}'

Or pick apple/foundation-model from the model dropdown in the Playground at playground.anyrouter.dev and chat with it in the browser. Streaming works the same way it does for any other model. If no paired device is online when the request lands, the call fails fast so a fallback model can answer instead — nothing hangs waiting on an offline Mac.

What to expect

Two honest caveats. First, the on-device model is small and built for on-device work — Apple caps its context window at 4,096 tokens, shared across your system instructions, any tool schemas, and the conversation. It's a genuine framework limit, not a setting we can raise. Treat it as a fast local model for focused tasks, not a long-context workhorse.

Second, throughput depends entirely on your Mac — the model runs on your silicon, so your hardware sets the pace. We're not going to quote latency numbers here, because the only ones that matter are the ones you measure on your own machine.

  • Private by design. Prompts and completions run on your device and are never sent to a third-party model vendor.
  • $0 per token. You supply the hardware, so there's nothing to bill — input and output are both priced at zero.
  • No vendor account. There's no API key to obtain from Apple; the model is part of Apple Intelligence on your Mac.

Share spare capacity, optionally

If your Mac has idle time, you can let it serve requests for other AnyRouter users the same way donated provider keys back the shared key pool. Add --pool when you connect:

anyr relay start --pool

Pool routing only ever falls back to a donor device when the requester's own machine is offline, and it's opt-in per device — leave the flag off and your Mac only ever answers your own requests. For the mechanics of the shared pool and how donors earn credits, see /blog/shared-key-pool.

Route your first request in 2 minutes

Start free with your own keys, or top up and pay per token. Get $4/mo in credits and free models on Go — $2/mo, or free when you donate a provider key.

Start free