What you're building
Hermes 3 405B is Nous Research's open-weight generalist — strong at reasoning, multi-turn chat, and tool calling, which makes it a great agent brain. AnyRouter serves it as nousresearch/hermes-3-llama-3.1-405b through one OpenAI-compatible endpoint, so any client or agent framework that speaks the OpenAI API can drive it.
The model runs on your own free Hugging Face token (BYOK), so inference is free to you and AnyRouter adds zero markup. You still get everything on top: one API, logs, analytics, and automatic failover. Total setup time: about three minutes.
1. Grab a free Hugging Face token
Sign in at huggingface.co and open Settings → Access Tokens. Create a token with the Inference Providers permission (the free tier is enough to start). Copy it — it starts with hf_.
2. Add it as a BYOK key
In the AnyRouter dashboard, open BYOK & Key Balancing at /byok, choose Hugging Face as the provider, paste your hf_ token, and save. Keys are encrypted at rest and write-only in the UI — after saving you only ever see a masked identifier.
Attach your free Hugging Face token in a few clicks.
Connect a key3. Mint an AnyRouter API key
Head to /keys and create an AnyRouter key (it starts with sk-ar-). This is the single key your agent authenticates with — AnyRouter maps the request onto your Hugging Face token behind the scenes, so your agent never touches the provider key directly.
4. Point your agent at AnyRouter
Use https://anyrouter.dev/api/v1 as the base URL and nousresearch/hermes-3-llama-3.1-405b as the model. A one-line curl to confirm it works:
curl https://anyrouter.dev/api/v1/chat/completions \
-H "Authorization: Bearer $ANYROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nousresearch/hermes-3-llama-3.1-405b",
"messages": [{ "role": "user", "content": "You are a helpful agent. Say hi." }]
}'Because it's OpenAI-compatible, the official OpenAI SDK works unchanged — just swap the base URL and key:
from openai import OpenAI
client = OpenAI(
base_url="https://anyrouter.dev/api/v1",
api_key="sk-ar-...", # your AnyRouter key
)
resp = client.chat.completions.create(
model="nousresearch/hermes-3-llama-3.1-405b",
messages=[{"role": "user", "content": "You are a helpful agent. Say hi."}],
)
print(resp.choices[0].message.content)Point Claude Code, Codex, Cursor, or your own loop at the same base URL and model id and you've got a Hermes agent running.
One more thing: earn while it sits idle
Your Hugging Face quota is mostly idle. Flip the donate toggle on the key and every request the shared pool routes through it adds up to 8% of that request's notional cost to a pending balance you can claim as credits (5% on Go, 8% on Pro/Max) — and donating one working key unlocks the Go plan for free, for as long as it stays active. You keep first call on your own key; the pool only ever borrows spare capacity, and you can opt out instantly.
Run Hermes as an agent through one API at $0 markup.
Get startedRelated posts
Point Claude Code at a custom base URL
Claude Code reads ANTHROPIC_BASE_URL and ANTHROPIC_API_KEY, so pointing it at AnyRouter takes two environment variables — no plugin, no fork. You get every model in the catalog behind the same key, automatic failover when a provider is rate-limited, and one audit log shared with Codex, Cursor, and every other agent you run.
MCP servers for AI agents: what they are and how to connect one
An MCP server lets an AI agent call a workspace's tools — keys, credits, models, status — over one standard interface instead of custom glue code per service. Here's what MCP is and how to connect AnyRouter's managed MCP server to Claude Desktop, Claude Code, Cursor, or your own agent.
Route your first request in 2 minutes
Start free with your own keys, or top up and pay per token. Get $4/mo in credits and free models on Go — $2/mo, or free when you donate a provider key.
Start free