Point any coding agent at ANT.
One base-URL swap gives your agent 20,000+ models across 30+ providers, automatic cross-provider failover, Anthropic prompt caching passthrough, and a full routing trace on every request — with no per-tool setup.
Works with Claude Code, Cline, Aider, Codex CLI, Continue, opencode, Kilo, Roo, Goose, Qwen Code & anything OpenAI- or Anthropic-compatible.
- base_url: https://api.openai.com/v1
+ base_url: https://antbase.ai/v1
- api_key: sk-...
+ api_key: ant_your_key
model: ant:coder # pick a strategy, not a modelOne concept, two surfaces
Whichever protocol your agent speaks, ANT already answers it.
ANT exposes two compatible surfaces on the same host. Your agent talks to the one it already knows — you change a URL, not your code.
https://antbase.ai/v1
The drop-in path every OpenAI-SDK agent uses. Chat completions for most tools, the Responses API for Codex CLI, plus catalog and generation-metadata endpoints.
https://antbase.ai/v1/messages
The native Messages API — this is what Claude Code speaks. ANT honors cache_control so your repeated system prompt and tool defs stay cached across the agent loop.
Pick a strategy, not a model
Set your agent's model to a virtual model and let ANT choose the actual model per request. Swap strategies without touching a single config again.
ant:autoBest model for each request, chosen live
ant:coderTuned for code generation & editing
ant:reasoningThinking models for hard problems
ant:bestMax quality, cost is no object
ant:fastLowest latency for tight loops
ant:freeFree-tier models only
Quick config
Copy the config for your agent. Paste. Ship.
Every snippet is the whole integration — swap in your ant_… key and you are routing through ANT.
Claude Code
Anthropic CLI
export ANTHROPIC_BASE_URL="https://antbase.ai"
export ANTHROPIC_AUTH_TOKEN="ant_your_key"
export ANTHROPIC_MODEL="ant:coder"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="ant:auto"Base URL is host-only — Claude Code appends /v1/messages itself. cache_control is passed straight through to the provider.
Included by default
What you get for free once everything routes through ANT.
Automatic failover
A 429 or 5xx on one provider silently retries the same logical model on the next healthy endpoint — across 30+ providers. Your agent loop never stalls on a single provider having a bad minute.
Cost routing
Routine turns (rename this variable, run the tests) go to cheap fast models; genuinely hard edits escalate to frontier models. You get frontier quality where it matters without paying frontier prices on every keystroke.
Prompt caching passthrough
Anthropic cache_control is honored end-to-end. In an agent loop the system prompt and tool definitions repeat every turn — caching them cuts up to ~90% off those tokens automatically.
Structured outputs
response_format: json_schema is enforced end-to-end — passed through natively on OpenAI-compatible providers and enforced on Anthropic via a forced tool. Your agent gets schema-valid JSON back as content, streaming or not, so parsing an action plan or tool payload never fails on stray prose.
The request trace
Every request carries a full decision trace: classification, candidates, fallbacks, corrections, cost. When an agent turn misbehaves you see exactly what routed where — not a black box.
The unique angle · routing observability
Every observability tool traces what happened inside your app. ANT traces the routing decision itself.
When an agent turn goes sideways, “the model was slow” is not an answer. ANT gives you the receipt: what it classified, which models it considered and rejected, every fallback attempt with its own latency and error, and any correction it applied before your agent ever saw a problem.
self-heal · 9ms · the receipt
We didn't just retry — we fixed the request. The correction is stored so the next call skips the failure entirely.
- "max_tokens": 4096
+ "max_completion_tokens": 4096
Self-healing corrections
max_tokens → max_completion_tokens applied inline. We didn’t just retry — we fixed it, and stored the fix.
Full fallback chain
Every attempt, its provider, its latency, and its exact error — in order, on one timeline.
Per-stage latency
Classify, route, execute, heal — each stage is its own span, so you see where the time actually went.
Cost attribution
The spend is attributed to the decision that produced it, not smeared across an opaque monthly total.
Try it in 10 seconds
One curl, and you are routing.
Drop in your ant_… key and hit ANT the same way you hit any OpenAI-compatible endpoint. Prefer a UI? The Playground lets you compare models and watch the routing trace update live.
curl https://antbase.ai/v1/chat/completions \
-H "Authorization: Bearer ant_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "ant:coder",
"messages": [
{ "role": "user", "content": "Refactor this loop into a map()" }
]
}'Compatibility & gotchas
The handful of things worth knowing.
Two OpenAI wire protocols
Almost every agent uses /v1/chat/completions. Codex CLI is the exception — it needs the Responses API, and ANT serves /v1/responses for it. opencode can speak either one.
Base-URL path conventions differ
Most agents want the full https://antbase.ai/v1. Goose splits host and path (OPENAI_HOST + OPENAI_BASE_PATH). Claude Code takes the host only and appends /v1/messages itself.
Claude Code is the Anthropic-protocol one
It is the only common agent that speaks the native Messages API, so it targets /v1/messages and passes cache_control through — which is exactly what keeps agent-loop prompt caching working.
Native tool calling passes through
ANT forwards native tools to every provider, including Anthropic tool-use. No prompt-injection emulation — your function schemas reach the model verbatim.
FAQ
Questions builders ask.
Give your coding agent every model.
Grab an API key, point your agent at https://antbase.ai/v1, and watch the routing trace do the rest.