Run Claude Code, Cline & Aider on Any of 20,000+ Models — With One Base URL
·7 min read
🤖

Run Claude Code, Cline & Aider on Any of 20,000+ Models — With One Base URL

Coding agents now dominate token spend, and each one is hardwired to a single provider — no failover, no cost routing, real lock-in. Point your agent's base URL at Antbase and it inherits automatic cross-provider failover, prompt caching, cost routing, and per-request routing traces across 20,000+ models. Here's how to wire up Claude Code, Cline, and Aider in a few lines.

Coding AgentsDevelopersGuide

Coding agents have quietly become the biggest thing running through AI proxies — well over 70% of the token volume we see now comes from agents, not chat. And nearly every one of them ships hardwired to a single provider: Claude Code talks to Anthropic, most Cline and Aider setups talk to OpenAI. That's fine until the provider rate-limits you mid-task, has an outage, or charges list price for the millions of tokens an agent loop burns through. You get no failover, no cost routing, and a codebase quietly locked to one vendor's roadmap.

One base URL fixes all of it. Point your agent at Antbase and it keeps working exactly as before — same wire format, same tool calls, same streaming — but now every request routes across 20,000+ models on 30+ providers, with automatic failover, prompt caching, cost routing, and a full trace of every decision.

The one idea

Every serious coding agent speaks one of two wire formats: OpenAI Chat Completions or Anthropic Messages. Antbase speaks both. So you don't need a plugin or a fork — you change the base URL your agent already points at, and it inherits failover, caching, cost routing, and tracing for free.

  • •OpenAI-compatible endpoint: https://antbase.ai/v1 — for Cline, Aider, Continue, Codex CLI, and anything that speaks the OpenAI API.
  • •Anthropic-compatible endpoint: https://antbase.ai/v1/messages — set the host to https://antbase.ai for Claude Code and any Anthropic-native tool.
  • •One key format for both: ant_... — issued from your dashboard, used as the bearer or auth token everywhere.

Pick a strategy, not a model

Instead of pinning your agent to one model and babysitting it through deprecations and price changes, point it at a virtual model that names an intent. Antbase routes each request to the best underlying model for that intent, and falls back automatically if a provider degrades.

  • •ant:coder — tuned for coding agents: strong code models, cost-aware, with escalation on hard turns.
  • •ant:auto — general best-fit routing; classifies each request and picks the strongest healthy model for it.
  • •ant:reasoning — routes to reasoning-grade models for planning, refactors, and multi-step problems.
  • •ant:best, ant:fast, ant:free — optimize for raw capability, latency, or zero cost respectively.

You can still pin an exact model any time — just use provider/model form, e.g. anthropic/claude-sonnet-5 or openai/gpt-5. Same API, same key; you're choosing a specific model instead of a strategy.

Set up the big three

Claude Code

Claude Code speaks the Anthropic wire format, so point its host at Antbase and hand it an ant_ token. Set three environment variables:

bash
export ANTHROPIC_BASE_URL="https://antbase.ai"
export ANTHROPIC_AUTH_TOKEN="ant_..."
export ANTHROPIC_MODEL="ant:coder"

claude

That's it — Claude Code now runs through Antbase's Anthropic-compatible endpoint at https://antbase.ai/v1/messages. Tool use, streaming, and prompt caching all pass straight through.

Cline

In Cline's settings, choose the OpenAI Compatible provider and fill in three fields:

  • •Base URL: https://antbase.ai/v1
  • •API Key: ant_...
  • •Model ID: ant:coder

Cline's plan/act loop, tool calls, and diffs work unchanged — you've just swapped the endpoint underneath it.

Aider

Aider uses the OpenAI-compatible path too. Set the base URL and key in the environment, then name the model with the openai/ prefix so Aider routes it through that provider:

bash
export OPENAI_API_BASE="https://antbase.ai/v1"
export OPENAI_API_KEY="ant_..."

aider --model openai/ant:coder

Continue, Codex CLI (which needs the Responses API — Antbase serves it), opencode, Kilo, Roo, Goose, and Qwen Code all follow the same one-base-URL pattern. Copy-paste configs for each live on the onboarding page:

What you get for free

Changing one base URL is the whole cost. In exchange, every request your agent makes picks up the things a single-provider setup can't give you:

  • •Failover across 30+ providers: a 429, timeout, or outage on one backend is retried on the next healthy provider that serves the model — mid-task, without your agent noticing.
  • •Cost routing: routine turns go to cheap, capable models; genuinely hard turns escalate to frontier models. You stop paying flagship prices for boilerplate edits.
  • •Prompt caching: agent loops resend the same big system prompt and tool definitions on every turn — Antbase passes Anthropic prompt caching straight through, taking up to ~90% off those repeated tokens.
  • •Native tool/function calling: tool and function calls pass through natively to every provider, including Anthropic — your agent's tool loop just works.
  • •Structured outputs: response_format json_object and json_schema are enforced end-to-end, so the JSON your agent parses back is always schema-valid — even on models without a native structured-output mode.
  • •Semantic caching: repeated or near-identical requests can be served from cache for $0.

Structured outputs that actually validate

Agents don't just chat — they emit plans, tool payloads, and typed data your code has to parse. Antbase supports OpenAI's response_format for both json_object and full json_schema (strict) mode, and it works the same across every provider. On OpenAI-compatible backends the schema is passed through natively; on Anthropic — which has no native structured-output mode — Antbase enforces it with a forced tool and unwraps the result straight back into message content. Either way your agent gets schema-valid JSON, streaming or not, with no stray prose to strip.

python
from openai import OpenAI

client = OpenAI(base_url="https://antbase.ai/v1", api_key="ant_...")

resp = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",  # or ant:coder — works on any provider
    messages=[{"role": "user", "content": "Plan: add a double-jump to the player."}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "plan",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {
                    "steps": {"type": "array", "items": {"type": "string"}},
                    "files": {"type": "array", "items": {"type": "string"}},
                },
                "required": ["steps", "files"],
                "additionalProperties": False,
            },
        },
    },
)

# resp.choices[0].message.content is guaranteed valid JSON for the schema
import json
plan = json.loads(resp.choices[0].message.content)

The unique part: routing observability

App-level tracing tools show you what your agent did — which prompt it sent, what came back. Antbase traces the routing decision itself. For every request you can see how it was classified, which candidate models were considered, the full fallback chain with per-attempt latency and cost, and any self-healing error corrections applied along the way.

That means when a turn is slow or expensive, you don't guess. You see not just what you spent, but which routing decision spent it — the model that answered, the ones it tried first, and why it moved on.

The cost angle

Agent loops are token-hungry in a very specific way: a large, mostly-static system prompt plus tool definitions on every turn, lots of routine edits, and a small fraction of genuinely hard reasoning. Route that shape well and the savings are large without downgrading the agent — cheap models handle the routine turns, prompt caching absorbs the repeated context, and escalation reserves frontier models for the hard 10%. In practice that's a 70–90% cut in agent token cost while the agent keeps reaching for a frontier model exactly when it needs one.

Get started

Grab a key, point your agent's base URL at Antbase, and pick ant:coder. Everything else — failover, caching, cost routing, tracing — is already on.

Or try the routing strategies your agent will use, side by side, in the playground:

Try ANT routing today

Drop-in OpenAI-compatible API. Change your base URL and every request gets intelligent routing across 30+ providers.