DeepSeek V4 Flash on Antbase — Nine Routes, One Model String, and a 31x Cache Discount
·4 min read
⚡

DeepSeek V4 Flash on Antbase — Nine Routes, One Model String, and a 31x Cache Discount

DeepSeek V4 Flash is served on Antbase by nine independent providers — DeepSeek direct, Fireworks, SiliconFlow, Vercel, Scaleway and RedPill among them. You write one model string; Antbase picks the cheapest healthy route and fails over when one degrades. Fireworks moves to $0.44 per 1M input on 21 August, with cached input at $0.014 — 31x cheaper.

AntbaseDeepSeekDeepSeek V4 FlashFireworksSiliconFlowVercel AI Gatewayroutingfailoverprompt cachingOpenAI-compatible

DeepSeek V4 Flash is the fast tier of the V4 family — built for the work where latency is the product: autocomplete, classification, extraction, routing, agent inner loops, anything you call thousands of times an hour rather than once. On Antbase it is not one endpoint. Nine independent providers serve it, and you reach all of them with a single model string through the OpenAI-compatible API you already point at.

Nine routes is the feature

Most of the value of a proxy shows up on the day a provider has a bad afternoon. DeepSeek V4 Flash is currently served on Antbase by DeepSeek direct, Fireworks, SiliconFlow (two variants), Vercel AI Gateway (two variants), Scaleway and RedPill. When one is rate-limited, out of capacity or simply slower, the router moves to the next healthy one mid-chain. You do not get paged, and you do not change code.

  • •deepseek/deepseek-v4-flash-0731 — DeepSeek direct, the reference route.
  • •fireworks/accounts/fireworks/models/deepseek-v4-flash-0731 — $0.44 per 1M input, $1.32 per 1M output from 21 August.
  • •siliconflow/deepseek-ai/DeepSeek-V4-Flash and …-0731 — two pinned variants.
  • •vercel/deepseek/deepseek-v4-flash and …-0731 — via the Vercel AI Gateway.
  • •scaleway/deepseek-v4-flash-0731 — EU-hosted.
  • •phala/deepseek/deepseek-v4-flash-0731 — RedPill, confidential compute.

The dated suffix matters. `-0731` pins the 31 July build, so a provider refreshing its default does not silently change your outputs. If you would rather always track the newest, ask for the undated name and let the catalog resolve it.

The cache discount is the number to design around

From 21 August, Fireworks prices V4 Flash at $0.44 per 1M uncached input and $1.32 per 1M output. The figure worth building around is the third one: cached input at $0.014 per 1M. That is roughly 31x cheaper than uncached, and it changes which architectures are sensible.

If your prompts share a long stable prefix — a system prompt, a tool schema, a document you ask many questions about — then the expensive part of every call after the first is already paid for. Put the stable material first and the variable part last, and you are billed the cache rate on the bulk of the tokens. Shuffle the order and you are not. It is the same content either way; only the layout differs.

A 31x gap between cached and uncached input is not a discount you claim. It is a constraint you design your prompt layout around.

Using it

python
from openai import OpenAI

client = OpenAI(base_url="https://antbase.ai/v1", api_key="ant-...")

# Pin the model — Antbase picks the cheapest healthy provider serving it.
resp = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash-0731",
    messages=[
        # Stable prefix first: this is the part that gets cached.
        {"role": "system", "content": LONG_STABLE_SYSTEM_PROMPT},
        # Variable part last.
        {"role": "user", "content": "Classify this ticket: ..."},
    ],
)

# Or pin one provider explicitly when you want a specific route:
resp = client.chat.completions.create(
    model="fireworks/accounts/fireworks/models/deepseek-v4-flash-0731",
    messages=[{"role": "user", "content": "Extract the invoice total."}],
)

Pinning a single provider is the right call when you are benchmarking, or when one route has a property you need — Scaleway for EU hosting, RedPill for confidential compute. The trade is that a pinned model has nowhere to fail over to: if that provider is briefly out of capacity, the request fails rather than quietly moving. Pin deliberately, not by default.

Build a NEST if you care about staying up

A NEST is a named list of models Antbase will route across. For V4 Flash, a nest holding three or four of the routes above gives you the model you asked for and the redundancy you want, without the router ever reaching for something you did not choose. One model, several providers, automatic failover.

The failure mode worth avoiding is a nest with one entry. It looks like pinning and behaves like a single point of failure — when that one route is busy, there is nothing to fall back to.

Where to find it

Or open it straight in the playground and compare routes side by side:

Try ANT routing today

Drop-in OpenAI-compatible API. Change your base URL and every request gets intelligent routing across 30+ providers.