Semantic Caching on Antbase — Never Pay for the Same Prompt Twice
·5 min read
♻️

Semantic Caching on Antbase — Never Pay for the Same Prompt Twice

Turn on semantic caching and Antbase remembers answers to prompts it has already seen. The next time the same — or a close-enough — prompt comes in, it returns the stored answer instantly and for free: no model is called, and your balance doesn't move. Per-workspace, private, and off by default.

Antbasesemantic cachingcost savingsembeddingslatencyOpenAI-compatible

A surprising share of production LLM traffic is repetitive: the same FAQ, the same classification prompt, the same boilerplate summary, over and over. You pay full price every single time. Semantic caching fixes that. Flip it on for a workspace and Antbase stores each answer keyed by the meaning of its prompt — so a repeat, exact or paraphrased, comes back instantly and costs nothing.

How it works

  • •On a request, Antbase first checks for an exact match of the prompt it has already answered — an instant hit.
  • •If there's no exact match, it embeds the prompt and compares it to recent cached prompts for your workspace. If one is similar enough (above your threshold), that answer is returned.
  • •On a hit, the stored response is returned with its original token counts — but no provider is called and the request is billed $0.
  • •On a miss, the request runs normally and the answer is added to the cache for next time.

"Similar enough" is a dial you control. The similarity threshold defaults to 0.92; raise it toward 1.0 to only reuse near-identical prompts, or lower it to catch looser paraphrases. Each cache entry also has a TTL, so answers naturally expire and refresh.

Private by design

A workspace's cache is its own. Entries are never shared across teams — one workspace can't ever be served another's answer. Caching is off until you turn it on, and it only applies to plain, single, text completions: requests with tool calls, image inputs, or multiple choices are never cached, because their outputs aren't safely reusable.

Turn it on

Enable semantic caching in your workspace settings, pick a threshold, and keep calling the same OpenAI-compatible endpoint — nothing about your request changes. The savings show up as $0 cache hits in your usage.

python
from openai import OpenAI

client = OpenAI(base_url="https://antbase.ai/v1", api_key="ant_...")

# First call: runs a model, billed normally, and cached.
client.chat.completions.create(
    model="ant:auto",
    messages=[{"role": "user", "content": "Summarize our refund policy in one sentence."}],
)

# Same (or a close paraphrase) again: served from cache, instantly, for $0.
client.chat.completions.create(
    model="ant:auto",
    messages=[{"role": "user", "content": "In one sentence, summarize the refund policy."}],
)

It's the cheapest correct answer there is: the one you already paid for.

Try ANT routing today

Drop-in OpenAI-compatible API. Change your base URL and every request gets intelligent routing across 30+ providers.