A surprising share of production LLM traffic is repetitive: the same FAQ, the same classification prompt, the same boilerplate summary, over and over. You pay full price every single time. Semantic caching fixes that. Flip it on for a workspace and Antbase stores each answer keyed by the meaning of its prompt — so a repeat, exact or paraphrased, comes back instantly and costs nothing.
How it works
- •On a request, Antbase first checks for an exact match of the prompt it has already answered — an instant hit.
- •If there's no exact match, it embeds the prompt and compares it to recent cached prompts for your workspace. If one is similar enough (above your threshold), that answer is returned.
- •On a hit, the stored response is returned with its original token counts — but no provider is called and the request is billed $0.
- •On a miss, the request runs normally and the answer is added to the cache for next time.
"Similar enough" is a dial you control. The similarity threshold defaults to 0.92; raise it toward 1.0 to only reuse near-identical prompts, or lower it to catch looser paraphrases. Each cache entry also has a TTL, so answers naturally expire and refresh.
Private by design
A workspace's cache is its own. Entries are never shared across teams — one workspace can't ever be served another's answer. Caching is off until you turn it on, and it only applies to plain, single, text completions: requests with tool calls, image inputs, or multiple choices are never cached, because their outputs aren't safely reusable.
Turn it on
Enable semantic caching in your workspace settings, pick a threshold, and keep calling the same OpenAI-compatible endpoint — nothing about your request changes. The savings show up as $0 cache hits in your usage.
from openai import OpenAI
client = OpenAI(base_url="https://antbase.ai/v1", api_key="ant_...")
# First call: runs a model, billed normally, and cached.
client.chat.completions.create(
model="ant:auto",
messages=[{"role": "user", "content": "Summarize our refund policy in one sentence."}],
)
# Same (or a close paraphrase) again: served from cache, instantly, for $0.
client.chat.completions.create(
model="ant:auto",
messages=[{"role": "user", "content": "In one sentence, summarize the refund policy."}],
)It's the cheapest correct answer there is: the one you already paid for.


