Gemini 3.7 Flash is now on Antbase. It is Google's fast-tier model, with gains across coding, agentic workflows, software engineering, knowledge work and web development — and it is unusually broad on input: text, image, video, audio and files all go in through the same endpoint. On Antbase it runs on a 1M-token context window with up to 65,536 output tokens, behind the OpenAI-compatible API you already point at.
What you get
- •1,048,576-token context window, up to 65,536 output tokens.
- •Five input modalities — text, image, video, audio and file — not just vision.
- •Streaming and function calling, through the same /v1 endpoint.
- •Reachable by name (google/gemini-3.7-flash) or via the pools it feeds: ant:fast, ant:vision and ant:auto.
Two routes, one model string
Antbase serves Gemini 3.7 Flash through both OpenRouter and the Vercel AI Gateway. You do not pick the backend: cost-weighted routing prefers the cheaper healthy path and fails over to the other if it degrades. Same model string, same dashboard, same usage tracking.
- •OpenRouter — $0.375 per 1M input, $1.875 per 1M output at the time of writing.
- •Vercel AI Gateway — a second independent path, used automatically when it is the better or the only healthy one.
- •OpenRouter is running Gemini 3.7 Flash at an extra 50% off through 27 August, exclusive to that route.
That promotional pricing is OpenRouter's and it has an end date, so treat the numbers above as a snapshot rather than a permanent rate. The live figure for any model is always the one in the Antbase catalog, which syncs from the provider.
from openai import OpenAI
client = OpenAI(base_url="https://antbase.ai/v1", api_key="ant-...")
# Pin Gemini 3.7 Flash by name — Antbase picks the cheaper healthy backend.
resp = client.chat.completions.create(
model="google/gemini-3.7-flash",
messages=[{"role": "user", "content": "Summarise this repo's architecture."}],
)
# Or let the router choose — Gemini 3.7 Flash now sits in these pools:
resp = client.chat.completions.create(
model="ant:fast", # also: ant:vision, ant:auto
messages=[{"role": "user", "content": "Draft the release notes."}],
)A million tokens is the interesting part
A 1M-token window at fast-tier pricing changes what is worth sending. An entire mid-sized codebase, a long meeting recording, or a stack of PDFs can go in whole rather than being chunked and re-assembled — and because audio and video are first-class inputs, transcription is not a separate step you have to bolt on in front.
The output ceiling matters too: 65,536 tokens is enough to generate a substantial file, a full migration, or a long structured document in one response instead of stitching continuations together.
Find it in the catalog, or try it side by side in the playground:
