LlamaIndex makes it deceptively easy to build document Q&A systems. Load your documents, build an index, call query() โ done. But under the hood, LlamaIndex is making multiple LLM calls per query: retrieval synthesis, response refinement, sometimes tree summarization across chunks. Each of those calls hits whatever model you configured, at whatever price that model charges. For a document Q&A system that handles 1,000 queries a day, this adds up fast.
The frustrating part is that most document queries are simple lookups. "What is the return policy?" does not need GPT-4o โ any decent model can synthesize an answer from a retrieved paragraph. But the occasional complex query โ "Compare the pricing strategies across all three competitor analysis documents" โ genuinely needs a capable model for multi-document reasoning. Static model configuration forces you to optimize for one or the other.
Antbase makes this a non-issue. Every LLM call LlamaIndex makes โ retrieval synthesis, refinement, tree summarization โ gets independently routed. Simple lookups get fast free models. Complex multi-document synthesis gets premium models. Your LlamaIndex application code stays exactly the same, but your cost-per-query drops dramatically for the majority of traffic while maintaining quality on the hard queries.
LlamaIndex is a data framework for building LLM applications over structured and unstructured data. It handles indexing, retrieval, and synthesis. By using Antbase as the LLM backend, your queries get routed to the best model โ simple lookups go to fast models, complex synthesis goes to premium ones.
Installation
pip install llama-index llama-index-llms-openai-likeConfigure the LLM
from llama_index.llms.openai_like import OpenAILike
from llama_index.core import Settings
llm = OpenAILike(
api_base="https://antbase.ai/v1",
api_key="ant_your-api-key",
model="auto",
is_chat_model=True,
)
Settings.llm = llmIndex and Query Documents
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
# Load documents from a directory
documents = SimpleDirectoryReader("./data").load_data()
# Build the index
index = VectorStoreIndex.from_documents(documents)
# Query
query_engine = index.as_query_engine()
response = query_engine.query("What are the key metrics in Q4?")
print(response)Streaming Responses
query_engine = index.as_query_engine(streaming=True)
response = query_engine.query("Summarize the revenue trends.")
for token in response.response_gen:
print(token, end="", flush=True)Tips
- โขLlamaIndex makes separate calls for retrieval synthesis and refinement. Each call is independently routed by Antbase.
- โขFor embedding, configure a separate OpenAI-compatible embedding model or use a dedicated embedding provider.
- โขUse response_mode='tree_summarize' for long documents โ it makes multiple LLM calls that Antbase routes efficiently.
- โขThe is_chat_model=True flag is required for Antbase since it serves chat completions, not legacy completions.



