Document Q&A with LlamaIndex and Antbase
ยท5 min read
๐Ÿฆ™

Document Q&A with LlamaIndex and Antbase

Build document question-answering systems with LlamaIndex and Antbase for intelligent model routing on every query.

llamaindexragdocuments

LlamaIndex makes it deceptively easy to build document Q&A systems. Load your documents, build an index, call query() โ€” done. But under the hood, LlamaIndex is making multiple LLM calls per query: retrieval synthesis, response refinement, sometimes tree summarization across chunks. Each of those calls hits whatever model you configured, at whatever price that model charges. For a document Q&A system that handles 1,000 queries a day, this adds up fast.

The frustrating part is that most document queries are simple lookups. "What is the return policy?" does not need GPT-4o โ€” any decent model can synthesize an answer from a retrieved paragraph. But the occasional complex query โ€” "Compare the pricing strategies across all three competitor analysis documents" โ€” genuinely needs a capable model for multi-document reasoning. Static model configuration forces you to optimize for one or the other.

Antbase makes this a non-issue. Every LLM call LlamaIndex makes โ€” retrieval synthesis, refinement, tree summarization โ€” gets independently routed. Simple lookups get fast free models. Complex multi-document synthesis gets premium models. Your LlamaIndex application code stays exactly the same, but your cost-per-query drops dramatically for the majority of traffic while maintaining quality on the hard queries.

LlamaIndex is a data framework for building LLM applications over structured and unstructured data. It handles indexing, retrieval, and synthesis. By using Antbase as the LLM backend, your queries get routed to the best model โ€” simple lookups go to fast models, complex synthesis goes to premium ones.

Installation

bash
pip install llama-index llama-index-llms-openai-like

Configure the LLM

python
from llama_index.llms.openai_like import OpenAILike
from llama_index.core import Settings

llm = OpenAILike(
    api_base="https://antbase.ai/v1",
    api_key="ant_your-api-key",
    model="auto",
    is_chat_model=True,
)

Settings.llm = llm

Index and Query Documents

python
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader

# Load documents from a directory
documents = SimpleDirectoryReader("./data").load_data()

# Build the index
index = VectorStoreIndex.from_documents(documents)

# Query
query_engine = index.as_query_engine()
response = query_engine.query("What are the key metrics in Q4?")
print(response)

Streaming Responses

python
query_engine = index.as_query_engine(streaming=True)
response = query_engine.query("Summarize the revenue trends.")

for token in response.response_gen:
    print(token, end="", flush=True)

Tips

  • โ€ขLlamaIndex makes separate calls for retrieval synthesis and refinement. Each call is independently routed by Antbase.
  • โ€ขFor embedding, configure a separate OpenAI-compatible embedding model or use a dedicated embedding provider.
  • โ€ขUse response_mode='tree_summarize' for long documents โ€” it makes multiple LLM calls that Antbase routes efficiently.
  • โ€ขThe is_chat_model=True flag is required for Antbase since it serves chat completions, not legacy completions.

Try ANT routing today

Drop-in OpenAI-compatible API. Change your base URL and every request gets intelligent routing across 30+ providers.