Llama 3.1 Nemotron 8B UltraLong 2M Instruct is a efficient model from Featherless with a 33K context window, supporting function calling. Available through ANT's unified API.
Limited routing data available for Llama 3.1 Nemotron 8B UltraLong 2M Instruct. Intelligence insights will appear as more requests are routed through ANT.
from openai import OpenAI
client = OpenAI(
base_url="https://api.antbase.ai/v1",
api_key="YOUR_ANT_API_KEY",
)
response = client.chat.completions.create(
model="featherless/nvidia/Llama-3.1-Nemotron-8B-UltraLong-2M-Instruct",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)Ask an AI about it