🚀 One API for every frontier model on FastInfra

Browse models

Transparent per-token pricing → compare providers on FastInfra

View pricing

OpenAI-compatible chat completions → start in minutes

Read the docs

⚡ Server Auction → Enterprise bare metal from $78.15/mo (173 in stock)

Browse deals

Llama Nemotron Embed Vl 1b v2 vs NVIDIA Nemotron Nano 9B v2

Live API pricing and availability, side by side. Both models run on FastInfra's OpenAI-compatible endpoint — switching between them is a one-line change.

Pricing

Price per 1M tokens

Prices are live and sync automatically from wholesale providers.

Llama Nemotron Embed Vl 1b v2 NVIDIA Nemotron Nano 9B v2
Input $0.01 $0.06
Output — $0.26
Default route route-21 route-03
Routes available 1 1

On input tokens, Llama Nemotron Embed Vl 1b v2 is currently 6× cheaper than NVIDIA Nemotron Nano 9B v2 on FastInfra. Output-token pricing and quality trade-offs differ per workload — test both with the free tier.

Quickstart

Try both in 30 seconds

Same endpoint, same SDK — only the model string changes.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

for model in ["nvidia/llama-nemotron-embed-vl-1b-v2", "nvidia/NVIDIA-Nemotron-Nano-9B-v2"]:
    response = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": "Hello!"}]
    )
    print(model, "->", response.choices[0].message.content)