🚀 One API for every frontier model on FastInfra

Browse models

Transparent per-token pricing → compare providers on FastInfra

View pricing

OpenAI-compatible chat completions → start in minutes

Read the docs

⚡ Server Auction → Enterprise bare metal from $78.15/mo (173 in stock)

Browse deals

Llama 4 Maverick 17B 128E Instruct FP8 API

Run meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 through FastInfra's API. Pay per token, no subscription, routed to the cheapest available provider.

Pricing

Llama 4 Maverick 17B 128E Instruct FP8 pricing

Billed per token used. Prices sync automatically from wholesale providers.

Direction Price per 1M tokens
Input$0.21
Output$0.84
Availability

1 route available

Requests route to route-21 by default (lowest cost). FastInfra fails over automatically if a route is unavailable.

Route label Model ID
route-21 meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
Quickstart

Call Llama 4 Maverick 17B 128E Instruct FP8 in 30 seconds

Works with any OpenAI SDK — change the base URL and API key only.

Python

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

response = client.chat.completions.create(
    model="meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)

curl

curl https://api.fastinfra.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
FAQ

Llama 4 Maverick 17B 128E Instruct FP8 — common questions

How much does the meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 API cost?

On FastInfra, meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 costs $0.21/1M input, $0.84/1M output tokens. Billing is per token used, with no subscription.

Is meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 compatible with the OpenAI SDK?

Yes. FastInfra exposes meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 through an OpenAI-compatible endpoint at https://api.fastinfra.ai/v1 — point any OpenAI SDK at that base URL with a FastInfra API key and keep your existing code.

How does routing work for meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8?

meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 is available on 1 route(s): route-21. Requests use route-21 by default (lowest cost). FastInfra fails over automatically if a route is unavailable.