Skip to main content
kRouter
All models

Llama 3.3 70B Instruct

Llama 3.3 70B Instruct is served by 2 providers that kRouter can route to, with a 128K-token context window. Because more than one provider offers it, the price you pay depends on which one you route to — the same request, the same model, different bills.

Where to get Llama 3.3 70B Instruct

ProviderRoute as
Nebius AInebius/meta-llama/Llama-3.3-70B-Instruct
Hyperbolichyperbolic/meta-llama/Llama-3.3-70B-Instruct

What Llama 3.3 70B Instruct is suited for

At 128K the window is modest, so it is better suited to focused edits than to long agent runs that accumulate history. It answers directly rather than reasoning at length, which keeps latency and output cost down on routine edits. Tool and function calling are supported, so it can drive an agent loop.

Specs

Context window
128K
Max output
64K
Input / 1M
Output / 1M

Llama 3.3 70B Instruct FAQ

How much does Llama 3.3 70B Instruct cost?

Llama 3.3 70B Instruct has no published per-token rate here — it is served through plans or free tiers rather than metered billing.

Which providers serve Llama 3.3 70B Instruct?

2 providers: Nebius AI, Hyperbolic. Through kRouter you can switch between them by changing the route prefix, without changing your tool.

What is Llama 3.3 70B Instruct's context window?

128K tokens, with up to 64K tokens of output. kRouter publishes this on /v1/models as context_length, so clients read the real number instead of guessing from the model name.

Route to Llama 3.3 70B Instruct from any tool

kRouter runs locally and gives Claude Code, Cursor, Codex and the rest one endpoint that reaches every provider above — with automatic failover. Free and MIT-licensed.

Install kRouter