Llama 3.3 70B Instruct
Llama 3.3 70B Instruct is served by 2 providers that kRouter can route to, with a 128K-token context window. Because more than one provider offers it, the price you pay depends on which one you route to — the same request, the same model, different bills.
Where to get Llama 3.3 70B Instruct
| Provider | Route as |
|---|---|
| Nebius AI | nebius/meta-llama/Llama-3.3-70B-Instruct |
| Hyperbolic | hyperbolic/meta-llama/Llama-3.3-70B-Instruct |
What Llama 3.3 70B Instruct is suited for
At 128K the window is modest, so it is better suited to focused edits than to long agent runs that accumulate history. It answers directly rather than reasoning at length, which keeps latency and output cost down on routine edits. Tool and function calling are supported, so it can drive an agent loop.
Specs
- Context window
- 128K
- Max output
- 64K
- Input / 1M
- —
- Output / 1M
- —
- Tool / function calling
Llama 3.3 70B Instruct FAQ
How much does Llama 3.3 70B Instruct cost?
Llama 3.3 70B Instruct has no published per-token rate here — it is served through plans or free tiers rather than metered billing.
Which providers serve Llama 3.3 70B Instruct?
2 providers: Nebius AI, Hyperbolic. Through kRouter you can switch between them by changing the route prefix, without changing your tool.
What is Llama 3.3 70B Instruct's context window?
128K tokens, with up to 64K tokens of output. kRouter publishes this on /v1/models as context_length, so clients read the real number instead of guessing from the model name.
Route to Llama 3.3 70B Instruct from any tool
kRouter runs locally and gives Claude Code, Cursor, Codex and the rest one endpoint that reaches every provider above — with automatic failover. Free and MIT-licensed.
Install kRouter