Skip to main content
kRouter
All providers

Pay-per-token

fireworks/

Fireworks AI

Fast Llama, Mixtral hosting.

fireworks.ai

Quick setup

  1. Open the kRouter dashboard at http://localhost:20128

  2. Go to Providers → click Fireworks AI

  3. Paste your API key, click Save

  4. Test the connection

Use in your IDE

Point any OpenAI-compatible client at kRouter using your local API key.

ide configbash
# Endpoint
http://localhost:20128/v1

# API key (from Dashboard → API Keys)
sk-krouter-XXXX

# Model
fireworks/llama-4-405b

How Fireworks AI billing works with kRouter

Fireworks AI is pay-per-token. You bring your own API key and kRouter routes to it, so you keep your own rate limits and billing relationship. kRouter authenticates it via apikey. 4 models are routable as fireworks/<model>.

Through kRouter, Fireworks AI serves chat and code completion and embeddings, exposed on /v1/embeddings. Requests reach it through the same local OpenAI-compatible endpoint as every other provider, so switching to or from Fireworks AI is a one-line model change in your client rather than a rewrite.

Fireworks AI models & pricing

These are the model ids kRouter routes to. This provider does not publish per-token rates.

ModelRoute as
DeepSeek V3.1fireworks/accounts/fireworks/models/deepseek-v3p1
Llama 3.3 70Bfireworks/accounts/fireworks/models/llama-v3p3-70b-instruct
Qwen3 235Bfireworks/accounts/fireworks/models/qwen3-235b-a22b
Nomic Embed Text v1.5fireworks/nomic-ai/nomic-embed-text-v1.5

Fireworks AI links

Fireworks AI FAQ

Is Fireworks AI free through kRouter?

Fireworks AI is pay-per-token: you bring your own API key and are billed by Fireworks AI at their rates. kRouter adds nothing on top and is free and MIT-licensed.

How do I use Fireworks AI with Claude Code or Cursor?

Connect Fireworks AI in the kRouter dashboard, then point your tool at kRouter's local endpoint (http://localhost:20128/v1) with a kRouter API key. Any OpenAI-compatible client works, and models are addressed as fireworks/<model>.

Which Fireworks AI models can kRouter route to?

4 models, including DeepSeek V3.1, Llama 3.3 70B, Qwen3 235B. Every one is addressed as fireworks/<model> from any client.

What happens when Fireworks AI is rate limited?

kRouter detects the limit, parks that account until its real reset time, and fails over to your next connected account or provider — so the request still completes instead of erroring.