Quick setup
Open the kRouter dashboard at http://localhost:20128
Go to Providers → click Ollama Cloud
Paste your API key, click Save
Test the connection
Use in your IDE
Point any OpenAI-compatible client at kRouter using your local API key.
# Endpoint
http://localhost:20128/v1
# API key (from Dashboard → API Keys)
sk-krouter-XXXX
# Model
ollama/llama-4How Ollama Cloud billing works with kRouter
Ollama Cloud gives you free starting credits, then bills per token. kRouter tracks the remaining quota and fails over to another provider before you hit a wall. kRouter authenticates it via apikey. 6 models are routable as ollama/<model>.
Through kRouter, Ollama Cloud serves chat and code completion. Requests reach it through the same local OpenAI-compatible endpoint as every other provider, so switching to or from Ollama Cloud is a one-line model change in your client rather than a rewrite.
Free tier: light usage, 1 cloud model at a time (limits reset every 5h & 7d). Pro $20/mo · Max $100/mo.
Ollama Cloud models & pricing
Prices are USD per 1M tokens, cheapest first. kRouter bills nothing on top; this is the provider's own rate.
| Model | Route as | Input /1M | Output /1M |
|---|---|---|---|
| Qwen3.5 | ollama/qwen3.5 | $0.50 | $2.00 |
| MiniMax M2.5 | ollama/minimax-m2.5 | $0.60 | $2.40 |
| GLM 4.7 Flash | ollama/glm-4.7-flash | $0.75 | $3.00 |
| GLM 5 | ollama/glm-5 | $1.00 | $4.00 |
| Kimi K2.5 | ollama/kimi-k2.5 | $1.20 | $4.80 |
| GPT OSS 120B | ollama/gpt-oss:120b | — | — |
Ollama Cloud links
Ollama Cloud FAQ
Is Ollama Cloud free through kRouter?
Ollama Cloud is pay-per-token: you bring your own API key and are billed by Ollama Cloud at their rates. kRouter adds nothing on top and is free and MIT-licensed.
How do I use Ollama Cloud with Claude Code or Cursor?
Connect Ollama Cloud in the kRouter dashboard, then point your tool at kRouter's local endpoint (http://localhost:20128/v1) with a kRouter API key. Any OpenAI-compatible client works, and models are addressed as ollama/<model>.
Which Ollama Cloud models can kRouter route to?
6 models, including GPT OSS 120B, Kimi K2.5, GLM 5. Every one is addressed as ollama/<model> from any client.
What happens when Ollama Cloud is rate limited?
kRouter detects the limit, parks that account until its real reset time, and fails over to your next connected account or provider — so the request still completes instead of erroring.