Local first, cloud when it can't: Ollama with cloud fallback
Use a local Ollama model and switch to a cloud model only when it fails. What triggers the switch, what silently does not, and how to set it up.
Your coding agent points at a model on your own machine, and it works until the laptop restarts and Ollama doesn't come back, or the desktop running it goes to sleep. The agent stops with a connection error, and checking by hand shows why:
$ curl http://localhost:11434/api/tags
curl: (7) Failed to connect to localhost port 11434 after 0 ms: Couldn't connect to serverThe setup people want: use the local model, and quietly use a cloud one when the local one can't answer. But "can't answer" covers several situations, and only some of them produce an error that anything in front of Ollama can act on. Here is which is which, how to build the chain with kRouter, and when you don't need a router at all.
When you do not need a router
Since v0.14.0, announced in January 2026, Ollama serves Anthropic's Messages API on its usual port (Ollama's announcement), next to its own API and an OpenAI-compatible one. For Claude Code, Ollama's guide is one command, ollama launch claude, or three variables:
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434Ollama also runs cloud models through the same local server: sign in and use a model name ending in :cloud (Ollama's cloud docs). But you pick the model for each request. Ollama's docs describe no automatic switch from local to cloud, and a :cloud model reached this way needs the local Ollama process running.
So if one tool talks to one local model and Ollama is always up, point the tool at Ollama and stop here. A router earns its place when you want the switch made for you, or want local models and paid subscriptions or API keys behind one endpoint.
And fallback means your prompt goes to a cloud provider whenever the local model fails. If the code must not leave the machine, don't build this chain.
What counts as "the local model can't"
kRouter's fallback reacts to failed requests. It does not judge speed or quality. Here is what each kind of local trouble does to a combo whose first entry is an Ollama model, as of v0.5.163:
| What goes wrong locally | Does the combo move on? |
|---|---|
| Ollama is not running | Yes, after three retries three seconds apart, so that request waits about nine seconds. The local model is then skipped for 30 seconds |
| A typo in the model name, or a model you never pulled (Ollama answers 404) | Yes, and the model is skipped for two minutes. The cloud answers, so the mistake is easy to miss |
| The model has no tool support (Ollama answers 400: the model "does not support tools") | Yes, and the model is skipped for 30 seconds. Coding agents send tools with every request, so all of them end up on the cloud |
| Ollama answers with another error, such as a model that fails to load | Yes, and the model is skipped for 30 seconds |
| No reply within 60 seconds: a large model loading cold, or a long prompt a slow machine is still reading | Treated as a failed connection and retried three times, so the request can wait about four minutes before the next entry |
| The model is generating, slowly | No, if the client streams, as coding agents do. A non-streamed reply sends nothing until it is finished, so one that takes over 60 seconds counts as a failure |
| The reply stalls for 60 seconds mid-stream | The stream ends with an error. No switch halfway through a reply |
| The conversation is longer than Ollama's context length | No. Ollama drops the oldest messages and answers normally (see below) |
| The answer is wrong | No |
While Ollama is down, the first request after each 30-second skip pays the nine seconds again, so if Ollama will be off for a while, point the tool at a cloud model. And "local first" is not "cloud when the task is hard": routing by difficulty is your choice per task, covered at the end.
Building the chain
1. Connect Ollama
Install and start kRouter:
npm install -g @sifxprime/krouter
krouter -tIn the dashboard at http://localhost:20128/dashboard, open Providers, choose Ollama Local and click Add API key. Despite the label, it asks for no key. In its place is Ollama Host URL: leave it blank for http://localhost:11434, and click Check to confirm kRouter can read Ollama's model list. Models are addressed as ollama-local/<tag>, for example ollama-local/gpt-oss:20b.
Ollama Cloud (ollama/...) is a separate provider that calls ollama.com with an API key: a cloud entry, not a local one.
2. Create a combo with the local model first
Open Combos, click Create Combo, and give it a name such as local-first (letters, numbers, -, _ and . only). Add the models in the order you want them tried:
Combo "local-first"
1. ollama-local/gpt-oss:20b your machine
2. ds/deepseek-v4-pro DeepSeek API key
3. ocg/kimi-k2.6 OpenCode Go subscriptionKeep the strategy on Fallback, which tries entries in order. The cloud entries are whatever you already have connected. Don't make the fallback a :cloud model reached through your local Ollama: when Ollama is down, so is that entry.
3. Point your tool at the combo
Any OpenAI-compatible tool takes http://localhost:20128/v1 as its base URL and local-first as the model. For Claude Code, the Claude Code card on the CLI Tools page writes ~/.claude/settings.json, and its model picker lists combos. By hand:
{
"env": {
"ANTHROPIC_BASE_URL": "http://localhost:20128/v1",
"ANTHROPIC_AUTH_TOKEN": "<your-krouter-key>",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "local-first",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "local-first",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "local-first",
"CLAUDE_CODE_MAX_CONTEXT_TOKENS": "65536"
}
}Claude Code sends background work to whatever the Haiku slot names, so give it the combo too. The last line is explained below.
4. Check that it falls over
curl -s http://localhost:20128/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"local-first","messages":[{"role":"user","content":"Say hi"}]}'Callers on the same machine can leave out the key unless Require API key is on (on the Endpoint page) or you set REQUIRE_API_KEY=true. Now stop Ollama (quit the app, or sudo systemctl stop ollama on Linux) and repeat the request. The answer should arrive after the retry pause, from the next entry. Run kRouter in the foreground with krouter -l to watch: the log shows Trying model 1/3: ollama-local/gpt-oss:20b, then failed, trying next.
Context: the failure that never becomes an error
Ollama runs each model with a fixed context length. Its docs set the default by GPU memory (4K under 24 GiB of VRAM, 32K up to 48 GiB, 256K above that) and recommend at least 64,000 tokens for agents and coding tools. When a conversation outgrows the window, Ollama's chat handler drops the oldest messages, keeping the system messages and the latest one, and answers with an ordinary success. (Models in safetensors format skip this step.)
kRouter never sees an error, so there is nothing to fall back on. The agent just forgets how the task started. Three settings prevent it:
- Raise Ollama's window with
OLLAMA_CONTEXT_LENGTH=65536 ollama serveor the context slider in the Ollama app's settings, and confirm withollama ps, whose CONTEXT column shows what the loaded model got. kRouter passes temperature, max tokens and top-p to Ollama, never a context size, so the size has to be set on Ollama's side. - Tell the agent the real size. Claude Code assumes 200K for a model ID it does not recognize, and a combo name is one. Set
CLAUDE_CODE_MAX_CONTEXT_TOKENSto Ollama's value as a plain integer, so Claude Code compacts before Ollama starts dropping messages. The dashboard's Max context menu starts at 200K, so write this one by hand. It is one number for the combo, so the cloud entries compact at the local size too.CLAUDE_CODE_AUTO_COMPACT_WINDOWcan't do this job: its minimum is 100K. - Keep tool output small. RTK, the only kRouter token saver on by default, compresses tool results such as diffs, grep output and file reads before a request reaches Ollama. That counts for more on a 64K window than on a 1M one.
A length error does not trigger fallback either: when a provider rejects a prompt as too long, kRouter returns the error at once instead of trying the next entry. The context window guide covers the Claude Code side in depth.
Two reorderings that can put the cloud first
kRouter adjusts a combo's order on every request (how combos work). Two adjustments matter here:
- Images. If your newest message carries an image and the local model can't read images, a model that can moves ahead of it: a vision-capable entry in your combo or, failing that, the vision adapter's pool, which is on by default and, until you add models to it, uses the free cloud model
mmf/mimo-auto. That happens even in a combo with no cloud entries. kRouter judges vision by model name, not by asking Ollama: Gemma and Qwen VL names count, whilellavaandllama3.2-visiondo not. The toggle is in the Vision Adapter section of the Combos page. - Remaining quota. Entries whose provider reports quota for that exact model move ahead of entries with no figure, and Ollama Local never has one. Antigravity reports quota per model, so once kRouter has used an Antigravity entry behind your local model, that entry can be tried first; providers that report per time window, such as Claude Code, keep their place. The log says
Quota auto-switchwhen it happens. To keep such a provider as a backup, put it in a second combo and make that combo's name your last entry: a nested combo has no quota figure, so it stays put.
Ollama on another machine, or kRouter in Docker
Ollama listens on 127.0.0.1:11434 by default. To reach it from elsewhere, set OLLAMA_HOST on the machine running it (Ollama FAQ) and put its address, such as http://192.168.1.10:11434, in Ollama Host URL. Ollama's local API asks for no key, so keep that port on a network you trust. With two Ollama Local connections, kRouter tries the other machine before the cloud, though its account picker, not the order you added them, decides which goes first. Give each connection its own Name: saving a second one under an existing name replaces the first.
If kRouter runs in Docker, localhost means the container: put http://host.docker.internal:11434 in Ollama Host URL. Docker Desktop on macOS and Windows provides that name. On Linux, add it with --add-host and make Ollama listen beyond loopback with OLLAMA_HOST:
KROUTER_PASSWORD="$(openssl rand -base64 18)" && echo "Dashboard password: $KROUTER_PASSWORD"
docker run -d -p 20128:20128 \
--add-host=host.docker.internal:host-gateway \
-e INITIAL_PASSWORD="${KROUTER_PASSWORD:?run the line above first}" \
-v "$HOME/.krouter:/app/data" --name krouter sifxprime/krouter:latestChoosing a setup
| Your situation | Set-up |
|---|---|
| One tool, one local model, Ollama always running | Point the tool at Ollama directly. No router |
| Local by default, keep working when Ollama is off or broken | A local-first combo |
| Code must never leave the machine | No cloud entries, and the Vision Adapter turned off |
| Local for Claude Code's background calls only | local-first in the Haiku slot, a cloud model in Sonnet and Opus |
| The local model is too weak for some tasks | local-first for routine work, a cloud model you pick for the hard parts |
Common questions
Does Ollama fall back to a cloud model on its own?
No. Ollama runs cloud models through the local server when you sign in and use a :cloud model name, but you choose the model for each request, and its docs describe no automatic switch. Those cloud models also need the local Ollama process, so they can't cover for it being down.
Will kRouter switch to the cloud when my local model is slow?
No. It moves on when a request fails: Ollama is not running, the model is missing, Ollama returns an error, or nothing comes back within 60 seconds. Once a streamed reply has started, it stays local however slowly it arrives.
Why does my local model forget the start of the task?
Usually Ollama's context length. When a conversation outgrows it, Ollama drops the oldest messages without an error. Raise OLLAMA_CONTEXT_LENGTH, check it with ollama ps, and set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the same number so Claude Code compacts first.
Why is a cloud model answering while Ollama is running?
Check the model name first: a typo or an unpulled model gets a 404, and the combo moves on. Then check that the model supports tools (ollama show <model> lists tools under Capabilities when it does), since every request from a coding agent carries them. After that, look for an image in the request, or a provider with per-model quota, such as Antigravity, behind the local entry. krouter -l shows each attempt.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- Save moneyCaveman vs Ponytail: two ways to cut output tokensCaveman cuts the words, Ponytail cuts the code. What each prompt does, what their benchmarks measured, and why you should run only one of them.
- Save moneyGemini CLI's free tier is gone: what still works for freeGoogle stopped serving Gemini CLI to free, AI Pro and Ultra users on June 18, 2026. Who kept access, what is still free, and how to keep using Gemini.
- Save moneyQwen Code's free tier is gone: keep the CLI, swap the modelQwen OAuth's free tier closed on April 15, 2026. What still works in Qwen Code, where Qwen models come from now, and how to keep the CLI answering.
Relevant docs
- Token savers: RTK, Caveman, Ponytail, Headroom, PXPIPEFive ways kRouter can shrink a request before it reaches the provider. Only RTK is on by default. What each saver does, what it needs, and its limits.
- Combos & fallbackPut several models behind one name. kRouter falls back from one to the next, rotates them, or asks a panel of models and merges the answers.
- ProvidersConnect 95+ AI providers to kRouter: OAuth subscriptions, free services, free-tier and pay-per-token API keys, and browser-cookie logins.