Skip to main content
kRouter
All posts
Save money

Local first, cloud when it can't: Ollama with cloud fallback

Use a local Ollama model and switch to a cloud model only when it fails. What triggers the switch, what silently does not, and how to set it up.

Kodelyth · The team behind kRouter
· Updated
9 min read

Your coding agent points at a model on your own machine, and it works until the laptop restarts and Ollama doesn't come back, or the desktop running it goes to sleep. The agent stops with a connection error, and checking by hand shows why:

$ curl http://localhost:11434/api/tags
curl: (7) Failed to connect to localhost port 11434 after 0 ms: Couldn't connect to server

The setup people want: use the local model, and quietly use a cloud one when the local one can't answer. But "can't answer" covers several situations, and only some of them produce an error that anything in front of Ollama can act on. Here is which is which, how to build the chain with kRouter, and when you don't need a router at all.

When you do not need a router

Since v0.14.0, announced in January 2026, Ollama serves Anthropic's Messages API on its usual port (Ollama's announcement), next to its own API and an OpenAI-compatible one. For Claude Code, Ollama's guide is one command, ollama launch claude, or three variables:

export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434

Ollama also runs cloud models through the same local server: sign in and use a model name ending in :cloud (Ollama's cloud docs). But you pick the model for each request. Ollama's docs describe no automatic switch from local to cloud, and a :cloud model reached this way needs the local Ollama process running.

So if one tool talks to one local model and Ollama is always up, point the tool at Ollama and stop here. A router earns its place when you want the switch made for you, or want local models and paid subscriptions or API keys behind one endpoint.

And fallback means your prompt goes to a cloud provider whenever the local model fails. If the code must not leave the machine, don't build this chain.

What counts as "the local model can't"

kRouter's fallback reacts to failed requests. It does not judge speed or quality. Here is what each kind of local trouble does to a combo whose first entry is an Ollama model, as of v0.5.163:

What goes wrong locallyDoes the combo move on?
Ollama is not runningYes, after three retries three seconds apart, so that request waits about nine seconds. The local model is then skipped for 30 seconds
A typo in the model name, or a model you never pulled (Ollama answers 404)Yes, and the model is skipped for two minutes. The cloud answers, so the mistake is easy to miss
The model has no tool support (Ollama answers 400: the model "does not support tools")Yes, and the model is skipped for 30 seconds. Coding agents send tools with every request, so all of them end up on the cloud
Ollama answers with another error, such as a model that fails to loadYes, and the model is skipped for 30 seconds
No reply within 60 seconds: a large model loading cold, or a long prompt a slow machine is still readingTreated as a failed connection and retried three times, so the request can wait about four minutes before the next entry
The model is generating, slowlyNo, if the client streams, as coding agents do. A non-streamed reply sends nothing until it is finished, so one that takes over 60 seconds counts as a failure
The reply stalls for 60 seconds mid-streamThe stream ends with an error. No switch halfway through a reply
The conversation is longer than Ollama's context lengthNo. Ollama drops the oldest messages and answers normally (see below)
The answer is wrongNo

While Ollama is down, the first request after each 30-second skip pays the nine seconds again, so if Ollama will be off for a while, point the tool at a cloud model. And "local first" is not "cloud when the task is hard": routing by difficulty is your choice per task, covered at the end.

Building the chain

1. Connect Ollama

Install and start kRouter:

npm install -g @sifxprime/krouter
krouter -t

In the dashboard at http://localhost:20128/dashboard, open Providers, choose Ollama Local and click Add API key. Despite the label, it asks for no key. In its place is Ollama Host URL: leave it blank for http://localhost:11434, and click Check to confirm kRouter can read Ollama's model list. Models are addressed as ollama-local/<tag>, for example ollama-local/gpt-oss:20b.

Ollama Cloud (ollama/...) is a separate provider that calls ollama.com with an API key: a cloud entry, not a local one.

2. Create a combo with the local model first

Open Combos, click Create Combo, and give it a name such as local-first (letters, numbers, -, _ and . only). Add the models in the order you want them tried:

Combo "local-first"
  1. ollama-local/gpt-oss:20b   your machine
  2. ds/deepseek-v4-pro         DeepSeek API key
  3. ocg/kimi-k2.6              OpenCode Go subscription

Keep the strategy on Fallback, which tries entries in order. The cloud entries are whatever you already have connected. Don't make the fallback a :cloud model reached through your local Ollama: when Ollama is down, so is that entry.

3. Point your tool at the combo

Any OpenAI-compatible tool takes http://localhost:20128/v1 as its base URL and local-first as the model. For Claude Code, the Claude Code card on the CLI Tools page writes ~/.claude/settings.json, and its model picker lists combos. By hand:

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:20128/v1",
    "ANTHROPIC_AUTH_TOKEN": "<your-krouter-key>",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "local-first",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "local-first",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "local-first",
    "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "65536"
  }
}

Claude Code sends background work to whatever the Haiku slot names, so give it the combo too. The last line is explained below.

4. Check that it falls over

curl -s http://localhost:20128/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local-first","messages":[{"role":"user","content":"Say hi"}]}'

Callers on the same machine can leave out the key unless Require API key is on (on the Endpoint page) or you set REQUIRE_API_KEY=true. Now stop Ollama (quit the app, or sudo systemctl stop ollama on Linux) and repeat the request. The answer should arrive after the retry pause, from the next entry. Run kRouter in the foreground with krouter -l to watch: the log shows Trying model 1/3: ollama-local/gpt-oss:20b, then failed, trying next.

Context: the failure that never becomes an error

Ollama runs each model with a fixed context length. Its docs set the default by GPU memory (4K under 24 GiB of VRAM, 32K up to 48 GiB, 256K above that) and recommend at least 64,000 tokens for agents and coding tools. When a conversation outgrows the window, Ollama's chat handler drops the oldest messages, keeping the system messages and the latest one, and answers with an ordinary success. (Models in safetensors format skip this step.)

kRouter never sees an error, so there is nothing to fall back on. The agent just forgets how the task started. Three settings prevent it:

  • Raise Ollama's window with OLLAMA_CONTEXT_LENGTH=65536 ollama serve or the context slider in the Ollama app's settings, and confirm with ollama ps, whose CONTEXT column shows what the loaded model got. kRouter passes temperature, max tokens and top-p to Ollama, never a context size, so the size has to be set on Ollama's side.
  • Tell the agent the real size. Claude Code assumes 200K for a model ID it does not recognize, and a combo name is one. Set CLAUDE_CODE_MAX_CONTEXT_TOKENS to Ollama's value as a plain integer, so Claude Code compacts before Ollama starts dropping messages. The dashboard's Max context menu starts at 200K, so write this one by hand. It is one number for the combo, so the cloud entries compact at the local size too. CLAUDE_CODE_AUTO_COMPACT_WINDOW can't do this job: its minimum is 100K.
  • Keep tool output small. RTK, the only kRouter token saver on by default, compresses tool results such as diffs, grep output and file reads before a request reaches Ollama. That counts for more on a 64K window than on a 1M one.

A length error does not trigger fallback either: when a provider rejects a prompt as too long, kRouter returns the error at once instead of trying the next entry. The context window guide covers the Claude Code side in depth.

Two reorderings that can put the cloud first

kRouter adjusts a combo's order on every request (how combos work). Two adjustments matter here:

  • Images. If your newest message carries an image and the local model can't read images, a model that can moves ahead of it: a vision-capable entry in your combo or, failing that, the vision adapter's pool, which is on by default and, until you add models to it, uses the free cloud model mmf/mimo-auto. That happens even in a combo with no cloud entries. kRouter judges vision by model name, not by asking Ollama: Gemma and Qwen VL names count, while llava and llama3.2-vision do not. The toggle is in the Vision Adapter section of the Combos page.
  • Remaining quota. Entries whose provider reports quota for that exact model move ahead of entries with no figure, and Ollama Local never has one. Antigravity reports quota per model, so once kRouter has used an Antigravity entry behind your local model, that entry can be tried first; providers that report per time window, such as Claude Code, keep their place. The log says Quota auto-switch when it happens. To keep such a provider as a backup, put it in a second combo and make that combo's name your last entry: a nested combo has no quota figure, so it stays put.

Ollama on another machine, or kRouter in Docker

Ollama listens on 127.0.0.1:11434 by default. To reach it from elsewhere, set OLLAMA_HOST on the machine running it (Ollama FAQ) and put its address, such as http://192.168.1.10:11434, in Ollama Host URL. Ollama's local API asks for no key, so keep that port on a network you trust. With two Ollama Local connections, kRouter tries the other machine before the cloud, though its account picker, not the order you added them, decides which goes first. Give each connection its own Name: saving a second one under an existing name replaces the first.

If kRouter runs in Docker, localhost means the container: put http://host.docker.internal:11434 in Ollama Host URL. Docker Desktop on macOS and Windows provides that name. On Linux, add it with --add-host and make Ollama listen beyond loopback with OLLAMA_HOST:

KROUTER_PASSWORD="$(openssl rand -base64 18)" && echo "Dashboard password: $KROUTER_PASSWORD"
docker run -d -p 20128:20128 \
  --add-host=host.docker.internal:host-gateway \
  -e INITIAL_PASSWORD="${KROUTER_PASSWORD:?run the line above first}" \
  -v "$HOME/.krouter:/app/data" --name krouter sifxprime/krouter:latest

Choosing a setup

Your situationSet-up
One tool, one local model, Ollama always runningPoint the tool at Ollama directly. No router
Local by default, keep working when Ollama is off or brokenA local-first combo
Code must never leave the machineNo cloud entries, and the Vision Adapter turned off
Local for Claude Code's background calls onlylocal-first in the Haiku slot, a cloud model in Sonnet and Opus
The local model is too weak for some taskslocal-first for routine work, a cloud model you pick for the hard parts

Common questions

Does Ollama fall back to a cloud model on its own?

No. Ollama runs cloud models through the local server when you sign in and use a :cloud model name, but you choose the model for each request, and its docs describe no automatic switch. Those cloud models also need the local Ollama process, so they can't cover for it being down.

Will kRouter switch to the cloud when my local model is slow?

No. It moves on when a request fails: Ollama is not running, the model is missing, Ollama returns an error, or nothing comes back within 60 seconds. Once a streamed reply has started, it stays local however slowly it arrives.

Why does my local model forget the start of the task?

Usually Ollama's context length. When a conversation outgrows it, Ollama drops the oldest messages without an error. Raise OLLAMA_CONTEXT_LENGTH, check it with ollama ps, and set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the same number so Claude Code compacts first.

Why is a cloud model answering while Ollama is running?

Check the model name first: a typo or an unpulled model gets a 404, and the combo moves on. Then check that the model supports tools (ollama show <model> lists tools under Capabilities when it does), since every request from a coding agent carries them. After that, look for an image in the request, or a provider with per-model quota, such as Antigravity, behind the local entry. krouter -l shows each attempt.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter