Claude Cowork with other models: third-party inference setup
Claude Desktop can run Cowork through any Messages API gateway. Setting it up with kRouter, what the app assumes about your models, and what you give up.
You enable Developer mode in Claude Desktop and find Developer → Configure Third-Party Inference…. You set the provider to Gateway, paste your gateway's URL and key, restart, and the model picker is empty. Or it shows only the few Claude models your gateway happens to list, and none of the Kimi, GLM or DeepSeek models you set it up for.
The setting works. It just makes assumptions that hold for Claude models on Bedrock and stop holding once something else answers. Four of them decide whether Cowork on other models is usable, and kRouter's Cowork card deals with two of them for you.
What third-party inference changes
Normally Claude Desktop signs in to your Anthropic account and sends every request to Anthropic. In third-party mode set up on one machine, you skip the Anthropic sign-in, and Chat, Cowork and Code send every model call to the provider you configure: Amazon Bedrock, Google Cloud's Agent Platform (Vertex), Microsoft Foundry, the Anthropic API, or a gateway. Conversation history stays on your disk. Anthropic's feature comparison describes the billing as token-based consumption charged by your provider, with no seat licensing. No Claude plan is involved.
The mode is built for companies that push configuration with MDM, but a single machine can be set up from the app: Help → Troubleshooting → Enable Developer Mode, then Developer → Configure Third-Party Inference…. Cowork itself needs macOS 14 or later, or Windows 10 version 2004 or later installed from the .msix package, plus working hardware virtualization.
A gateway has one hard requirement: POST /v1/messages with streaming and tool use. GET /v1/models is optional. Anthropic's page is written for Claude models, but nothing in the protocol checks what sits behind the gateway. If it translates for another provider, that provider answers.
1. The picker only lists Claude-looking IDs
When you give Claude Desktop no model list, it calls /v1/models and shows, in Anthropic's words, "only models whose IDs are recognizably Claude." A gateway that serves kimi-k2.6 and glm-5.3 therefore produces an empty picker. LiteLLM's Cowork guide describes the same filter: the ID has to contain claude or anthropic.
The fix: write the list yourself in inferenceModels. An explicit list replaces discovery, so the filter no longer decides what you see, and the first entry becomes the default model. kRouter's Cowork card always writes this list, and every combo it creates gets a claude- prefix. Use plain names such as claude-work. The app sizes the context window from the model ID (see the next section), so a name that contains a real Claude model ID, such as claude-opus-4-6-mix, may be sized as that Claude model rather than with the 200K default.
2. The context window is a guess
Through a gateway, a session cannot ask the provider how large the model's window is. Claude Desktop assumes 200K tokens for any model ID it does not recognize and compacts the conversation as it nears that size. To compact, it sends the whole conversation in one request and asks for a summary.
How to tell: a backend with a smaller window starts rejecting requests before Cowork decides to compact. If Cowork then tries to compact, the summary request carries the same oversized conversation and fails too, and Cowork shows:
This conversation is too long to continue. Start a new session, or remove some tools to free up space.The fix: put only models with a window of at least 200K behind a combo Cowork uses. Larger windows work too; Cowork just compacts at 200K. If every model in a combo really accepts 1M-token requests, adding "supports1m": true to that entry in the config file adds a 1M option to the picker. The card does not write that field, so add it by hand after each Apply. The context window post covers the same guess in Claude Code.
3. Web search is a server tool
Cowork's built-in Web Search is not run by the app. Anthropic documents it as a server-side tool executed by your inference provider, and a gateway that translates to another API has to run the search itself. kRouter does not run searches on a provider's behalf, so with a non-Claude backend, plan on a search MCP server instead.
The fix: that is what the Web Search & Fetch (Exa) option on kRouter's Cowork card adds: Exa's hosted MCP server. When Exa's or Tavily's tools are in a request bound for a non-Claude provider, kRouter removes Cowork's built-in WebSearch and WebFetch from the tool list, so the model sees one way to search instead of two.
Claude Desktop has its own option too: a built-in Web search server under Connectors in its configuration window, which runs the search from the app with your Brave, Tavily or Exa key. It is saved in the same config file, so a later Apply from kRouter removes it.
4. The agent loop was built around Claude
Cowork plans, calls tools, reads the results and keeps going for many steps. Models differ in how reliably they do that, and one that is fine in a chat can stall halfway through a task. No setting fixes this. Try candidate models on your own tasks first; the Kimi, GLM and DeepSeek comparison covers the same trade-off in Claude Code. A combo's fallback covers requests that fail, such as rate limits and outages. It cannot catch a model that answers badly.
Setting it up with kRouter
Install kRouter and start it in the tray:
npm install -g @sifxprime/krouter
krouter -t --host 127.0.0.1kRouter listens on every interface by default; --host 127.0.0.1 keeps it on this computer, which is all a local Claude Desktop needs.
- Open
http://127.0.0.1:20128/dashboardand connect the providers you want under Providers. - Go to CLI Tools and open the Claude Cowork card. If it says Claude Desktop (Cowork mode) is not detected, install Claude Desktop, launch it once, and reopen the card.
- Select Endpoint: keep the default,
http://127.0.0.1:20128/v1. - API Key: pick any key from the Endpoint page. With none, the card writes the placeholder
sk_krouter, which works only from this machine and only while Require API key on the Endpoint page is off. - Models: click + Combo, name it (the card adds the
claude-prefix) and add models in fallback order. - MCP and Tools: Exa and Tavily are preselected, and Tavily needs an OAuth sign-in. Browser MCP drives your running Chrome through the Browser MCP extension; kRouter starts its server with
npx -y @browsermcp/mcp@latest, so the first use downloads it from npm. + Browse lists servers from Anthropic's MCP registry, minus those that run through claude.ai or need per-organization setup. - Click Apply. Quit Claude Desktop completely, not just its window, and reopen it. If the sign-in screen appears, choose the third-party option instead of signing in to Anthropic.
The card writes configLibrary/<id>.json inside Claude's Claude-3p folder, the local configuration that Claude's own window also edits. Manual Config shows the core of it:
{
"inferenceProvider": "gateway",
"inferenceGatewayBaseUrl": "http://127.0.0.1:20128/v1",
"inferenceGatewayApiKey": "<your-krouter-key>",
"inferenceModels": [{ "name": "claude-work" }]
}You can type the same values into Claude's window instead: provider Gateway, that URL, your key, auth scheme Bearer, and your combo names under Models. Claude Desktop appends /v1/messages itself, and kRouter accepts the base URL with or without /v1.
The card writes files on the computer kRouter runs on, and kRouter accepts that request only from the same machine. If kRouter runs on a server, enter the values in Claude's window yourself, using the Tunnel or Tailscale URL and a real API key. Remote callers always need one.
What Apply changes besides the model
Apply does more than set the gateway. It rewrites the whole file, so anything you added in Claude's own window, such as Claude's built-in web search or a supports1m entry, is gone afterwards. And every time, it writes a fixed, permissive profile:
coworkEgressAllowedHosts: ["*"]. By default the Cowork sandbox can reach only your inference endpoint. With*, Web Fetch and the agent's shell (curl,pip install) can reach any host.- No approval prompts for the MCP servers it adds. Each tool it knows of gets a
toolPolicyofallow. - Desktop extensions on. Claude's default is off. The profile also leaves unsigned extensions and user-added local MCP servers allowed, which are Claude's defaults anyway.
- Crash reports, analytics and non-essential services off. Anthropic's telemetry page says the last switch makes artifact previews static and shows connector widgets as plain text, and that with crash reports off, support becomes manual.
If you want Claude's defaults for any of these, change those keys in the file after applying, or configure Claude's window by hand. The next Apply writes the profile again.
Check that it works
From the same machine, no key needed:
curl -s http://127.0.0.1:20128/v1/messages \
-H "content-type: application/json" \
-d '{"model":"claude-work","max_tokens":50,"messages":[{"role":"user","content":"Reply with ok"}]}'If Require API key is on (kRouter asks for it before it opens a tunnel), add -H "x-api-key: <your-krouter-key>". A reply with the model's text means the combo routes. On the Claude side, Help → Troubleshooting → Generate Diagnostic Report → Export to file gives you a zip: provider-status.txt says whether the provider settings are valid, and deployment-mode.txt whether the app runs in third-party mode.
What does not work
An OpenAI-only endpoint. Claude Desktop sends Messages API requests. A gateway that serves only /v1/chat/completions fails every one.
Changing the configuration while the app runs. It is read at launch, so quit fully and reopen.
A work laptop with a Claude profile from IT. A managed configuration wins, and Claude ignores locally written values, so neither kRouter nor the in-app window changes anything.
Features that need Anthropic's servers. Anthropic's feature table lists no mobile app, no claude.ai web access, no voice mode, no sharing between users and no sessions in Anthropic's cloud in third-party mode. Claude in Chrome works there only for organizations on the Enterprise Admin Console. To drive your own Chrome, use Browser MCP on the card; Claude Desktop's built-in browser pane is another option once its builtinBrowserEnabled key is on.
Choosing
| You want | Use | Why |
|---|---|---|
| Cowork on Kimi, GLM, DeepSeek or a mix with fallback | Third-party mode pointed at kRouter | It translates the Messages API and writes the model list |
| Cowork on one provider that already has an Anthropic-compatible endpoint | Point Claude Desktop at that provider | For example DeepSeek's /anthropic endpoint; no router needed |
| Your company already runs Bedrock, Vertex or Foundry | That provider in Claude's window | kRouter adds nothing |
| Keep your Claude account and swap only some models | MITM mode | Needs a local root CA and admin rights |
| Claude Code in a terminal | ANTHROPIC_BASE_URL | Claude Desktop is not involved |
If you have a Claude plan and only want Claude models, sign in normally. kRouter is not the answer there.
Common questions
Do I need a Claude subscription to use Cowork with third-party inference?
No. In third-party mode set up on your own machine, Claude Desktop does not sign in to Anthropic, and you pay whatever the backends behind your gateway cost. If you connect a subscription account to kRouter, such as Claude Code, Codex, GitHub Copilot, Antigravity or Kiro, the dashboard shows a Risk Notice first: those sessions are not licensed for proxy use, and the account can be restricted.
Why is the model picker empty after I point Claude Desktop at my gateway?
With no model list, Claude Desktop discovers models from /v1/models and shows only IDs that look like Claude. List the models explicitly in inferenceModels. kRouter's Cowork card does this and names its combos with a claude- prefix.
Can Claude Desktop use kRouter on localhost?
Yes, when both run on the same computer. The card's default endpoint is http://127.0.0.1:20128/v1, and local callers need no API key unless Require API key is on. For a kRouter on another machine, use its Tunnel or Tailscale URL with a real key.
How do I switch back to my normal Claude account?
Choose the Anthropic sign-in option on Claude Desktop's sign-in screen; the two modes keep separate histories, and switching deletes neither. Reset on the card empties the configuration kRouter wrote and stops auto-approving its MCP servers. It does not remove the "deploymentMode": "3p" line it added to claude_desktop_config.json, so delete that by hand if you want no trace of third-party mode left.
Does Cowork work as well on other models as on Claude?
Not necessarily. Cowork is a long agent loop, and models differ in how reliably they call tools. Test on your own tasks, and keep each backend's window at 200K or more.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- How kRouter worksOpenCode Go in Claude Code and Codex: three protocolsOpenCode Go serves its models on three APIs. Which model needs which, why the wrong one fails with ModelProtocolUnsupported, and how to fix it.
- How kRouter worksXcode with any model: chat, Claude Agent and CodexXcode 26 and 27 can run chat, Claude Agent and Codex on other models, but each reads its own config. How to point all three at one local router.
- How kRouter worksLive model catalog: new models without a kRouter updatekRouter asks each provider what it serves right now, so a model launched today can be added from the dashboard without waiting for a kRouter release.
Relevant docs
- Core conceptsProviders, combos, account routing, token savers, MITM mode, quota tracking, the response cache and remote access: the eight ideas behind kRouter.
- The Zenith Routing EngineHow kRouter picks which of your accounts serves a request: Zenith scoring, conversation stickiness, cooldowns, and the round-robin, P2C and random options.
- Token savers: RTK, Caveman, Ponytail, Headroom, PXPIPEFive ways kRouter can shrink a request before it reaches the provider. Only RTK is on by default. What each saver does, what it needs, and its limits.