Skip to main content
kRouter
All posts
Comparisons

Claude Code with GLM, Kimi or DeepSeek: setup and trade-offs

GLM-5.3, Kimi K3 and DeepSeek V4 can all drive Claude Code. What each costs, which slot each suits, what breaks, and when you need a router.

Kodelyth · The team behind kRouter
· Updated
9 min read

Z.ai, Moonshot and DeepSeek all publish a Claude Code setup. You set a base URL, a key and a few model variables, run claude, and the first prompt works. The failures come later, and most have nothing to do with how clever the model is:

400 invalid thinking: only type=enabled is allowed for this model

That is Kimi K2.7 Code with thinking switched off. The others are just as mundane: a subagent asks for a Claude model the endpoint does not serve, a screenshot goes to a model that reads only text, or a plan's 5-hour window runs out mid-afternoon.

This post does not crown a winner. We have not run a controlled comparison of the three as Claude Code's agent, and each vendor's benchmark tables are self-reported. What can be checked is below, followed by a way to test them on your own work.

What each vendor offers Claude Code

Each model is served on an Anthropic-compatible endpoint by its own vendor. Figures are from Z.ai's, Kimi's and DeepSeek's pricing pages in October 2026:

Model (ID)Claude Code endpointContextImagesPer 1M tokens, in / out
GLM-5.3 (glm-5.3)https://api.z.ai/api/anthropic1MNo$1.40 / $4.40
GLM-5.3-Flash (glm-5.3-flash)same1MYes$0.15 / $0.50
Kimi K3 (kimi-k3)https://api.moonshot.ai/anthropic1,048,576Yes$3.00 / $15.00
Kimi K2.7 Code (kimi-k2.7-code)same262,144Yes$0.95 / $4.00
DeepSeek V4 Pro (deepseek-v4-pro)https://api.deepseek.com/anthropic1MNo$0.66 / $1.98 off-peak, double at peak
DeepSeek V4.1 Flash (deepseek-flash)same1MYes$0.15 / $0.60 off-peak, double at peak

DeepSeek's peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, Chinese public holidays excepted. Cached input matters too, because Claude Code sends the whole conversation with every request: per million cached tokens, Kimi K3 charges $0.30, GLM-5.3 $0.26 and DeepSeek V4 Pro $0.022 off-peak. Kimi K3 also bills cache writes, $3.00 per million at the default 5-minute lifetime.

Flat-rate plans exist too:

  • GLM Coding Plan, from $18 a month, for GLM-5.3 and GLM-5.3-Flash. Usage is counted in credits per 5 hours and per week: 2,000 and 10,000 on Lite, up to 28,000 and 140,000 on Max (Z.ai's plan page).
  • OpenCode Go, $10 a month (Go Plus $40), with an allowance per model. OpenCode's own estimate for one 5-hour window on Go is about 110 requests on Kimi K3, 220 on GLM-5.3, 1,050 on DeepSeek V4 Pro and 26,000 on DeepSeek V4.1 Flash. The catch: Go serves Kimi, GLM and DeepSeek only on Chat Completions, which Claude Code does not speak (details).
  • Kimi Code, part of Kimi's membership, also works in Claude Code, but through its own endpoint and with keys that do not work on the Moonshot platform. Everything below assumes a platform key.

What the vendors put in each slot

Claude Code asks for models through aliases: opus, sonnet, haiku (which also runs background work such as summarizing) and fable if you select it. Subagents run on the main model unless something assigns one: CLAUDE_CODE_SUBAGENT_MODEL for the general-purpose subagent, or an alias in a subagent's definition, such as haiku for the built-in claude-code-guide. Each has to resolve to something your endpoint serves, and the vendors' guides fill them differently:

Vendor guideOpus and SonnetHaikuAlso sets
Z.aiglm-5.3[1m]glm-5.3-flash[1m]1M compact window, 3,000-second API timeout
Moonshotkimi-k3[1m]kimi-k2.7-codeFable and subagents on kimi-k3[1m], effort max, 1M compact window
DeepSeekdeepseek-flash[1m]deepseek-flashSubagents on deepseek-flash, effort max, 786,432-token compact window

Two things stand out. DeepSeek's guide puts V4 Pro in no slot at all: it runs Flash everywhere, and V4 Pro answers only a request for a claude-opus name, through DeepSeek's own name mapping. Moonshot fills every slot, and its FAQ says why: an empty slot makes background tasks and subagents ask for a model name its endpoint does not recognize.

The [1m] suffix tells Claude Code to assume a 1M-token window. Without it, Claude Code assumes 200K for a model ID it does not know and compacts early; the context window post covers the details.

What breaks

An empty slot. Whatever asks for that alias, such as a subagent defined with model: haiku or /model opus, gets the alias's Claude model ID. Connected directly, DeepSeek and Z.ai map Claude names to a model of their own; Moonshot does not. Through kRouter, a bare claude- name goes to its Anthropic provider, and with no Anthropic key connected you get No active credentials for provider: anthropic (the fix). Map Opus, Sonnet and Haiku, plus Fable and subagents if you use them.

Thinking switched off. Kimi K2.7 Code rejects it with the 400 above. Z.ai documents the same rule for GLM-5.3 and GLM-5.3-Flash: thinking is always on, and a request that disables it fails. Keep thinking on (Option+T on macOS, Alt+T elsewhere).

A screenshot to a text-only model. GLM-5.3 and DeepSeek V4 Pro do not accept images. The model you paste into needs to be Kimi K3, Kimi K2.7 Code, GLM-5.3-Flash or DeepSeek's deepseek-flash.

The expensive model in the background. The Haiku slot also runs background work. On OpenCode Go, where Kimi K3 gets about 110 requests per 5 hours, K3 does not belong there.

Testing them on your own work

A ranking that matters comes from your repository, and it takes an afternoon:

  1. Pick three tasks you have already finished, so you know what a correct result looks like.
  2. Run each on each candidate. claude --model <model-id> starts a session on one model without touching your settings.
  3. Note whether it finished without your help, how often a tool call failed, and what it cost in the vendor's console.
  4. Run the close calls twice. One run is an anecdote.

One endpoint for all three: where kRouter fits

Claude Code has one ANTHROPIC_BASE_URL, so connected directly, every slot goes to the same vendor. Mixing vendors across slots, falling back when a plan window runs out, or reaching OpenCode Go's models needs something in between. In kRouter the routes look like this:

Model ID in kRouterGoes toUpstream format
glm/glm-5.3, glm/glm-5.3-flashZ.ai's Anthropic endpointAnthropic, not translated
ds/deepseek-v4-pro, ds/deepseek-v4-flashDeepSeek's Chat CompletionsTranslated
ocg/kimi-k3, ocg/glm-5.3, ocg/deepseek-v4-proOpenCode Go's Chat CompletionsTranslated
moonshot/kimi-k3Moonshot's Anthropic endpoint, as a provider you addAnthropic, passed through

DeepSeek now serves the old deepseek-v4-flash name with V4.1 Flash. glm-5.3-flash is not in kRouter's GLM list as of 0.5.163, so type the ID rather than picking it. kRouter's built-in Kimi provider points at Kimi's coding endpoint rather than the platform one and lists no K3, which is why the last row is a custom provider.

The translated rows differ from the Anthropic ones in a way that matters: kRouter's conversion to Chat Completions does not carry Claude Code's thinking setting or effort level. On ds/ and ocg/, the thinking toggle and the CLAUDE_CODE_EFFORT_LEVEL=max that Moonshot and DeepSeek recommend do nothing, and the model runs at its provider's default, which on DeepSeek's own API is thinking on at high effort. For DeepSeek's maximum, put ds/deepseek-v4-pro-max in the slot: kRouter sends it to V4 Pro with reasoning_effort set to max.

Install and start kRouter:

npm install -g @sifxprime/krouter
krouter -t

Open http://localhost:20128/dashboard, go to Providers and add your key to GLM Coding (Z.ai), DeepSeek or OpenCode Go. For Kimi K3, click Add Anthropic Compatible, give it a name, set the prefix moonshot and the base URL https://api.moonshot.ai/anthropic/v1, click Create, then use Add API Key on the new provider.

Then open CLI Tools, expand Claude Code, enter a provider/model-id for each slot and click Apply. kRouter merges the result into ~/.claude/settings.json:

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:20128/v1",
    "ANTHROPIC_AUTH_TOKEN": "<your-krouter-key>",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "moonshot/kimi-k3",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "glm/glm-5.3",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm/glm-5.3-flash",
    "CLAUDE_CODE_SUBAGENT_MODEL": "glm/glm-5.3"
  }
}

The card writes the three slots. Add the subagent line, and ANTHROPIC_DEFAULT_FABLE_MODEL if you use Fable, by hand. This shows the wiring, not a recommendation. The vendors' [1m] suffix works on these IDs too, because kRouter strips it before routing (since 0.5.150); declare 1M only for a route that serves it. Confirm a route with one real request before starting Claude Code:

curl -s http://localhost:20128/v1/messages \
  -H "Content-Type: application/json" \
  -d '{"model": "moonshot/kimi-k3", "max_tokens": 32, "stream": false, "messages": [{"role": "user", "content": "Reply with the word ok"}]}'

Any JSON message back, even one cut short by thinking, means the route and key work; an error names what is wrong.

For fallback, create a combo on the Combos page, for example open-main holding glm/glm-5.3, ocg/kimi-k3 and ds/deepseek-v4-pro, and put open-main in a slot. When one entry hits its limit, the next answers (combos). A switch mid-session hands your conversation to a different model, so test the pair first, and do not start a combo name with claude- (the context window post explains why).

Three more pages help. Quota Tracker shows the GLM Coding Plan's 5-hour and weekly use with reset times; OpenCode Go and DeepSeek balances stay in their own consoles. Usage, on its Details tab, logs the last 1,000 requests by default, each with the client request, the translated provider request and the raw provider response. A part over 5 KB is kept only as a 200-character preview, and a Claude Code request is nearly always bigger than that, but a provider's error reply fits, so this is where to read exactly what an upstream rejected. And Token Saver holds RTK, the one saver on by default, which compresses tool output such as diffs and logs on all of these routes.

Where kRouter is not the answer

One vendor for every slot. Connect directly with that vendor's guide. It is one hop fewer, and the vendor tests its own endpoint.

DeepSeek, when its reasoning matters. Translated routes drop Claude Code's thinking blocks. Where DeepSeek requires reasoning_content in tool-calling conversations, kRouter fills in a one-character placeholder, for ds/ and for DeepSeek and Kimi models on OpenCode Go. That avoids a 400, but the model does not see its earlier reasoning. DeepSeek's own Anthropic endpoint speaks Claude Code's format natively, and DeepSeek says it supports Claude Code's web search there; connecting directly is the documented way to get both. To keep the native format behind kRouter, add https://api.deepseek.com/anthropic/v1 as an Anthropic Compatible provider with a prefix no built-in provider uses, such as dsa; reusing a built-in one like ds makes the two routes collide. Claude Code's requests to that kind of provider pass through with their thinking blocks intact.

Plan terms. Z.ai's usage policy limits the GLM Coding Plan to officially supported tools. Claude Code is one; a gateway in between is not on the list, so that risk is yours. OpenCode Go lists Claude Code as a validated client and asks for a stable session header, which kRouter sends.

Support. Anthropic says it does not support routing Claude Code to non-Claude models through any gateway, so do not expect help from that side.

Choosing

If you wantMain slotsHaiku slotConnect
The lowest per-token pricesdeepseek-flash or glm-5.3-flashThe sameThat vendor directly
One flat bill for all three familiesGLM-5.3, Kimi K3 sparinglyDeepSeek V4.1 FlashkRouter, ocg/
A GLM subscriptionGLM-5.3GLM-5.3-FlashZ.ai directly; kRouter glm/ at your own risk under the plan terms
Screenshots in the loopKimi K3, GLM-5.3-Flash or deepseek-flashAnyEither
To keep working when a window runs outA combo across vendorsA Flash modelkRouter

Common questions

Can Claude Code use GLM, Kimi or DeepSeek models?

Yes. Z.ai, Moonshot and DeepSeek each run an Anthropic-compatible endpoint and publish the settings: ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN and a model for each slot. Anthropic does not support this setup, so the vendor's guide is the reference.

Which open model is best for Claude Code?

No published table can answer that for your code; each vendor's benchmarks are self-reported. What differs measurably is price, image input, context and plan limits. Run three tasks you have already solved on each model and compare how often each finishes without help.

Why do background tasks or subagents fail when the main chat works?

Usually a slot is empty. Whatever asks for that alias, such as a subagent defined with model: haiku, then gets a Claude model ID the endpoint does not serve; Moonshot's FAQ gives the same cause. Set the Haiku slot, plus CLAUDE_CODE_SUBAGENT_MODEL and the Fable slot if you use them.

Can GLM-5.3 or DeepSeek V4 Pro read screenshots?

No. Both are text-only on their vendors' APIs. Kimi K3, Kimi K2.7 Code, GLM-5.3-Flash and DeepSeek V4.1 Flash accept images.

Do I need a router to use these models with Claude Code?

Not for a single vendor. A router earns its place when you want different vendors in different slots, fallback when a plan window runs out, or OpenCode Go's Kimi, GLM and DeepSeek models, which Claude Code cannot reach directly.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter