Claude Code context window with non-Claude models: the fix
Claude Code assumes 200K for a model it doesn't recognize, so it compacts too early or overruns the real limit. How to declare the real window.
You point Claude Code at DeepSeek, Kimi or GLM through a gateway, and it manages context as if it were running some other model. It goes wrong in one of two directions:
- It compacts far too early. DeepSeek's own API gives
deepseek-v4-proa 1M-token window. Claude Code summarizes your session and drops detail before it reaches 200K, less than a fifth of the way in. - It does not compact in time. The model's real window is smaller than Claude Code believes, the provider rejects the request as too long, and the session usually stops on an error instead of compacting.
Neither the model nor the gateway is broken. Claude Code needs a model's window to decide when to compact, and for a model it does not recognize, it guesses.
How Claude Code decides the window
Behind a custom ANTHROPIC_BASE_URL, Claude Code gives each Claude model it recognizes the window that model has on Anthropic's API. For an ID it does not recognize, Anthropic's gateway reference is specific: it assumes 200K, or 1M when the ID contains [1m], and compacts at that size. kRouter's provider/model IDs land in different buckets:
| Model ID Claude Code sends | How Claude Code reads it | How to correct the window |
|---|---|---|
cc/claude-opus-4-8 | Contains a Claude model it knows, so it uses that model's window | Only lower it, with CLAUDE_CODE_AUTO_COMPACT_WINDOW |
ds/deepseek-v4-pro, kimi/kimi-k2.6, ocg/glm-5.3 | Unrecognized, so 200K | CLAUDE_CODE_MAX_CONTEXT_TOKENS |
A combo named long-context | Unrecognized, so 200K | CLAUDE_CODE_MAX_CONTEXT_TOKENS |
A combo named claude-cheap | A bare claude- name | Rename the combo |
An unrecognized ID containing [1m] | Assumes 1M | The variable plus CLAUDE_CODE_DISABLE_1M_CONTEXT=1 |
The fourth row catches people out: if a combo of non-Claude models has a name starting with claude-, the override below is ignored unless you switch off compaction entirely.
Check what Claude Code thinks
- Run
/autocompactwith no argument (Claude Code 2.1.221 or later). It opens a dialog showing the current auto-compact window. - Look for the unrecognized-model line (2.1.233 or later). Start with
claude --debugand search~/.claude/debug/<session-id>.txtfor[claude-code:unrecognized_model]; in-pmode it goes to stderr. There is one line per unrecognized ID, including IDs only a subagent or background task uses, which settles spellings you cannot predict, such as Kiro's dottedkr/claude-sonnet-4.5.
The fix: declare the real window
Set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the window Claude Code should assume (Anthropic's reference). The env block of ~/.claude/settings.json is the dependable place: it applies every time claude runs, and in most sessions it replaces the same variable exported in your shell.
{
"env": {
"ANTHROPIC_BASE_URL": "http://localhost:20128/v1",
"ANTHROPIC_AUTH_TOKEN": "<your-krouter-key>",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "ds/deepseek-v4-pro",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "kimi/kimi-k2.6",
"CLAUDE_CODE_MAX_CONTEXT_TOKENS": "260000"
}
}Since Claude Code v2.1.193, how the variable applies depends on how Claude Code reads the model ID. Three details decide whether it works:
- It is one number, not one per model. It applies to whichever untagged, unrecognized model is active. DeepSeek V4 Pro has 1M on DeepSeek's API, but Kimi K2.6 is listed at 262,144 tokens on Kimi's, so the safe value sits just under the smaller one. The same goes for a combo: any entry can answer, so size for its smallest window.
- It only touches unrecognized IDs. A slot on a real Claude model, such as the dashboard's default Haiku slot
cc/claude-haiku-4-5-20251001, keeps Claude's window. - Write plain integers. Anthropic documents the trap for the sibling variable
CLAUDE_CODE_AUTO_COMPACT_WINDOW, where500kreads as 500 and clamps to the 100K minimum. Write260000, not260k, here too.
Then check the output reservation. For an ID it cannot resolve, Claude Code defaults CLAUDE_CODE_MAX_OUTPUT_TOKENS to 32,000 (cap 128,000), and the reservation comes out of the window, so raising it means earlier compaction.
Use the window of the route you actually use: a reseller or coding plan can serve less than a model's maximum, and a local server runs whatever context length it was started with.
The shortcut for 1M models: the [1m] tag
If the model really has about 1M on your route, you can tag the ID instead: type ds/deepseek-v4-pro[1m] into the slot. Claude Code assumes 1M for an unrecognized ID containing [1m], and kRouter strips a trailing [1m] before it resolves the model (v0.5.150 and later), so the provider sees the plain ID. DeepSeek's own Claude Code guide tags its model IDs the same way (deepseek-flash[1m]) and adds CLAUDE_CODE_AUTO_COMPACT_WINDOW=786432, so sessions compact well before the window runs out.
The tag is what makes a mixed setup work. CLAUDE_CODE_MAX_CONTEXT_TOKENS does not apply to a tagged ID on its own, so ds/deepseek-v4-pro[1m] in the Opus slot keeps 1M while 260000 sizes the untagged kimi/kimi-k2.6 in the Sonnet slot.
The tag claims exactly 1M, so it is wrong for Kimi K2.6 or anything smaller. To override the window on a tagged ID, you also need CLAUDE_CODE_DISABLE_1M_CONTEXT=1, which has two side effects: real Claude models with a native 1M window, such as Opus 4.8 in a cc/ slot, are held to 200K, and a declared window above 200K brings a startup warning that the 200K limit isn't enforced. Anthropic documents that warning as expected in this setup.
Setting it from the kRouter dashboard
kRouter writes the same file for you. Open http://localhost:20128/dashboard, go to CLI Tools and click the Claude Code card. Fill the Opus, Sonnet and Haiku slots, pick a Max context value, and click Apply. The control has been there since v0.5.132.
It offers Default, 200K, 300K, 500K and 1M, and writes values 2K under each label: 198000, 298000, 498000 and 998000. Default writes nothing. Two details:
- Only those five sizes. For Kimi K2.6's 262,144 tokens, 300K is too big. Pick 200K, or edit the number in the file yourself.
- Clearing it is manual. Apply merges into your existing settings, so choosing Default later does not delete a saved value. In v0.5.163 and earlier, Reset removes the base URL, token and model slots but leaves
CLAUDE_CODE_MAX_CONTEXT_TOKENSbehind. Delete that line by hand.
To look a window up, curl -s http://localhost:20128/v1/models shows context_length for each connected chat model (since v0.5.138). Treat it as a hint: it comes from kRouter's built-in table, which can lag behind new releases, and an unknown model shows 200K. Claude Code does not read it.
When the provider rejects the request anyway
When the API rejects a request as too long, Claude Code compacts and retries on its own, but only for wording it recognizes: Anthropic's Prompt is too long, Amazon Bedrock's Input is too long for requested model., and a Claude apps gateway's capability_rejected: prompt_too_long.
Other providers say it their own way. Kiro answers Input content length exceeds threshold., and kRouter passes the provider's message back with the HTTP status in front, not rewritten into Anthropic's phrasing. As Anthropic's gateway troubleshooting explains, Claude Code then neither compacts nor retries. Run /compact to recover, then fix the window.
kRouter does nothing clever here, on purpose. Since v0.5.94, a too-long error it recognizes, such as Kiro's above or OpenAI's maximum context length, comes straight back: no other account is tried, since each would reject the same request, and the account gets no cooldown. The same rule stops a combo, even when its next model has a bigger window.
A route serving a Claude model can also enforce less than Anthropic's API, and Claude Code cannot see that limit. If yours rejects requests above 200K, set CLAUDE_CODE_AUTO_COMPACT_WINDOW=200000.
What does not work
Turning compaction off. DISABLE_AUTO_COMPACT=1 turns automatic compaction off. It does not tell Claude Code the real window, and nothing compacts on its own any more: Anthropic's advice for that setup is to run /compact yourself before the window fills. DISABLE_COMPACT=1 removes /compact too, leaving only /clear.
Overriding a Claude-looking ID. For an ID that resolves to a Claude model, or a bare claude- name, the window variable only works with DISABLE_COMPACT set. Rename the combo instead.
Silencing the warning with modelOverrides. Anthropic's error reference suggests it to stop the unrecognized-model line. It works by making Claude Code treat your ID as that Claude model, so the window variable stops applying.
Raising the window with /autocompact. It moves the compaction point inside the model's window and is capped at it. On an unrecognized ID, that cap is the window Claude Code assumes: 200K unless you declare another.
Compacting only after a rejection. CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1 (2.1.223 or later) waits for a too-long error. Through kRouter, that error carries the provider's own wording, which Claude Code usually does not recognize, so nothing compacts and the session stops on the error.
A bigger window is not free
Claude Code sends the whole conversation with every request, so a 1M window means longer, costlier requests; prompt caching softens that where the provider supports it. To compact earlier, set CLAUDE_CODE_AUTO_COMPACT_WINDOW=400000. kRouter's RTK, the only token saver on by default, compresses tool output before it reaches the provider. Headroom compresses more of the context, but it is off by default and needs its own Python proxy.
Working out which case you are in
| What you see | What is happening | Do this |
|---|---|---|
| Compaction just under 200K on a model with a much bigger window | Unrecognized ID, 200K assumed | Set CLAUDE_CODE_MAX_CONTEXT_TOKENS, or [1m] for a true 1M model |
| The provider rejects the request as too long and nothing compacts | Real window smaller than assumed, wording not recognized | /compact now, then set the variable |
| You set the variable and nothing changed | Claude-looking ID, a [1m] tag, a modelOverrides entry, or a project settings file with its own value | Rename, untag or remove the override |
| A Claude model through another route rejects requests above 200K | The route enforces a lower limit | CLAUDE_CODE_AUTO_COMPACT_WINDOW=200000 |
/compact answers Not enough messages to compact. | One exchange is too big on its own | /clear and resend with less pasted content |
Where kRouter is not the answer
The fix is a Claude Code setting; kRouter writes it but cannot change what Claude Code assumes. If you use one provider that speaks Anthropic's API itself, such as DeepSeek at https://api.deepseek.com/anthropic, you can skip the router; the same window rules apply.
Anthropic also says it does not support routing Claude Code to non-Claude models through any gateway, and window handling is where that shows. Which open model belongs in which slot is a separate question; this comparison covers it.
Common questions
What context window does Claude Code assume for a non-Claude model?
For a model ID it does not recognize, 200K tokens, or 1M if the ID contains [1m], whatever the model's real window is. An ID containing a Claude model name it knows, such as cc/claude-opus-4-8, gets that Claude model's window instead.
Why is CLAUDE_CODE_MAX_CONTEXT_TOKENS being ignored?
Usually because Claude Code treats the ID as a Claude model: it starts with claude-, contains a Claude model name, or a modelOverrides entry maps it to one. Those IDs only take the variable with all compaction disabled. Other causes are a [1m] tag, which also needs CLAUDE_CODE_DISABLE_1M_CONTEXT=1, and a project or managed settings file that sets its own value.
Does Claude Code read context_length from my gateway's /v1/models?
No. Its gateway model discovery is off by default, reads only id, display_name and description, and keeps only IDs containing "claude" or "anthropic". kRouter publishes context_length for other clients and for you to look up.
Will kRouter switch to a bigger-context model when a prompt is too long?
No. When it recognizes the error as too long, it returns it straight away, without trying other accounts or the next model in a combo. The fix is on the Claude Code side: a correct window, or /compact.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- Fix an errorClaude Code weekly limit reached: keep working until resetClaude Code says you've hit your weekly limit. What the caps count, what does not help, and three ways to keep working until the reset.
- Fix an errorScreenshots and text-only models like GLM-5.3: what worksGLM-5.3 can't read a pasted screenshot. Three fixes -- a model that reads images, a vision MCP server, a per-turn switch -- and where each falls short.
- Fix an errorClaude Code "rate limit exceeded": every cause and the fixClaude Code stops with a 429 or a usage limit. Four different limits cause it, each needs a different fix, and the one everyone tries first makes it worse.
Relevant docs
- TroubleshootingFixes for common kRouter problems: banned or rate-limited accounts, OAuth sign-ins, MITM certificates, localhost and Cursor, Docker logins and the CLI.
- Error referenceWhat the common AI provider errors mean, why they happen, and how to get past them.
- Combos & fallbackPut several models behind one name. kRouter falls back from one to the next, rotates them, or asks a panel of models and merges the answers.