Skip to main content
kRouter
All posts
Fix an error

Claude Code context window with non-Claude models: the fix

Claude Code assumes 200K for a model it doesn't recognize, so it compacts too early or overruns the real limit. How to declare the real window.

Kodelyth · The team behind kRouter
· Updated
9 min read

You point Claude Code at DeepSeek, Kimi or GLM through a gateway, and it manages context as if it were running some other model. It goes wrong in one of two directions:

  • It compacts far too early. DeepSeek's own API gives deepseek-v4-pro a 1M-token window. Claude Code summarizes your session and drops detail before it reaches 200K, less than a fifth of the way in.
  • It does not compact in time. The model's real window is smaller than Claude Code believes, the provider rejects the request as too long, and the session usually stops on an error instead of compacting.

Neither the model nor the gateway is broken. Claude Code needs a model's window to decide when to compact, and for a model it does not recognize, it guesses.

How Claude Code decides the window

Behind a custom ANTHROPIC_BASE_URL, Claude Code gives each Claude model it recognizes the window that model has on Anthropic's API. For an ID it does not recognize, Anthropic's gateway reference is specific: it assumes 200K, or 1M when the ID contains [1m], and compacts at that size. kRouter's provider/model IDs land in different buckets:

Model ID Claude Code sendsHow Claude Code reads itHow to correct the window
cc/claude-opus-4-8Contains a Claude model it knows, so it uses that model's windowOnly lower it, with CLAUDE_CODE_AUTO_COMPACT_WINDOW
ds/deepseek-v4-pro, kimi/kimi-k2.6, ocg/glm-5.3Unrecognized, so 200KCLAUDE_CODE_MAX_CONTEXT_TOKENS
A combo named long-contextUnrecognized, so 200KCLAUDE_CODE_MAX_CONTEXT_TOKENS
A combo named claude-cheapA bare claude- nameRename the combo
An unrecognized ID containing [1m]Assumes 1MThe variable plus CLAUDE_CODE_DISABLE_1M_CONTEXT=1

The fourth row catches people out: if a combo of non-Claude models has a name starting with claude-, the override below is ignored unless you switch off compaction entirely.

Check what Claude Code thinks

  • Run /autocompact with no argument (Claude Code 2.1.221 or later). It opens a dialog showing the current auto-compact window.
  • Look for the unrecognized-model line (2.1.233 or later). Start with claude --debug and search ~/.claude/debug/<session-id>.txt for [claude-code:unrecognized_model]; in -p mode it goes to stderr. There is one line per unrecognized ID, including IDs only a subagent or background task uses, which settles spellings you cannot predict, such as Kiro's dotted kr/claude-sonnet-4.5.

The fix: declare the real window

Set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the window Claude Code should assume (Anthropic's reference). The env block of ~/.claude/settings.json is the dependable place: it applies every time claude runs, and in most sessions it replaces the same variable exported in your shell.

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:20128/v1",
    "ANTHROPIC_AUTH_TOKEN": "<your-krouter-key>",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "ds/deepseek-v4-pro",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "kimi/kimi-k2.6",
    "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "260000"
  }
}

Since Claude Code v2.1.193, how the variable applies depends on how Claude Code reads the model ID. Three details decide whether it works:

  • It is one number, not one per model. It applies to whichever untagged, unrecognized model is active. DeepSeek V4 Pro has 1M on DeepSeek's API, but Kimi K2.6 is listed at 262,144 tokens on Kimi's, so the safe value sits just under the smaller one. The same goes for a combo: any entry can answer, so size for its smallest window.
  • It only touches unrecognized IDs. A slot on a real Claude model, such as the dashboard's default Haiku slot cc/claude-haiku-4-5-20251001, keeps Claude's window.
  • Write plain integers. Anthropic documents the trap for the sibling variable CLAUDE_CODE_AUTO_COMPACT_WINDOW, where 500k reads as 500 and clamps to the 100K minimum. Write 260000, not 260k, here too.

Then check the output reservation. For an ID it cannot resolve, Claude Code defaults CLAUDE_CODE_MAX_OUTPUT_TOKENS to 32,000 (cap 128,000), and the reservation comes out of the window, so raising it means earlier compaction.

Use the window of the route you actually use: a reseller or coding plan can serve less than a model's maximum, and a local server runs whatever context length it was started with.

The shortcut for 1M models: the [1m] tag

If the model really has about 1M on your route, you can tag the ID instead: type ds/deepseek-v4-pro[1m] into the slot. Claude Code assumes 1M for an unrecognized ID containing [1m], and kRouter strips a trailing [1m] before it resolves the model (v0.5.150 and later), so the provider sees the plain ID. DeepSeek's own Claude Code guide tags its model IDs the same way (deepseek-flash[1m]) and adds CLAUDE_CODE_AUTO_COMPACT_WINDOW=786432, so sessions compact well before the window runs out.

The tag is what makes a mixed setup work. CLAUDE_CODE_MAX_CONTEXT_TOKENS does not apply to a tagged ID on its own, so ds/deepseek-v4-pro[1m] in the Opus slot keeps 1M while 260000 sizes the untagged kimi/kimi-k2.6 in the Sonnet slot.

The tag claims exactly 1M, so it is wrong for Kimi K2.6 or anything smaller. To override the window on a tagged ID, you also need CLAUDE_CODE_DISABLE_1M_CONTEXT=1, which has two side effects: real Claude models with a native 1M window, such as Opus 4.8 in a cc/ slot, are held to 200K, and a declared window above 200K brings a startup warning that the 200K limit isn't enforced. Anthropic documents that warning as expected in this setup.

Setting it from the kRouter dashboard

kRouter writes the same file for you. Open http://localhost:20128/dashboard, go to CLI Tools and click the Claude Code card. Fill the Opus, Sonnet and Haiku slots, pick a Max context value, and click Apply. The control has been there since v0.5.132.

It offers Default, 200K, 300K, 500K and 1M, and writes values 2K under each label: 198000, 298000, 498000 and 998000. Default writes nothing. Two details:

  • Only those five sizes. For Kimi K2.6's 262,144 tokens, 300K is too big. Pick 200K, or edit the number in the file yourself.
  • Clearing it is manual. Apply merges into your existing settings, so choosing Default later does not delete a saved value. In v0.5.163 and earlier, Reset removes the base URL, token and model slots but leaves CLAUDE_CODE_MAX_CONTEXT_TOKENS behind. Delete that line by hand.

To look a window up, curl -s http://localhost:20128/v1/models shows context_length for each connected chat model (since v0.5.138). Treat it as a hint: it comes from kRouter's built-in table, which can lag behind new releases, and an unknown model shows 200K. Claude Code does not read it.

When the provider rejects the request anyway

When the API rejects a request as too long, Claude Code compacts and retries on its own, but only for wording it recognizes: Anthropic's Prompt is too long, Amazon Bedrock's Input is too long for requested model., and a Claude apps gateway's capability_rejected: prompt_too_long.

Other providers say it their own way. Kiro answers Input content length exceeds threshold., and kRouter passes the provider's message back with the HTTP status in front, not rewritten into Anthropic's phrasing. As Anthropic's gateway troubleshooting explains, Claude Code then neither compacts nor retries. Run /compact to recover, then fix the window.

kRouter does nothing clever here, on purpose. Since v0.5.94, a too-long error it recognizes, such as Kiro's above or OpenAI's maximum context length, comes straight back: no other account is tried, since each would reject the same request, and the account gets no cooldown. The same rule stops a combo, even when its next model has a bigger window.

A route serving a Claude model can also enforce less than Anthropic's API, and Claude Code cannot see that limit. If yours rejects requests above 200K, set CLAUDE_CODE_AUTO_COMPACT_WINDOW=200000.

What does not work

Turning compaction off. DISABLE_AUTO_COMPACT=1 turns automatic compaction off. It does not tell Claude Code the real window, and nothing compacts on its own any more: Anthropic's advice for that setup is to run /compact yourself before the window fills. DISABLE_COMPACT=1 removes /compact too, leaving only /clear.

Overriding a Claude-looking ID. For an ID that resolves to a Claude model, or a bare claude- name, the window variable only works with DISABLE_COMPACT set. Rename the combo instead.

Silencing the warning with modelOverrides. Anthropic's error reference suggests it to stop the unrecognized-model line. It works by making Claude Code treat your ID as that Claude model, so the window variable stops applying.

Raising the window with /autocompact. It moves the compaction point inside the model's window and is capped at it. On an unrecognized ID, that cap is the window Claude Code assumes: 200K unless you declare another.

Compacting only after a rejection. CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1 (2.1.223 or later) waits for a too-long error. Through kRouter, that error carries the provider's own wording, which Claude Code usually does not recognize, so nothing compacts and the session stops on the error.

A bigger window is not free

Claude Code sends the whole conversation with every request, so a 1M window means longer, costlier requests; prompt caching softens that where the provider supports it. To compact earlier, set CLAUDE_CODE_AUTO_COMPACT_WINDOW=400000. kRouter's RTK, the only token saver on by default, compresses tool output before it reaches the provider. Headroom compresses more of the context, but it is off by default and needs its own Python proxy.

Working out which case you are in

What you seeWhat is happeningDo this
Compaction just under 200K on a model with a much bigger windowUnrecognized ID, 200K assumedSet CLAUDE_CODE_MAX_CONTEXT_TOKENS, or [1m] for a true 1M model
The provider rejects the request as too long and nothing compactsReal window smaller than assumed, wording not recognized/compact now, then set the variable
You set the variable and nothing changedClaude-looking ID, a [1m] tag, a modelOverrides entry, or a project settings file with its own valueRename, untag or remove the override
A Claude model through another route rejects requests above 200KThe route enforces a lower limitCLAUDE_CODE_AUTO_COMPACT_WINDOW=200000
/compact answers Not enough messages to compact.One exchange is too big on its own/clear and resend with less pasted content

Where kRouter is not the answer

The fix is a Claude Code setting; kRouter writes it but cannot change what Claude Code assumes. If you use one provider that speaks Anthropic's API itself, such as DeepSeek at https://api.deepseek.com/anthropic, you can skip the router; the same window rules apply.

Anthropic also says it does not support routing Claude Code to non-Claude models through any gateway, and window handling is where that shows. Which open model belongs in which slot is a separate question; this comparison covers it.

Common questions

What context window does Claude Code assume for a non-Claude model?

For a model ID it does not recognize, 200K tokens, or 1M if the ID contains [1m], whatever the model's real window is. An ID containing a Claude model name it knows, such as cc/claude-opus-4-8, gets that Claude model's window instead.

Why is CLAUDE_CODE_MAX_CONTEXT_TOKENS being ignored?

Usually because Claude Code treats the ID as a Claude model: it starts with claude-, contains a Claude model name, or a modelOverrides entry maps it to one. Those IDs only take the variable with all compaction disabled. Other causes are a [1m] tag, which also needs CLAUDE_CODE_DISABLE_1M_CONTEXT=1, and a project or managed settings file that sets its own value.

Does Claude Code read context_length from my gateway's /v1/models?

No. Its gateway model discovery is off by default, reads only id, display_name and description, and keeps only IDs containing "claude" or "anthropic". kRouter publishes context_length for other clients and for you to look up.

Will kRouter switch to a bigger-context model when a prompt is too long?

No. When it recognizes the error as too long, it returns it straight away, without trying other accounts or the next model in a combo. The fix is on the Claude Code side: a correct window, or /compact.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter