Skip to main content
kRouter
All posts
Save money

Qwen Code's free tier is gone: keep the CLI, swap the model

Qwen OAuth's free tier closed on April 15, 2026. What still works in Qwen Code, where Qwen models come from now, and how to keep the CLI answering.

Kodelyth · The team behind kRouter
· Updated
8 min read

You start qwen the way you have for months, and instead of an answer you get this:

Qwen OAuth free tier was discontinued on 2026-04-15. Run /auth to switch to Coding Plan, OpenRouter, Fireworks AI, or another provider.

Nothing is wrong with your install. Alibaba closed the free sign-in that came with Qwen Code. A cached token may work for a short while, but new requests are rejected.

The CLI itself is fine: still open source under Apache 2.0, still installed with npm install -g @qwen-code/qwen-code. What ended is the free model access behind the qwen.ai login.

What ended on April 15

The announcement is issue #3203, opened April 13, 2026. It cut the daily free quota from 1,000 requests to 100 immediately and closed the Qwen OAuth free entry point on April 15. The stated reason was "a product strategy adjustment to better manage the free tier usage and costs."

Qwen OAuth is no longer an option in /auth, and the qwen auth command is gone (auth docs). What /auth does offer:

  • Alibaba ModelStudio: Coding Plan, Token Plan or a standard API key.
  • Third-party providers: DeepSeek, Grok, MiniMax, Z.AI, Kimi, Idealab, ModelScope, OpenRouter and Requesty.
  • Custom provider: any OpenAI, Anthropic or Gemini-compatible endpoint, including a local server.

What does not work

Signing in again, or with another qwen.ai account. The sign-in no longer grants access, whichever account you use.

Pinning an older Qwen Code. The cutoff is on Alibaba's side, and an older client calls the same backend.

OpenRouter's free Qwen Coder. Older guides point at qwen/qwen3-coder:free. OpenRouter no longer lists it and no provider serves it, and none of the free models on its list in early October 2026 is a Qwen model.

OpenCode's free models. None of them are Qwen, and requests from outside OpenCode now get OpenCode's free tier can only be used from within OpenCode. That includes kRouter's OpenCode Free provider.

A proxy that reuses the Qwen login. Same backend, same refusal. kRouter hid its own Qwen OAuth provider in v0.4.55 (May 2026) for that reason.

Where Qwen models come from now

The Model Studio new-user quota

This is the closest thing to free Qwen that Alibaba still offers, and it is a trial rather than a tier. New Model Studio users in the Singapore region get a free quota per model, typically 1 million tokens, valid for 90 days (free quota). Quotas cannot be pooled across models, and re-registering does not grant another.

Turn on Free Quota Only for each model you use. It is off by default, and without it calls carry on at pay-as-you-go rates once the quota is gone; with it on, they stop with AllocationQuota.FreeTierOnly. In Qwen Code, pick Alibaba ModelStudio → Standard API Key.

Coding Plan

Alibaba's Coding Plan is the flat fee the error message suggests. As of early October 2026:

  • Lite is closed. New subscriptions stopped on March 20, 2026, and renewals on April 13.
  • Pro is $50 a month, sold in limited slots restocked daily. It allows up to 6,000 requests per 5 hours, 45,000 per week and 90,000 per month, and whichever you reach first pauses calls.
  • It counts model calls, not tokens. Alibaba estimates 5 to 10 calls for a simple task and 10 to 30 or more for a complex one.
  • Only listed models work: Qwen 3.7 Plus, 3.6 Plus, 3.5 Plus, Qwen3 Max, Qwen3 Coder Plus and Coder Next, plus GLM-5, GLM-4.7, Kimi K2.5 and MiniMax M2.5.
  • Interactive use only. The terms say: "Do not use the plan's API key for automated scripts, application backends, or other non-interactive scenarios."

Coding Plan keys start with sk-sp-, and Alibaba says they are not interchangeable with standard Model Studio keys and base URLs.

Token Plan

The Token Plan page notes that Coding Plan Pro is no longer available once sold out. The Token Plan is metered in credits. On Alibaba's international site it is sold for the Singapore region, and personal plans start at $8 a month at list price. Qwen Code is on its supported-tools list.

Elsewhere

  • OpenCode Go ($10 a month, $40 for Go Plus) serves Qwen 3.8 Max, 3.8 Flash and 3.7 Plus, but only on its Anthropic Messages endpoint (Go docs).
  • OpenRouter sells Qwen models per token.
  • Kiro's free plan (50 credits a month) includes Qwen3 Coder Next among its rate-limited open-weight models. Kiro speaks its own format, so it reaches Qwen Code only through a translator.
  • A local model through Ollama or vLLM costs nothing per token if you have the hardware.

The structural fix: one endpoint, several sources

If one of these covers you, set it up in /auth and stop there. Qwen Code talks to most of them directly, and you do not need a router.

A router earns its place when you want more than one source behind a single model name: the Coding Plan's 5-hour cap pauses you mid-task, a free quota runs out, or you want OpenCode Go's Qwen models alongside the rest.

Install kRouter and start it in the background:

npm install -g @sifxprime/krouter
krouter -t

Then, in the dashboard at http://localhost:20128/dashboard:

  1. Providers: connect what you have. A Coding Plan key goes under Alibaba Intl (or Alibaba for Beijing), a Token Plan key under Alibaba Token Plan, then OpenCode Go or OpenRouter. If a model you are entitled to is missing, such as qwen3.7-plus on Alibaba Intl, add it with Add Model.
  2. Combos → Create Combo, named qwen-chain, with models in order of preference:
Combo "qwen-chain" (fallback)
  1. alicode-intl/qwen3-coder-plus           Coding Plan key
  2. ocg/qwen3.7-plus                        OpenCode Go
  3. openrouter/qwen/qwen3-coder-next        OpenRouter, paid per token
  1. Endpoint: copy an API key. A Default Key is created on your first visit.

Then point Qwen Code at kRouter in ~/.qwen/settings.json:

{
  "env": { "KROUTER_API_KEY": "<key from the Endpoint page>" },
  "modelProviders": {
    "openai": [
      {
        "id": "qwen-chain",
        "name": "Qwen chain (kRouter)",
        "envKey": "KROUTER_API_KEY",
        "baseUrl": "http://localhost:20128/v1",
        "generationConfig": { "contextWindowSize": 262144 }
      },
      {
        "id": "ocg/qwen3.7-plus",
        "name": "Qwen 3.7 Plus via OpenCode Go",
        "envKey": "KROUTER_API_KEY",
        "baseUrl": "http://localhost:20128/v1"
      }
    ]
  },
  "security": { "auth": { "selectedType": "openai" } },
  "model": { "name": "qwen-chain" }
}

Both entries appear in /model. Four details matter:

  • Set contextWindowSize on the combo entry. Qwen Code sizes the window from the model name, and the window decides when it compacts the conversation. A combo name says nothing about the models behind it: qwen-chain happens to match Qwen Code's catch-all Qwen rule, and a name it does not recognise gets its 200,000-token default. Use the smallest window in the combo: OpenRouter lists 262,144 for Qwen3 Coder Next, and if another provider documents a lower limit, use that.
  • The key sits in the env block on purpose. Qwen Code reads only the first .env file it finds, searching up from the current directory before it tries your home directory, so a key in ~/.qwen/.env is skipped in any repo with its own .env. The env block fills any variable not set elsewhere. Keep the file out of published dotfiles.
  • The base URL is the /v1 root, without /chat/completions.
  • kRouter's Qwen Code card only shows a guide. Its JSON block (CLI Tools → Qwen Code) uses the older security.auth.apiKey and baseUrl fields, which Qwen Code still reads but has deprecated. Prefer the modelProviders form.

When an entry fails with a rate limit, quota or server error, the request moves to the next one. The combo does not add quota: when every source is spent, requests fail until one resets.

Kiro's free plan can join the combo as kr/qwen3-coder-next, with a caveat: kRouter shows a risk notice before you connect it, because Kiro's sign-in is not licensed for router use and the account can be restricted.

RTK, the one kRouter token saver on by default, compresses tool output such as diffs, grep results and file reads before it is sent. That helps on per-token billing. It does nothing for the Coding Plan, which counts calls.

OpenCode Go's Qwen models

Go serves its Qwen models only on /zen/go/v1/messages, the Anthropic format, and asks every client to send a stable x-opencode-session header per conversation. Qwen Code is not on its list of validated clients.

The direct route is Qwen Code's own Anthropic provider, with "x-opencode-session": "${session_id}" in the model's generationConfig.customHeaders. Qwen Code's docs give that placeholder for OpenCode Go, and it expands to the current session id on every request. Test it before you rely on it, since OpenCode has not validated Qwen Code.

Through kRouter, the OpenAI-style entry above works as it is:

  • kRouter translates. kRouter 0.5.164 and later send every Go Qwen model to the Messages endpoint. On 0.5.163, set the model's Protocol to Messages on the OpenCode Go provider page.
  • kRouter adds the session header, taken from a conversation id in the request body when there is one, otherwise a stable id per connection.

Go's limits are dollar amounts per model, with 20% of the monthly figure available per 5 hours and 50% per week. On the $10 plan, OpenCode lists Qwen 3.8 Max at $15 a month, 3.8 Flash at $30 and 3.7 Plus at $60, so Max gets the fewest requests. The OpenCode Go protocols post covers the rest of its catalog.

Choosing

Your situationWhat to do
You want Qwen models free for a whileModel Studio new-user quota, with Free Quota Only on
You code on Qwen daily and want a flat feeCoding Plan Pro if a slot is open, otherwise the Token Plan
You already pay for OpenCode GoIts Qwen models through kRouter, or a direct Anthropic entry with the session header added by hand
You have a capable GPUA local Qwen model through Ollama or vLLM
One source covers your usageConfigure it in /auth, no router
A 5-hour cap keeps stopping you mid-taskA kRouter combo with a second source behind the first

Common questions

Is Qwen Code still free?

The CLI is: it is open source under Apache 2.0 and installs from npm. The free model access that came with the qwen.ai sign-in is not. Alibaba cut it to 100 requests a day on April 13, 2026, and closed it on April 15.

Can I get the free tier back by signing in again or downgrading?

No. The cutoff is server-side, so a fresh sign-in, a second account or an older version all reach the same closed endpoint. Run /auth and pick another source.

Is there any free way left to use Qwen models in Qwen Code?

Three, each limited. The Model Studio new-user quota lasts 90 days. Kiro's free plan includes Qwen3 Coder Next with rate limits, and using it through a router risks the account. A local open-weight model costs nothing per token but needs the hardware. OpenRouter had no free Qwen model in early October 2026.

Can I use my Alibaba Coding Plan key through kRouter?

kRouter has the providers: Alibaba Intl for the international endpoint, Alibaba for Beijing. The plan's terms limit the key to the subscriber's interactive use in coding tools and do not mention routers. Sharing the endpoint with other people, or pointing a script at it, is clearly outside them. kRouter listens on every network interface by default, so start it with krouter -t --host 127.0.0.1 to keep it on your own machine.

Does kRouter bring back the Qwen OAuth free tier?

No. Only Alibaba decides who it serves, and kRouter hid its own Qwen OAuth provider in v0.4.55 after the sign-in stopped working. What kRouter can do is put the sources that still work behind one model name and move to the next when one runs out.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter