Skip to main content
kRouter
All posts
Save money

GLM, MiniMax and Alibaba coding plans: stack two of them

What the GLM, MiniMax and Alibaba coding plans meter, where their 5-hour and weekly windows bite, and how to fall back from one plan to the next.

Kodelyth · The team behind kRouter
· Updated
9 min read

Z.ai documents the reply you get when a GLM Coding Plan window runs out. It comes back as HTTP 429:

1308  Usage limit reached for {number} {unit}. Your limit will reset at {next_flush_time}

The weekly (or monthly) version is code 1310. Z.ai's FAQ says calls from supported tools never fall back to your account balance, so work simply stops until the reset. The usual answer is a second plan from another vendor. That works, but the plans count different things, their keys and endpoints are separate from pay-as-you-go, and two of the vendors restrict where the plan may be used.

This post covers three plan families that kRouter has dedicated providers for: Z.ai's GLM Coding Plan, MiniMax's M Plan (the successor to its Token Plan), and Alibaba's Coding Plan and Token Plan. Prices and limits come from each vendor's documentation in October 2026. All three changed their plans this year, so check again before you subscribe.

What each plan actually counts

GLM Coding Plan. Since July 30, 2026, usage is counted in credits worked out from tokens: input, cached input and output tokens, each times a multiplier, divided by 10,000. For GLM-5.3 the multipliers are 6.9, 1.7 and 24; GLM-5.3-Flash costs about a third of that. Lite gets 2,000 credits per 5 hours and 10,000 per week, Pro 12,000 and 60,000, Max 28,000 and 140,000. Prices start at $18 a month. Weekday use from 14:00 to 18:00 Singapore time is charged at the full rate, and any other time at half (Z.ai's plan page). Plans bought before July 30 keep their old limits until the end of the current billing cycle (Z.ai's notice).

To make a credit concrete, take one request late in a session: 60,000 tokens of conversation the vendor already has cached, 4,000 new input tokens and 1,000 output tokens. On GLM-5.3 at the full rate that is 15.4 credits, so Lite's 5-hour window holds about 130 such requests, or 260 off-peak. With nothing cached, the same request costs 46.6. An agent sends a request for every tool round-trip, which is why the window empties faster than your prompt count suggests.

MiniMax M Plan. MiniMax announced on September 29, 2026 that its Token Plan is closed to new purchases; subscribers who keep auto-renew on stay on it (announcement). The M Plan that replaces it keeps the same prices and text quotas: Go is $22 a month, Explore $55 and Build $132, with Explore giving three times Go's usage and Build 7.5 times (M Plan overview). How much a request uses depends on the model, the context length and the task. There is a 5-hour window and a weekly one, each starting at your first request. Text models need room in both, and the two reset independently. MiniMax warns that peak-hour limits can block calls before either window is used up, and purchased Credits cover usage past the plan (usage rules). The text model listed for every tier is M3.1 Flash Preview.

Alibaba Coding Plan. Counted in model calls. Pro allows up to 6,000 per 5 hours, 45,000 per week and 90,000 per month for $50 a month, and whichever cap you reach first pauses you. Pro is sold in limited slots that Alibaba restocks daily. Lite closed to new subscribers on March 20, 2026 and to renewals on April 13 (Coding Plan overview).

Alibaba Token Plan. What Alibaba now recommends instead. It is a separate product, so a Coding Plan cannot be moved into it. On Alibaba Cloud's international site it is sold in the Singapore region, metered in Credits per 30-day cycle with no 5-hour or weekly window. Personal tiers list at $8 to $80 a month and are currently discounted to $6 to $68. When the Credits run out, the service pauses until the next cycle unless you buy an Extra Bundle; nothing is billed pay-as-you-go (Token Plan overview).

PlanWhat it countsWindowsPer month
GLM Coding PlanCredits from tokens; off-peak costs half5 hours and weeklyFrom $18
MiniMax M PlanUsage by model, context and task5 hours and weekly$22, $55 or $132
Alibaba Coding Plan ProModel calls5 hours, weekly and monthly$50, limited daily slots
Alibaba Token PlanCredits30 days$6 to $68 (discounted)

Plan keys only work on plan endpoints

Each plan is tied to its own endpoint, its own key, or both. Get either wrong and the request misses the plan:

  • Z.ai applies the plan only on three endpoints: https://api.z.ai/api/anthropic (Anthropic Messages, the one for Claude Code), https://api.z.ai/api/coding/paas/v4 (Chat Completions) and https://api.z.ai/api/v1 (Responses) (endpoint list). It covers only GLM-5.3 and GLM-5.3-Flash; requests for GLM-5.2 or GLM-5.1 are served by GLM-5.3, and GLM-4.7 by GLM-5.3-Flash. Miss the endpoint and you get 1113 Insufficient balance, or a charge on your pay-as-you-go balance (FAQ).
  • MiniMax issues a Subscription Key on the plan page of its console. It is not interchangeable with a pay-as-you-go key. Claude Code's endpoint is https://api.minimax.io/anthropic.
  • Alibaba Coding Plan keys go to coding-intl.dashscope.aliyuncs.com and Token Plan keys to token-plan.ap-southeast-1.maas.aliyuncs.com. Both kinds start with sk-sp-, so the prefix does not tell you which plan a key belongs to. Alibaba says each plan's keys and URLs must be used in matching pairs, and a Token Plan key on the Coding Plan's URL gets a 401 (Token Plan quick start).

What each plan lets you connect

  • Z.ai limits the plan to the tools on its list. Claude Code, Codex, OpenCode, Cursor, Cline, Kilo Code and Roo Code are on it; a router in between is not. The usage policy allows rate limiting or freezing an account that breaks the rules, and a ban after more than three violations.
  • Alibaba's Coding Plan is for interactive use in coding tools such as Claude Code. Scripts, application backends and other non-interactive use are prohibited, and a key used that way can be revoked.
  • Alibaba's Token Plan accepts any tool that takes a custom base URL and key. Each tier caps how many agents can run at once, from 1-2 on Lite to 6-8 on Pro.
  • MiniMax names Claude Code, Codex, Cursor, OpenCode and others for the M Plan, and every tool on one Subscription Key draws from the same quota.

What does not work

Waiting out the wrong window. A GLM 1310 is the weekly window, and the next 5-hour reset will not clear it. MiniMax's windows are independent in the same way.

Alternating plans on every request. Each vendor caches only the requests it receives, so each switch sends the other one a stretch of conversation it has never seen. On GLM, uncached input costs about four times as many credits as cached. Use one plan until it fails, then the other.

Expecting an individual GLM plan to run into overage. Past the limit it stops, and your balance does not cover it. Z.ai's Team Plan can enable paid overage; MiniMax sells Credits.

Two plans behind one endpoint

Your coding tool has one base URL, so switching plans by hand means changing the URL and key and restarting. A local router can hold both plans and fail over between them. With kRouter:

npm install -g @sifxprime/krouter
krouter -t

Open http://localhost:20128/dashboard and go to Providers. Add your Z.ai key to GLM Coding (Z.ai), your MiniMax Subscription Key to Minimax Coding, and your Alibaba key to Alibaba Intl (Coding Plan) or Alibaba Token Plan. China-site accounts use GLM (China), Minimax (China) and, for the Coding Plan, Alibaba; there is no provider for the China-site Token Plan. Each provider already points at its plan endpoint, so you never type a base URL.

The built-in model lists trail the vendors. As of 0.5.163, glm-5.3-flash is missing from the GLM list, MiniMax-M3.1-Flash-Preview from the MiniMax list, and qwen3.7-plus and qwen3.6-plus from the Coding Plan list. Add each one with Add Model on the provider's page.

Then open Combos, click Create Combo, name it coding-plans, and add glm/glm-5.3 followed by minimax/MiniMax-M3.1-Flash-Preview. On a Token Plan, minimax/MiniMax-M3 from the built-in list also works. Leave the strategy on Fallback. Check it with one real request:

curl -s http://localhost:20128/v1/messages \
  -H "Content-Type: application/json" \
  -d '{"model": "coding-plans", "max_tokens": 32, "stream": false, "messages": [{"role": "user", "content": "Reply with the word ok"}]}'

For Claude Code, open CLI Tools, click Claude Code, put coding-plans in the Opus, Sonnet and Haiku slots and click Apply. kRouter merges this into ~/.claude/settings.json:

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:20128/v1",
    "ANTHROPIC_AUTH_TOKEN": "<your-krouter-key>",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "coding-plans",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "coding-plans",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "coding-plans"
  }
}

Claude Code also uses the Haiku slot for background work. To spend fewer GLM credits there, give it a second combo that starts with glm/glm-5.3-flash, the model Z.ai's own Claude Code guide puts in that slot. The same combo names work on the Codex and OpenCode pages (Codex setup). kRouter translates each tool's format into the plan's: Anthropic Messages for Z.ai's international endpoint and for MiniMax, Chat Completions for GLM (China) and both Alibaba plans.

How the failover behaves:

  • GLM's 429 sends the same request to MiniMax. Your tool gets one answer. GLM goes on a cooldown that starts at 2 seconds and doubles up to 5 minutes. kRouter does not read the reset time in Z.ai's error message, so unless the response also carries a standard Retry-After header, it keeps retrying GLM on that backoff while the window is empty. Each retry costs one failed round trip before MiniMax answers, and within about five minutes of the reset, requests go back to GLM.
  • The order is the order you set. kRouter can reorder a combo by remaining quota, but only for providers that report quota per model. GLM and MiniMax report windows for the whole plan, and kRouter reads nothing from Alibaba, so put first the plan you would rather spend.
  • A switch changes the model mid-session, and switches back when GLM resets. Try the pair on a task you have already solved before relying on it.
  • Quota Tracker shows GLM's 5-hour and weekly windows with reset times. For MiniMax it reads the Token Plan usage endpoint and shows the 5-hour and 7-day windows; if an M Plan key shows nothing there, use MiniMax's console. Neither Alibaba plan is tracked.

Where kRouter is not the answer

One plan that rarely runs out. Connect your tool directly with the vendor's guide. It is one hop fewer, and the vendor tests its own endpoint.

GLM, if account risk matters. A router is not on Z.ai's supported list. If that risk is too much, connect Claude Code to Z.ai directly and switch by hand when a window runs out.

Staying on MiniMax. If MiniMax's model is the one you want, purchased Credits keep you on it past the limit, with no change of model halfway through a task.

Choosing

Your situationDo this
GLM's 5-hour window runs out, weekly has roomMove heavy work off-peak, or add a fallback plan
GLM's weekly window runs outA second plan; the 5-hour reset will not help
You want one monthly pool, no 5-hour windowsAlibaba Token Plan
You cannot accept risk under plan termsConnect each plan directly, switch by hand

Common questions

Can I use the GLM Coding Plan outside Claude Code?

Yes, in the tools Z.ai lists, which include Codex, OpenCode, Cursor and Cline. The plan applies only on its plan endpoints and only to GLM-5.3 and GLM-5.3-Flash. Z.ai says use in unsupported tools can restrict your plan benefits.

Why does my plan key return 401 or "insufficient balance"?

Almost always the endpoint. Z.ai bills the plan only on api.z.ai/api/anthropic, api.z.ai/api/coding/paas/v4 and api.z.ai/api/v1. MiniMax needs the Subscription Key, not a pay-as-you-go key. Alibaba's Coding Plan and Token Plan each have their own base URL, both use sk-sp- keys, and a key on the other plan's URL is rejected.

Does the 5-hour reset also reset the weekly limit?

No. On GLM, error 1308 is the 5-hour window and 1310 the weekly one. MiniMax's two windows reset independently, and its text models need room in both.

Which plan gives the most for the money?

The monthly price cannot tell you, because the plans count different things: credits from tokens, usage weighted by model and context, or model calls. Run a normal week of your own work on one, read its usage page, and compare from there.

Can kRouter switch plans automatically?

Yes. Put both in a fallback combo and the next plan answers when the first returns a limit error. The order is the one you set, because these plans do not report quota per model. Quota Tracker shows GLM's windows and reads MiniMax's from its Token Plan usage endpoint; Alibaba's are not tracked.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter