Skip to main content
kRouter

Changelog

Every release.

What shipped, what was fixed, what was verified. The last 10 entries — pulled live from the kRouter repo CHANGELOG.md.

Latestv0.5.159
Releases on GitHub

September 2026

  1. v0.5.159

    the crash-loop safety valve now actually closes

    v0.5.158 fixed the directory the MITM crash-loop recovery looked in. It did not fix the larger problem, which only surfaced when the router was finally run end to end against a real, isolated data directory: **the recovery was writing to a file the app abandoned.** When the server crashes `MAX_RESTARTS` times in a row, the CLI supervisor assumes MITM is the cause, disables it, and restarts. It did that by patching `settings.mitmEnabled` in `DATA_DIR/db.json`. But `src/lib/db/migrate.js` imports that JSON into SQLite once, writes a `.migrated-from-json` marker, and keeps the JSON only as a rollback artifact: - on a **fresh install**, `db.json` never exists — `existsSync` is false and the write is skipped - on a **migrated install**, it exists but nothing reads it — the write lands in a dead file Either way the safety valve did nothing and the server kept crash-looping. The `catch { /* best effort */ }` wrapped around it is why this stayed invisible through many releases. ## The supervisor no longer knows how settings are stored That assumption is what rotted here, twice — first the directory, then the file itself. So it is gone. The supervisor writes a marker, `DATA_DIR/.mitm-recovery`, and the app disables MITM itself through `updateSettings()` — the same path every other caller uses. Storage can change again; this cannot silently rot with it. Three details matter, and each is pinned by a test: - The marker is consumed at the top of `autoStartMitm()`, **before** settings are read, so the change takes effect on the boot that handles it rather than one boot later. - The marker is removed **only after** the write succeeds. A failed write leaves it in place so the next start retries, instead of dropping the request silently. - A failed write is **logged, not swallowed**. `updateSettings()` was verified rather than assumed: it is an upsert (`ON CONFLICT(id) DO UPDATE`) performing an atomic `{...current, ...updates}` merge inside a transaction. Both properties were load-bearing — a fresh install has no settings row at all, so a plain `UPDATE` would have matched zero rows and reproduced the same class of bug a third time. ## Honest limit This is verified by contract tests and by reading the write path, **not** by observing a live recovery. `initializeApp()` only runs when a page renders through the root layout, and in a fresh unauthenticated install every route either redirects or is prerendered, so the boot path could not be driven in a sandbox. The corollary is worth knowing: this recovery — in its old design and its new one — only fires once a page has rendered. An API-only deployment would never reach it.
  2. v0.5.158

    DATA_DIR is honoured everywhere, including where it was silently failing

    **Upgrade if you set `DATA_DIR`.** `cli.js` ignored it and resolved `~/.krouter` unconditionally, so a single process disagreed with itself: auth and machine-id came from `DATA_DIR`, while `db.json`, the tunnel directory and the mitm pidfile came from the home directory. `DATA_DIR` is documented — the README lists it as *"Data directory (SQLite, certs, cache)"* and `docs/ARCHITECTURE.md` states `db.json` lives at `${DATA_DIR}/db.json` — so this was a broken promise, not an undocumented edge. The consequence worth the release was silent. When the server crash-loops past `MAX_RESTARTS`, the CLI clears `settings.mitmEnabled` in `db.json` to break the loop. Looking in the wrong directory made `existsSync` return false, the surrounding `catch { /* best effort */ }` swallowed it, and the server kept crash-looping with the safety valve doing nothing at all. Worse, a stale `~/.krouter/db.json` left from before `DATA_DIR` was set would be written to instead of the real one. Mostly this hit Docker and multi-instance users — the people most likely to set `DATA_DIR`, and the least likely to be watching a terminal when the loop began. ## The cause was duplication Three hand-synced copies of the same resolver, each carrying a *"kept in sync with…"* comment, and one had drifted. They are now one module, `cli/src/lib/dataDir.js`, shipped inside the package. It cannot import the app's own `src/lib/dataDir.js` — that is ESM and lives outside the CLI package — so it mirrors it, which is the relationship the old comments described, minus the drift. Two smaller bugs came out with it: - `cli.js` fell back to `process.env.APPDATA || ""` on Windows, so an unset `APPDATA` produced a **relative** path and wrote into the current working directory. It now falls back to the home directory, as the app already did. - `api/client.js` honoured `DATA_DIR` but without the Windows-path and writability guards, so one setting could resolve three different ways inside one process. The `mkdir` fires only for an explicitly configured `DATA_DIR`; the default path stays a pure lookup, with a test pinning that so it cannot become a surprise side effect of importing a module. Seven new tests cover both behaviour and shape — two of them assert the deleted copies stay deleted, since a drifted duplicate is what caused this in the first place. ## Also - Removed two dead `gitbook` rules from `.gitignore`, finishing the removal started in v0.5.157.
  3. v0.5.157

    an unauthenticated RCE on Windows hosts, and a docs app that never deployed

    **Upgrade if you run kRouter on Windows.** Next.js published two critical advisories against the 16.3.1 this release was built on, and one of them reaches the dashboard: - **[GHSA-p293-qw3h-jr36](https://github.com/advisories/GHSA-p293-qw3h-jr36)** — unauthenticated remote code execution on Windows-hosted servers, fixed in 16.3.3. `server.js` binds `0.0.0.0` by default, so on Windows the dashboard is reachable from the local network by anyone who can route to the machine. This one is real and it is why the release exists. - **[GHSA-2xp9-vwfh-vxw4](https://github.com/advisories/GHSA-2xp9-vwfh-vxw4)** — unauthenticated RCE in the Image Optimization API when AVIF files are used. **Not reachable here.** `next.config.mjs` has set `images.unoptimized` since long before the advisory, so there is no `/_next/image` endpoint to attack. Recorded rather than omitted, because "we ship a vulnerable version" deserves a reason and not silence. Both are now closed by Next.js 16.3.4, which also brings `sharp` to 0.35.4 and clears [GHSA-rgj7-g3m4-5g8c](https://github.com/advisories/GHSA-rgj7-g3m4-5g8c). `npm audit --omit=dev` goes from 2 findings, one critical, to zero. Neither advisory was reported by Dependabot. It reads manifests, and the vulnerable version was a lockfile resolution of a `^16.2.9` range — so nine alerts sat open against a different directory while the two that mattered went unmentioned. ## The nine alerts were noise, and the directory is gone Every one of those nine was the same package, `next`, pinned at 16.1.1 in `gitbook/` — a second Next.js app inherited from the fork. All nine require a running Next server. `gitbook` had none: it was `output: "export"` with no middleware, no route handlers and no server actions. It was also dead. Its workflow ran exactly once, on 2026-06-14, and failed. It deployed to `9router.github.io`, a domain this project does not own. Nothing had touched it since the initial fork commit. It has been removed along with `gitbook-pages.yml` and the two `next.config.mjs` entries that existed only to keep it out of the bundle — 121 files, and the entire source of the alert noise that was hiding the real problem. ## The support cards line up now The two contact cards sized themselves to their own text, so `[email protected]` made the email card 28.79px wider than the WhatsApp one, with left edges 28.79px apart. Only their right edges agreed, because `align-items: flex-end` was the only thing aligning them. The container now carries one definite width and both cards fill it, so neither can size to its label: measured live, width, height, left and right all agree to 0.0000px. The marks inside them are matched too. The official WhatsApp path is drawn edge-to-edge in its 24-unit box while the envelope stops at 94%, so at an identical box size WhatsApp rendered 0.86px wider and 4px taller and read as the heavier of the pair. Its viewBox is padded to inset it to the envelope's exact fraction; the two ink widths now land 0.009px apart, with the official path data untouched. ## Also - A rejected `krouter-web` dispatch fails the job again. The website's own cron is declared every 10 minutes but measures at a 212-minute median across 85 runs, with a 751-minute worst case, so a silent dispatch failure means the site lags a release by hours rather than minutes.
  4. v0.5.156

    the two support marks now read at equal weight

    The envelope in the support widget looked bigger than the WhatsApp mark. Measured on the live widget, it is not: the envelope carries 177px² of ink against WhatsApp's 255px², and the two boxes are within a pixel of each other in width. What differs is contrast. The envelope rendered in near-black `rgb(11, 13, 18)` on a pale badge while WhatsApp sat at mid-green `rgb(37, 211, 102)`, and the darker mark dominates regardless of size — which is why the earlier attempt at this, resizing the glyph, did not fix it. Each channel now carries its own colour with a 10% wash of it behind the glyph, so the badge takes some of the weight and neither mark has to carry it alone. WhatsApp keeps its brand green; email takes a blue. They read as a pair rather than one loud and one quiet.
  5. v0.5.155

    the last of the upstream backlog, and 24KB of dead code that was shipping

    The end of the upstream triage. Of the 19 commits still marked worth porting, **11 are in**; the other 8 depend on subsystems this fork does not have and are listed below with the reason rather than left unexplained. **A stale backup file was shipping inside every install.** `chatCore.js.bak`, a copy of the chat handler taken at 0.5.40, had been tracked ever since — 24KB of dead code inside the npm package, 243 diff-lines behind the live file. Anyone who opened it looking for the real handler would have been reading code with no SSRF guards, no usage accounting, and no clientTool wiring. Deleted, with `*.bak` added to `.gitignore` so the next one cannot be committed by accident. **Claude requests poisoned by a foreign tool id.** Anthropic validates `server_tool_use` ids against `^srvtoolu_` and rejects the whole request with a 400 when one does not match. A combo that falls back to a provider with its own built-in tools — z.ai/glm emits OpenAI-style `call_` ids for its `analyze_image` tool — leaves those blocks in the history, so every later Claude turn carried a poisoned id and kept failing. Those blocks are dropped now, along with the result blocks that referenced them (an orphaned `tool_result` is rejected just as firmly) and any message left with no content at all, since an empty text block is its own 400. (upstream `ed1bd0c5`) **GLM credit plans showed no quota at all.** Quota parsing accepted only `TOKENS_LIMIT` and wrote every limit to the same `session` key. An account on a `CREDIT_LIMIT` plan therefore saw nothing, and an account with more than one window kept only whichever arrived last — a weekly figure displayed as the session one. Both types are accepted now, each window keyed by its own type and period. (upstream `fcfcced4`) **OpenCode Go quota was not tracked at all.** A subscriber saw nothing on the usage dashboard for that provider, including when a window was exhausted and their requests had already started failing. Rolling, weekly and monthly windows are read now, and a missing subscription is reported distinctly from plain forbidden — those fail identically at the transport level, and only the error type separates them. (upstream `0da803ee`) **Antigravity onboarding was tripping Google's anti-abuse limiter.** It retried five times in rapid succession; with several accounts refreshing at once those calls arrived as a burst from one IP and the whole set got rate-limited. Two attempts now, a twelve-second base delay, and jitter so concurrent refreshes do not stay in lockstep. Both knobs are env-overridable for a slow network. (upstream `1442cc73`, the `projectId` half) **Also in:** `GET /v1/models/{provider}/{model}` for single-model lookup, OpenAI-compatible, with a proper `model_not_found` 404 — this replaces the `[kind]` route rather than sitting beside it, since a provider-prefixed id contains a slash and only a catch-all can capture it (`5caa72f5`). The GPT-5.3-Codex-Spark quota window, which is metered separately and previously left users staring at unexplained 429s with no row to explain them (`40eed186`). Antigravity image sizes now resolve to the aspect-ratio model suffix that provider actually expects (`2a9213c5`). A completed Responses call no longer logs a disconnect — codex and droid close the socket on every one, which made a normal finish look like a client hanging up and buried the disconnects that are real (`c4af43fa`). And system-prompt injection is exact-idempotent, closing a latent way to double a persona and bill for it inside a feature meant to save tokens (`cadef6c4`, dedup only — ours has been format-aware across six formats all along). **Not ported, and why.** `7e5f5a88` re-anchors Claude cache breakpoints through a subsystem this fork does not have, and our passthrough deliberately preserves the client's body byte-for-byte so Anthropic's prompt cache is not busted — the opposite strategy. `ac98dd9d` and `1a3db1ef` need an `antigravityQuota` service we do not carry. `ab044e6d` and `acb5c34c` route Muse Spark through a `transports[]` mechanism we do not have. `2ab6a4c9`, `cec672d9`, `e014cb53` and `ed963931` are catalog refreshes that land in a provider registry we do not have. `c24a8542` extends an endpoint- presets module we never had. 1957 tests pass, and the real end-to-end suite passes against live provider credentials.
  6. v0.5.154

    three regressions from 0.5.153, and an end-to-end suite that never ran

    An adversarial pass over the previous release's own diff found three defects it had introduced. All three shipped. All three are fixed here, each verified against a real runtime rather than by reading the code. **SQLite silently downgraded for most Node 22 users.** 0.5.153 chose the better-sqlite3 build by Node major version: `NODE_MAJOR >= 22` selected 13.x. But 13.x declares Node-API 10, and Node only supports that from **22.14.0**. Node 22.11.0 — the release that became Node 22 LTS, and the most widely installed 22.x — reports Node-API 9 and cannot load the addon at all. The failure was invisible. The binary probe found the prebuild on disk, the magic-number check passed, and the installer printed "SQLite engine ready" — while the driver quietly fell through to a slower path. Those users had a working 12.6.2 until 0.5.153 took it away and told them otherwise. The gate now keys on `process.versions.napi`, which is the thing that actually decides whether the addon loads, rather than on a version number that only correlates with it. Verified against four installed runtimes: 22.11.0 and 22.13.1 report napi 9 and get 12.6.2; 22.14.0 and 24.14.0 report napi 10 and get 13.0.3. **Dimmed icons stopped being dim, and chevrons stopped rotating.** The icon-reveal rule from 0.5.153 was written outside any cascade layer. Unlayered rules beat every layered one, and Tailwind's utilities live in `@layer utilities` — so `.material-symbols-outlined { opacity: … }` outranked them. Icons carrying `opacity-20` or `opacity-50` rendered at full strength: the faint empty-state glyphs on the usage cards became full-strength 48–64px icons, and a hover affordance meant to fade from 50% to 100% was simply always lit. The same rule used the `transition` shorthand, which resets `transition-property`. That clobbered `transition-transform` on every element that is also an icon — the expand chevrons on 15 CLI tool cards, the usage table row expander, and the sidebar caret all snapped instead of rotating. Both are fixed by moving the rules into `@layer base` and setting transition longhands. Verified in a browser against the compiled stylesheet: an icon with `opacity-20` now computes `0.2`, one with `opacity-50` computes `0.5`, `transition-transform` reports its own properties again, and a plain icon still reveals to full opacity once the font loads. **The real end-to-end suite has never run.** There is a smoke test that drives the full production path — `handleChatCore`, real credentials, real network — against every provider with an active connection. It looked for the database at `~/.9router`, the *upstream* directory. This fork stores it in `~/.krouter`, so the path never existed; a bare `catch` swallowed the ENOENT and the suite reported "no active providers" as if the machine simply had none configured. It has been silently testing nothing since the fork renamed its data directory. Fixed to resolve the data directory the way the CLI does, Windows branch included, and to say *why* it found nothing — an unreadable database and an empty one are different problems. With the path corrected it ran, and failed for a second reason: `translator/index.js` registers translators with `require()`, a bundler-only pattern that no-ops under vitest. With an empty registry `translateRequest` returns the body untouched, so a raw OpenAI payload was posted to providers that expect their own envelope — Antigravity answered `Unknown name "messages"`, Kiro answered `Improperly formed request`. `tests/translator/registerAll.js` exists for exactly this case and the suite did not import it. Both fixed, and the suite now passes against live Antigravity and Kiro credentials. 1913 tests pass.
  7. v0.5.153

    fifteen upstream fixes, and a claim the dashboard should never have made

    The remaining upstream backlog, triaged commit by commit against our own code. Of 83 candidates, 15 were worth porting; the rest depend on subsystems this fork does not have (the provider registry, the usage submodule, the resolveSessionId family), add models we do not carry, or are changelog and dependency churn. Every port below was checked against our tree rather than applied on the strength of its commit message, and each is covered by a test verified against the *unfixed* code. **Tool schemas that Gemini and Antigravity reject outright.** Two shapes fail hard: `prefixItems` (2020-12 tuple validation) comes back as `Unknown name prefixItems`, and a `type: "array"` with no `items` is rejected for a missing required field. Zod's `z.tuple()` emits the first and a great many MCP server tool schemas emit the second, so any client with MCP tools pointed at one of those connections lost the whole turn rather than degrading. The sanitiser now converts `prefixItems` to `items` — one variant becomes `items` directly, several become an `anyOf`, a `null` variant is dropped — fills a permissive `items` on any bare array, and strips both keywords wherever they survive conversion. (upstream `f6c59d30`) **A failed web search took the account offline for chat.** `markAccountUnavailable` with no model key writes `modelLock___all`, an account-level lock that `isModelLockActive()` honours for every model. The search handler passed no key — and `providerId` there is an ordinary `AI_PROVIDERS` entry, so search draws on the same connections chat does. One failed search disabled that connection for chat until the cooldown expired. Locks are now scoped to `websearch:<provider>`, read back under the same key, and verified to survive persistence: connection rows already carry dynamic keys like `modelLock_claude-opus-4-6-thinking` in a JSON column, so the colon is legal. A genuine account ban still forces the account-wide lock, which is correct. (upstream `ec669280`, adapted — upstream scopes via a `credentialFallback` path this fork does not have; here the same collision arrives through the shared provider id) **The dashboard told remote viewers their data was on their machine.** "Local Mode - All data stored on your machine" was shown to everyone, including anyone reaching the dashboard over a tunnel, Tailscale, or a LAN address — all of which this fork actively supports. It is exactly the claim someone leans on before pasting a provider key. Upstream fixed the footer line; our page repeats it in a prominent card, so that is corrected too. Then a second pass on our own port: the flag can only be read after mount, so initialising it to `false` still painted the false claim for one frame. It now starts from `true` — a local user sees "Remote Mode" for that frame instead, understating safety rather than overstating it. (upstream `28cfd9fa`, plus a fix of our own) **Anthropic rejected proxied Claude requests.** An `anthropic-compatible-*` node serving a real Claude model sits in front of Anthropic itself — a rotating multi-account proxy, a corporate gateway — and needs the beta flags the `claude` provider sends. Without `context-management-2025-06-27`, Anthropic rejects the `context_management` block Claude Code puts in every request with "Extra inputs are not permitted" (400) and the combo falls through to the next model with nothing surfaced. The model id gates it, so a node fronting Kimi or GLM never matches and is left alone. (upstream `fb9fab02`) Also: Anthropic requires an explicit `type` on every tool, and strict gateways fronting it answer 400 without one — Claude-format tools are now defaulted to `type: "custom"`, with the spread ordered so a falsy `type` is overridden rather than surviving to 400 anyway (`e08ac6da`). And a client already speaking the Responses API that declares a no-argument tool as `{type:"object"}` got a hard 400 from the Codex and Grok CLI backends; our translator fills that in, but a Responses-native body short-circuits past it, so both executors now fill it themselves (`11222eff`, that hunk only). **SQLite needed build tools it could not assume.** better-sqlite3 12.6.2 has no N-API build, so on Node 22+ npm runs node-gyp against it — and a user without build tools, the common case on Windows and on a clean macOS with no Xcode CLT, failed the native install and dropped silently to the much slower sql.js fallback. 13.x is N-API and ships per-platform prebuilds inside the package. Scripts are skipped for that build only, since npm injects an implicit `node-gyp rebuild` for anything carrying a `binding.gyp`; doing the same on 12.x would leave it with no binary at all. The validity check reads both layouts and tells musl from glibc, so an Alpine container looks for `linuxmusl`. (upstream `90a00058`) **Smaller things that were still real.** Connection tests had no deadline at all, so a provider whose endpoint blackholes traffic left the Test button spinning forever and pinned a socket per retry — 15s now, and only when the caller set no signal of its own (`df85e16d`). Antigravity flags competing-client branding in a system prompt and answers 429 Quota Exhausted; we stripped Zed's line but sent OpenCode's naming verbatim, and the ZWJ obfuscation never covered it because it only rewrites `contents` (`dff64849`). An OpenAI Responses usage body matched the Claude branch of usage extraction, which read neither cache field, so every `/v1/responses` and codex request logged a zero cache read (`e7dd72a8`). Dark-mode users saw a white flash on every load, since the theme was only applied from an effect (`925cb4aa`). Icons rendered blank or as raw ligature text on a cold cache, because `document.fonts.ready` settles on the text faces alone while the 3.8MB icon font has not begun downloading (`14401c43`). `"•".repeat()` threw a RangeError on an API key shorter than the 8-character prefix, crashing the whole card (`bb3cb43e`). A dead OpenCode model is no longer suggested (`44e4b80b`). And the rtk headroom path now records a reason instead of returning null silently (`548e32aa`). The adversarial verification pass for this release could not run — every agent hit a session limit — so the risky parts were checked by hand instead: the colon in the new lock key survives persistence, the rewrite regex is global (`replaceAll` throws on one that is not), the schema walker does not mistake a property named `items` or `type` for the keyword, the sqlite binary probe returns false rather than throwing when neither layout exists, the icon font family string matches what the package ships, and the theme script survives a `localStorage` that throws. 1909 tests pass.
  8. v0.5.152

    MiniMax M3 on the right endpoint, and a way to reach a human

    **MiniMax M3 was being sent to the wrong endpoint.** Upstream's provider registry lists `minimax-m3` with `supportedFormats: ["openai", "claude"]` — the same pair as `minimax-m2.5` and `m2.7`, both of which this fork already routes to `/zen/go/v1/messages`. M3 was missing from our catalog altogether, so it carried no `targetFormat` and fell through to `/chat/completions` with an OpenAI-shaped body. That choice lives in two places that have to agree: `targetFormat: "claude"` in `config/providerModels.js` decides how the request *body* is translated, and `CLAUDE_FORMAT_MODELS` in the executor decides which *endpoint and auth header* it goes out with. A model in one list but not the other posts a Claude body to `/chat/completions` with a bearer token, or an OpenAI body to `/messages` with `x-api-key` — and neither failure names its real cause. Both are updated, and a test now asserts the two lists are identical, so the next model can't be added to one and forgotten in the other. This one is not verified against the live provider: there's no OpenCode Go connection on the build machine to probe with. The evidence is upstream's registry plus the identical treatment of its two siblings. **There was no way to reach a human from inside the product.** A user who hits a provider-side change — the kind that breaks every request at once, as OpenCode Go's session header did yesterday — had to go looking for an address off-site. The dashboard sidebar now has a Support control, pinned below the nav so it stays reachable from any page and doesn't scroll away. It's collapsed by default so it never competes with navigation; clicking it opens two channels, WhatsApp and email. The website gets the same thing as a bottom-right widget, so a reader who never installs anything can still ask a question. Both marks are inline SVG rather than the icon font the nav uses — that font has no WhatsApp glyph, and a recognisable brand mark is the entire point of the affordance. Destinations live in one constants module per repo rather than inline in the markup, because a second copy of the number is exactly how the link and the label a user reads drift apart; the test asserts the `wa.me` form actually resolves (digits only — a leading `+` gives a broken link) and that the displayed number matches the one it dials. Keyboard and screen-reader behaviour is part of the contract rather than an afterthought: `aria-expanded`/`aria-controls` on the trigger, Escape closes and returns focus to it instead of dropping the user at the top of the document, and the links leave the tab order while collapsed so Tab can't land on invisible targets. Verified in a browser at desktop and 375px. 1851 tests pass.
  9. v0.5.151

    OpenCode Go started requiring a session header

    A user on `ocg/glm-5.2` reported every request failing: 400 MissingSessionID — "Request is missing x-opencode-session and cannot be routed efficiently." OpenCode Go began enforcing `x-opencode-session` on every request. This fork never sent it, so the whole provider stopped working the moment they turned it on — not a subset of models, not a degraded mode. Every `ocg/*` request, 400. Upstream added the header in `81f4f930`, through a `resolveSessionId()` helper in `sessionManager.js` that this fork does not have. Rather than port that whole subsystem, the precedence is walked by hand — the same shape `resolveGrokCliSessionId()` already uses in `grok-cli.js` for exactly this reason: an explicit conversation id from the client wins, since it is the only thing that actually tracks a thread, then the workspace, then a stable id derived from the connection. The value is namespaced by client tool before hashing to `ses_<32 hex>`, so two tools that both call their thread `1` don't collide into a single upstream session. A caller already speaking OpenCode's own protocol keeps its header verbatim — rewriting it would split one conversation across two sessions upstream. It is computed in `execute()` and carried on a copy of the per-request credentials, never on the executor itself: that object is a singleton, and a field on `this` would leak between concurrent requests. `buildHeaders()` applies it to both transports — `/chat/completions` and the Claude-format `/messages` branch — and falls back on its own, so no path can produce a header-less request. **The request path wasn't the only caller.** Connection validation and the dashboard's test-connection button build their own `fetch`, and were sending header-less requests too. Neither showed up as a failure, because both grade anything that isn't 401/403 as a healthy key — so they were reporting success on a 400 and validating nothing. Both now use the same exported helper rather than a second copy of the hashing. Verified that `prompt_cache_key` survives `openai → openai` translation, so the explicit branch is genuinely reachable for clients that send one; clients that send nothing get a stable per-connection session, which is coarser but never changes mid-conversation. Against the shipped 0.5.150 code, 7 of the 11 new tests fail. 1838 tests pass.
  10. v0.5.150

    a token in the logs, a billing hole, and five broken client paths

    The second half of the upstream triage. v0.5.149 shipped the security findings; this ships the seven high-value functional ones. Every commit was checked against our own code rather than applied on the strength of its message, and each fix was verified against the *unfixed* build — if reverting the change did not turn a test red, the test was wrong and got rewritten. **A live token pair, written to stdout.** `refreshCline` in `open-sse/executors/default.js` logged the refresh response: console.log('[DEBUG] Cline refresh payload:', JSON.stringify(payload).substring(0, 200)) That response is `{ data: { accessToken, refreshToken, expiresAt } }` — the shape the very next line parses. 200 characters is more than enough for both tokens in full, so every Cline token refresh printed working credentials to the terminal and into any log capture behind it. Four more `[DEBUG]` lines around it logged the token length and the raw error body. All five are gone, and a test now fails if the string `[DEBUG]` reappears in that executor or if any captured console output contains a token. This was found while porting upstream `88676b30`, which turned out to be a fix we already had — our fork already spoke the extension JSON contract. What we did not have was the hardening, now adopted: a failed refresh returns `null` instead of throwing out of the caller and aborting the whole request, a body with no `accessToken` bails instead of returning `{ accessToken: undefined }` for something downstream to trip over, expiry falls back to `expiresIn`/`expires_in`/3600 rather than `undefined`, and the refresh URL is read from `PROVIDERS.cline` instead of a second hardcoded copy of it. **Requests that billed nothing at all.** The Responses API has no `[DONE]` sentinel, so codex closes the socket the moment `response.completed` arrives. That cancels the reader, `flush()` never runs — and every usage side effect lived only in `flush()`. Not an estimate, not a zero row: no row. The usage tail is now a once-guarded `finalizeStream()` called from the terminal event as well as from flush. Placement is the whole fix. It fires *after* the terminal chunk's usage is extracted and handed to the client, not where the event is detected; anchored at the detection point it would have estimated instead of using the real numbers the provider had just sent. Verified against the unfixed code: zero usage records before, one after. (upstream `d7f7d70d`) **Codex could not authenticate at all.** Codex reads a custom model provider's credentials from `env_key`, `http_headers`, `env_http_headers` or a token command. `auth.json` is read only by its *built-in* openai provider — so writing `OPENAI_API_KEY` there left every request unauthenticated with `401 Missing API key`, while overwriting the user's existing ChatGPT login on the way past. The key now goes in `[model_providers.krouter.http_headers]` and `auth.json` is no longer written. `DELETE` still clears it, to repair machines the previous version already configured. The subagent model moved to the `agents.default_subagent_model` scalar, because `agents.<role>` now declares a custom role and requires a description — the old `[agents.subagent]` table was being discarded with a startup warning. The card's status regex reads the new key too, or the dashboard showed it blank. Verified the emitted TOML round-trips with the header intact, preserves unrelated `[agents]` keys, and orders the sub-table after the provider scalars — wrong order and Codex refuses the file. (upstream `9c45b27c`) **Ollama's last chunk — the one with the token counts — was dropped.** `createSSEStream` splits on `\n` and keeps the remainder for `flush()` to parse, but that call omitted `targetFormat`, so `parseSSELine` demanded a `data: ` prefix and discarded whatever an NDJSON provider left without a closing newline. The `!parsed.done` guard compounded it: the SSE sentinel and an Ollama final chunk both carry `done:true`, but the latter is the real last chunk, holding `done_reason` and the token counts. The sentinel check is now scoped to formats that emit one. (upstream `f9d82c65`) **Claude requests that 400'd before failover could try the next hop.** Anthropic rejects a tool carrying both `defer_loading:true` and `cache_control`. MCP clients put deferred tools at the tail — exactly where we anchored the 1h cache breakpoint. The anchor now lands on the last tool that *can* be cached, so caching is kept for the tools that can use it instead of dropped wholesale. (upstream `6ab9ca9e`, #3567) **Parallel tool calls arriving as one malformed call.** Responses-to-chat translation attributed every arguments delta to a positional index that only advanced on `output_item.done`. When an upstream emits all `output_item.added` events before any dones — normal for parallel tool calls — every delta landed on index 0, and the client concatenated N JSON payloads into a single tool input and failed validation. Indices are now keyed off the server item id, assigned when the item is added; a duplicate added (a retry) reuses its index rather than allocating a new one. (upstream `e74db4d0`, taking only this half — the `muse-spark` model it also adds needs two modules this fork does not have) **CommandCode errors shown to users as assistant output.** CommandCode reports failures as a `type:"error"` event inside an HTTP 200 NDJSON stream. Combo and account fallback key off `response.status`, so a 200 never triggered them and the error text was streamed to the client as if the model had said it. The leading events are now peeked before the stream is committed; an error event becomes a real 4xx/5xx that the existing fallback can see, with the status taken from the event or inferred from its text so a rate limit fails over differently than an auth failure. Normal streams replay losslessly and the peek stops at the first event proving real content. (upstream `67d9182e`) Two things came out of porting that one. Upstream's tests exercise the helper directly, so they stay green even if `execute()` never calls it — reverting the wiring did not turn them red, so coverage for `execute()` itself was added. That new test then caught a defect upstream carries too: the response headers were built from an object literal holding both `Content-Type` and `content-type`. Those are two distinct JS keys and the `Headers` constructor appends rather than replaces, so every CommandCode response went out with a doubled `text/event-stream, text/event-stream`. **The peek that fix needed, fixed in turn.** An adversarial sweep over this release's own diff, run before publishing, found that the new CommandCode peek rebuilt the stream it had already consumed out of parsed, trimmed text lines. Three defects came from that, every one of them dependent on where the network split the body — which is exactly why the suite stayed green: it fed one line per read, the single framing that never exercised the replay. Lines sharing a read with the first content event were dropped, because the peek stopped there and the rest of that chunk existed nowhere else — a short answer arriving in one chunk lost everything past its first token, terminal event included. A final line with no trailing newline was emitted twice, the done branch having pushed it into the replay list without clearing the text buffer. And a multi-byte character split across a read came out corrupted, because the peek and the wrapper each held their own `TextDecoder` and the bytes stranded in the first were never handed to the second. The peek now keeps each raw chunk and replays those bytes verbatim; nothing is re-derived, so there is nothing to lose, duplicate or mis-decode. Reproduced independently before fixing — identical bytes framed as one read versus one line per read gave `AAA` against `AAABBBCCC`, `ONCEONCE` against `ONCE`, and `AA你好世界` arriving with a replacement character where its third character should be — and the new tests vary the framing rather than the content. **`npm test` failed after `npm run build`.** `next build` copies the tree, tests included, into `.next/standalone`, and vitest had no exclude — so it ran every test twice and went red on stale build artifacts rather than on anything in source. Excluded, and verified by planting a stale copy and watching the count stay put. 1824 tests pass, up from 1777.