Skip to main content
kRouter
All posts
How kRouter works

Codex CLI with any model: model_providers and wire_api

Codex now speaks only the Responses API. How to run it on Claude, GLM, Kimi or DeepSeek through config.toml, and what breaks along the way.

Kodelyth · The team behind kRouter
· Updated
8 min read

You copy a Codex config from a guide written last year, start Codex, and it refuses to load. Under "Error loading config.toml" and the offending line, the message says:

`wire_api = "chat"` is no longer supported.
How to fix: set `wire_api = "responses"` in your provider config.
More info: https://github.com/openai/codex/discussions/7782

Until early 2026, Codex could talk to a provider over Chat Completions, and any guide written before then may still tell you to. OpenAI announced the deprecation on December 9, 2025 (discussion #7782) and removed the code in Codex 0.95.0, released February 4, 2026. Today responses is the only value wire_api accepts, and it is the default when you leave the key out.

Changing one word is the easy part. The real constraint is that Codex now speaks only the OpenAI Responses API, while "OpenAI-compatible" in a provider's docs usually means Chat Completions. Point Codex straight at one of those and it still fails, just later and with a different error.

How Codex picks a provider

Two keys in ~/.codex/config.toml decide where requests go: model and model_provider. Anything other than a built-in provider is defined as a [model_providers.<id>] table:

model = "some-model"
model_provider = "myrouter"
 
[model_providers.myrouter]
name = "My router"
base_url = "http://localhost:20128/v1"
wire_api = "responses"
env_key = "MYROUTER_API_KEY"

Codex then sends POST {base_url}/responses. Three details catch people out:

  • The ids openai, ollama and lmstudio belong to built-in providers and cannot be reused. Codex refuses the config and asks you to rename the table.
  • By default a custom provider skips Codex's login screen and does not use codex login or auth.json, so no OpenAI sign-in is needed. Its key comes from env_key (the name of an environment variable), http_headers, env_http_headers, an auth command, or the discouraged experimental_bearer_token. There is no api_key field; a key written that way is never sent.
  • codex -m <model> overrides the model for one run, and -c key=value overrides any other key.

What does not work

wire_api = "chat". A hard error since February 2026, as above.

Putting the provider in the repository. Codex ignores model_provider, model_providers and openai_base_url in a project's .codex/config.toml, because repository contents should not decide where your credentials go. Put them in your user config or pass them with -c.

openai_base_url for a multi-provider router. It points the built-in OpenAI provider at another host, but that provider still requires an OpenAI sign-in and sends those credentials. Fine for a proxy in front of OpenAI; a router with its own provider keys belongs in its own [model_providers] table.

A Chat Completions URL with wire_api = "responses". The config loads, then the first request goes to /responses on a server that has nothing there.

The in-session /model picker. With a custom provider it has been reported to list only OpenAI's models, and picking one rewrites model in the active config file to an OpenAI name (openai/codex#35487, reported against 0.142.2 and still open). Use -m instead.

Do you need a translator at all?

A Responses request carries instructions and typed input items instead of a messages array, and its stream only counts as finished when a response.completed event arrives. A backend that does not speak this needs both directions converted on every turn, tool calls included (Responses API vs Chat Completions covers the differences). Whether kRouter should do that depends on the backend:

Your backendWhat to do
Serves the Responses API itself (the OpenAI API, Azure OpenAI)Codex's built-in OpenAI provider, or one [model_providers] block for Azure. A router only adds a hop.
One local model in Ollama or LM StudioCodex's own --oss mode.
Speaks only Chat Completions or Anthropic Messages, a subscription, or several providers at onceA translator between Codex and the providers.

The rest of this post covers the third row.

Pointing Codex at kRouter

kRouter accepts Responses requests on /v1/responses and converts them to and from whatever format the provider behind the model speaks.

npm install -g @openai/codex        # if Codex is not installed yet
npm install -g @sifxprime/krouter
krouter -t

Open http://localhost:20128/dashboard and connect at least one provider. Then open CLI Tools and the OpenAI Codex CLI / App card. Choose the endpoint, an API key and a model in provider/model-id form, optionally a different subagent model, and click Apply. kRouter merges this into ~/.codex/config.toml:

model = "cc/claude-sonnet-4-6"
model_provider = "krouter"
 
[model_providers.krouter]
name = "kRouter"
base_url = "http://localhost:20128/v1"
wire_api = "responses"
 
[model_providers.krouter.http_headers]
Authorization = "Bearer <your-krouter-key>"
 
[agents]
default_subagent_model = "cc/claude-sonnet-4-6"

Your other settings stay, but the file is rewritten, so its comments do not survive. Manual Config on the same card shows the block if you would rather paste it. Because the key travels in http_headers, your ChatGPT login in auth.json is left alone. To keep the key out of the file, use env_key = "KROUTER_API_KEY" instead of the header table and export that variable; applying from the card again puts the header back.

On the same machine, kRouter does not require a key unless you turn on Require API key on the Endpoint page or set REQUIRE_API_KEY=true. If no key exists yet, the card writes the placeholder sk_krouter, which only works locally and only while that requirement is off. From another machine -- Codex on a laptop, kRouter on a VPS or behind its built-in Cloudflare or Tailscale tunnel -- every request needs a real key from the Endpoint page. Without one, kRouter answers 401 "API key required for remote API access".

Before blaming Codex for anything, test the route on its own, streamed the way Codex calls it:

curl -sN http://localhost:20128/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <your-krouter-key>" \
  -d '{"model": "cc/claude-sonnet-4-6", "input": "Reply with the word ok", "stream": true}'

If text events arrive and the stream ends with a response.completed event, routing and translation work. Then try Codex itself with codex exec "reply with the word ok".

Which models to use

Any model on kRouter's 95+ providers can go in model. A few that suit Codex:

model valueComes fromWhat kRouter does
cx/gpt-5.5A ChatGPT plan, signed in to kRouter's Codex providerForwards the Responses request without translating it
cc/claude-sonnet-4-6A Claude subscription (Claude Code login)Translates Responses to Claude Messages and back
gh/claude-sonnet-4.6A GitHub Copilot subscriptionTranslates
kr/claude-sonnet-4.5Kiro, signed in with a device codeTranslates
ocg/kimi-k3, ocg/glm-5.3, ocg/deepseek-v4-proAn OpenCode Go subscriptionTranslates to Chat Completions and back
A combo name, no slashSeveral of the aboveTries the next model when one fails

Before you connect a Claude, Codex, GitHub Copilot or Kiro account, kRouter shows a risk notice: those subscription sessions are not officially licensed for proxy use, and the account may be restricted or banned. Read it before you decide.

For long sessions a combo is usually the better model: Codex keeps one stable name, and when the first model hits a rate limit or runs out of quota, the next one answers. Do not start the combo's name with a GPT model id (see below). OpenCode Go serves different models on different endpoints; see OpenCode Go's three protocols.

What breaks, and why

Six things behave differently once Codex is not talking to OpenAI.

The context window is a guess. For a name it does not recognise, Codex logs "Unknown model ... This will use fallback model metadata" and assumes a 272,000-token window. If the model you route to has less room, Codex does not compact the conversation in time, and the provider rejects the request instead. Set the real figure:

model_context_window = 128000            # the window of the model you actually use
model_auto_compact_token_limit = 100000  # compact a little before that

In Codex 0.160.1, an unknown model's window can be lowered this way but not raised above 272,000.

Reasoning effort does not travel. Codex sends model_reasoning_effort in the request's reasoning field. kRouter passes it through on cx/ but drops it when it translates to Claude or Chat Completions. Where a provider exposes thinking as a separate model id, choose it by name instead, such as Kiro's kr/claude-sonnet-4.5-thinking.

Hosted tools stay with OpenAI. Codex's built-in web search, on by default, runs on OpenAI's servers. On a translated route nothing can run it, so kRouter drops it.

MCP tools do not survive translation. Current Codex sends each MCP server's tools as one namespace tool, a Responses-only shape that groups several functions under one name. kRouter passes it through untouched to its Codex provider. On a translated route, as of 0.5.163, the whole group becomes a single function with no parameters, and the individual tools inside it are lost. If a session depends on MCP servers, run it on cx/, or point Codex directly at a backend that supports namespace tools.

GPT model names change what Codex sends. When Codex does not recognise a name, it retries the lookup with one leading provider/ segment removed, and it also matches any name that starts with a model id it knows. So kr/gpt-5.6-sol, or a combo called gpt-5.5-backup, gets OpenAI's settings for that model. In Codex 0.160.1, gpt-5.6-sol and the other GPT-5.6 and GPT-6 models use a newer request layout that lists the tools inside the conversation input instead of the request's tools field. kRouter 0.5.163 does not translate that layout, so on a translated route the model receives no tools at all. gpt-5.5 keeps the usual layout, but its apply_patch is a freeform tool that takes plain text instead of JSON; translated, it arrives without parameters, and the model cannot fill in a patch. Use GPT models in Codex through cx/, and keep the translated routes for Claude, GLM, Kimi and DeepSeek.

Streams that end early. Codex reports a stream that closes before response.completed as "stream disconnected before completion". By default it reconnects up to 5 times (stream_max_retries) and treats 300 seconds without data as a dropped connection (stream_idle_timeout_ms); both can be set in the [model_providers.krouter] table. The causes, and how to tell them apart, are in Codex "stream disconnected before completion".

Pin what works

Codex moves quickly: twelve stable releases between September 18 and October 5, 2026, from 0.155.1 to 0.160.1. Custom providers are not its main path, and updates occasionally break them. After 0.147.0, a LiteLLM user saw every request fail until going back to 0.146.0 (#37425), and an Azure user got Invalid 'input[0].tools[0].description': empty string (#37380). If an update breaks a setup that worked yesterday, go back one version:

npm install -g @openai/codex@<last-version-that-worked>

On the kRouter side, run 0.5.163 or newer (npm i -g @sifxprime/krouter@latest). It stopped dropping the developer messages Codex sends when the model is one GitHub Copilot serves only through its Responses API. kRouter does not update itself, so it stays on the version you installed until you run that command.

Common questions

Does Codex CLI still support wire_api = "chat"?

No. Chat Completions support was deprecated in December 2025 and removed in Codex 0.95.0 in February 2026. responses is now the only accepted value, and the default when the key is missing. A provider that only speaks Chat Completions needs a translator in front of it.

Can Codex CLI use Claude models?

Yes, through something that converts the Responses API to Anthropic's Messages API and back, such as kRouter with model = "cc/claude-sonnet-4-6" or gh/claude-sonnet-4.6. Expect four differences from GPT models: you set model_context_window yourself, Codex's reasoning-effort setting does not reach the model, its built-in web search is not available, and as of kRouter 0.5.163 tools from MCP servers do not work through the translation.

Where does the API key go for a custom provider?

In the provider table, as an Authorization header under http_headers or as env_key naming an environment variable. A custom provider does not use auth.json or codex login by default, and there is no api_key field. For kRouter on the same machine the key is optional unless you require one; from any other machine it is required.

Why does Codex warn about fallback model metadata?

It does not recognise the model name, so it uses generic settings, including a 272,000-token context window. If the model has less room, set model_context_window and model_auto_compact_token_limit so Codex compacts before the provider refuses the request.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter