Codex CLI with any model: model_providers and wire_api
Codex now speaks only the Responses API. How to run it on Claude, GLM, Kimi or DeepSeek through config.toml, and what breaks along the way.
You copy a Codex config from a guide written last year, start Codex, and it refuses to load. Under "Error loading config.toml" and the offending line, the message says:
`wire_api = "chat"` is no longer supported.
How to fix: set `wire_api = "responses"` in your provider config.
More info: https://github.com/openai/codex/discussions/7782Until early 2026, Codex could talk to a provider over Chat Completions, and any guide written before then may still tell you to. OpenAI announced the deprecation on December 9, 2025 (discussion #7782) and removed the code in Codex 0.95.0, released February 4, 2026. Today responses is the only value wire_api accepts, and it is the default when you leave the key out.
Changing one word is the easy part. The real constraint is that Codex now speaks only the OpenAI Responses API, while "OpenAI-compatible" in a provider's docs usually means Chat Completions. Point Codex straight at one of those and it still fails, just later and with a different error.
How Codex picks a provider
Two keys in ~/.codex/config.toml decide where requests go: model and model_provider. Anything other than a built-in provider is defined as a [model_providers.<id>] table:
model = "some-model"
model_provider = "myrouter"
[model_providers.myrouter]
name = "My router"
base_url = "http://localhost:20128/v1"
wire_api = "responses"
env_key = "MYROUTER_API_KEY"Codex then sends POST {base_url}/responses. Three details catch people out:
- The ids
openai,ollamaandlmstudiobelong to built-in providers and cannot be reused. Codex refuses the config and asks you to rename the table. - By default a custom provider skips Codex's login screen and does not use
codex loginorauth.json, so no OpenAI sign-in is needed. Its key comes fromenv_key(the name of an environment variable),http_headers,env_http_headers, anauthcommand, or the discouragedexperimental_bearer_token. There is noapi_keyfield; a key written that way is never sent. codex -m <model>overrides the model for one run, and-c key=valueoverrides any other key.
What does not work
wire_api = "chat". A hard error since February 2026, as above.
Putting the provider in the repository. Codex ignores model_provider, model_providers and openai_base_url in a project's .codex/config.toml, because repository contents should not decide where your credentials go. Put them in your user config or pass them with -c.
openai_base_url for a multi-provider router. It points the built-in OpenAI provider at another host, but that provider still requires an OpenAI sign-in and sends those credentials. Fine for a proxy in front of OpenAI; a router with its own provider keys belongs in its own [model_providers] table.
A Chat Completions URL with wire_api = "responses". The config loads, then the first request goes to /responses on a server that has nothing there.
The in-session /model picker. With a custom provider it has been reported to list only OpenAI's models, and picking one rewrites model in the active config file to an OpenAI name (openai/codex#35487, reported against 0.142.2 and still open). Use -m instead.
Do you need a translator at all?
A Responses request carries instructions and typed input items instead of a messages array, and its stream only counts as finished when a response.completed event arrives. A backend that does not speak this needs both directions converted on every turn, tool calls included (Responses API vs Chat Completions covers the differences). Whether kRouter should do that depends on the backend:
| Your backend | What to do |
|---|---|
| Serves the Responses API itself (the OpenAI API, Azure OpenAI) | Codex's built-in OpenAI provider, or one [model_providers] block for Azure. A router only adds a hop. |
| One local model in Ollama or LM Studio | Codex's own --oss mode. |
| Speaks only Chat Completions or Anthropic Messages, a subscription, or several providers at once | A translator between Codex and the providers. |
The rest of this post covers the third row.
Pointing Codex at kRouter
kRouter accepts Responses requests on /v1/responses and converts them to and from whatever format the provider behind the model speaks.
npm install -g @openai/codex # if Codex is not installed yet
npm install -g @sifxprime/krouter
krouter -tOpen http://localhost:20128/dashboard and connect at least one provider. Then open CLI Tools and the OpenAI Codex CLI / App card. Choose the endpoint, an API key and a model in provider/model-id form, optionally a different subagent model, and click Apply. kRouter merges this into ~/.codex/config.toml:
model = "cc/claude-sonnet-4-6"
model_provider = "krouter"
[model_providers.krouter]
name = "kRouter"
base_url = "http://localhost:20128/v1"
wire_api = "responses"
[model_providers.krouter.http_headers]
Authorization = "Bearer <your-krouter-key>"
[agents]
default_subagent_model = "cc/claude-sonnet-4-6"Your other settings stay, but the file is rewritten, so its comments do not survive. Manual Config on the same card shows the block if you would rather paste it. Because the key travels in http_headers, your ChatGPT login in auth.json is left alone. To keep the key out of the file, use env_key = "KROUTER_API_KEY" instead of the header table and export that variable; applying from the card again puts the header back.
On the same machine, kRouter does not require a key unless you turn on Require API key on the Endpoint page or set REQUIRE_API_KEY=true. If no key exists yet, the card writes the placeholder sk_krouter, which only works locally and only while that requirement is off. From another machine -- Codex on a laptop, kRouter on a VPS or behind its built-in Cloudflare or Tailscale tunnel -- every request needs a real key from the Endpoint page. Without one, kRouter answers 401 "API key required for remote API access".
Before blaming Codex for anything, test the route on its own, streamed the way Codex calls it:
curl -sN http://localhost:20128/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <your-krouter-key>" \
-d '{"model": "cc/claude-sonnet-4-6", "input": "Reply with the word ok", "stream": true}'If text events arrive and the stream ends with a response.completed event, routing and translation work. Then try Codex itself with codex exec "reply with the word ok".
Which models to use
Any model on kRouter's 95+ providers can go in model. A few that suit Codex:
model value | Comes from | What kRouter does |
|---|---|---|
cx/gpt-5.5 | A ChatGPT plan, signed in to kRouter's Codex provider | Forwards the Responses request without translating it |
cc/claude-sonnet-4-6 | A Claude subscription (Claude Code login) | Translates Responses to Claude Messages and back |
gh/claude-sonnet-4.6 | A GitHub Copilot subscription | Translates |
kr/claude-sonnet-4.5 | Kiro, signed in with a device code | Translates |
ocg/kimi-k3, ocg/glm-5.3, ocg/deepseek-v4-pro | An OpenCode Go subscription | Translates to Chat Completions and back |
| A combo name, no slash | Several of the above | Tries the next model when one fails |
Before you connect a Claude, Codex, GitHub Copilot or Kiro account, kRouter shows a risk notice: those subscription sessions are not officially licensed for proxy use, and the account may be restricted or banned. Read it before you decide.
For long sessions a combo is usually the better model: Codex keeps one stable name, and when the first model hits a rate limit or runs out of quota, the next one answers. Do not start the combo's name with a GPT model id (see below). OpenCode Go serves different models on different endpoints; see OpenCode Go's three protocols.
What breaks, and why
Six things behave differently once Codex is not talking to OpenAI.
The context window is a guess. For a name it does not recognise, Codex logs "Unknown model ... This will use fallback model metadata" and assumes a 272,000-token window. If the model you route to has less room, Codex does not compact the conversation in time, and the provider rejects the request instead. Set the real figure:
model_context_window = 128000 # the window of the model you actually use
model_auto_compact_token_limit = 100000 # compact a little before thatIn Codex 0.160.1, an unknown model's window can be lowered this way but not raised above 272,000.
Reasoning effort does not travel. Codex sends model_reasoning_effort in the request's reasoning field. kRouter passes it through on cx/ but drops it when it translates to Claude or Chat Completions. Where a provider exposes thinking as a separate model id, choose it by name instead, such as Kiro's kr/claude-sonnet-4.5-thinking.
Hosted tools stay with OpenAI. Codex's built-in web search, on by default, runs on OpenAI's servers. On a translated route nothing can run it, so kRouter drops it.
MCP tools do not survive translation. Current Codex sends each MCP server's tools as one namespace tool, a Responses-only shape that groups several functions under one name. kRouter passes it through untouched to its Codex provider. On a translated route, as of 0.5.163, the whole group becomes a single function with no parameters, and the individual tools inside it are lost. If a session depends on MCP servers, run it on cx/, or point Codex directly at a backend that supports namespace tools.
GPT model names change what Codex sends. When Codex does not recognise a name, it retries the lookup with one leading provider/ segment removed, and it also matches any name that starts with a model id it knows. So kr/gpt-5.6-sol, or a combo called gpt-5.5-backup, gets OpenAI's settings for that model. In Codex 0.160.1, gpt-5.6-sol and the other GPT-5.6 and GPT-6 models use a newer request layout that lists the tools inside the conversation input instead of the request's tools field. kRouter 0.5.163 does not translate that layout, so on a translated route the model receives no tools at all. gpt-5.5 keeps the usual layout, but its apply_patch is a freeform tool that takes plain text instead of JSON; translated, it arrives without parameters, and the model cannot fill in a patch. Use GPT models in Codex through cx/, and keep the translated routes for Claude, GLM, Kimi and DeepSeek.
Streams that end early. Codex reports a stream that closes before response.completed as "stream disconnected before completion". By default it reconnects up to 5 times (stream_max_retries) and treats 300 seconds without data as a dropped connection (stream_idle_timeout_ms); both can be set in the [model_providers.krouter] table. The causes, and how to tell them apart, are in Codex "stream disconnected before completion".
Pin what works
Codex moves quickly: twelve stable releases between September 18 and October 5, 2026, from 0.155.1 to 0.160.1. Custom providers are not its main path, and updates occasionally break them. After 0.147.0, a LiteLLM user saw every request fail until going back to 0.146.0 (#37425), and an Azure user got Invalid 'input[0].tools[0].description': empty string (#37380). If an update breaks a setup that worked yesterday, go back one version:
npm install -g @openai/codex@<last-version-that-worked>On the kRouter side, run 0.5.163 or newer (npm i -g @sifxprime/krouter@latest). It stopped dropping the developer messages Codex sends when the model is one GitHub Copilot serves only through its Responses API. kRouter does not update itself, so it stays on the version you installed until you run that command.
Common questions
Does Codex CLI still support wire_api = "chat"?
No. Chat Completions support was deprecated in December 2025 and removed in Codex 0.95.0 in February 2026. responses is now the only accepted value, and the default when the key is missing. A provider that only speaks Chat Completions needs a translator in front of it.
Can Codex CLI use Claude models?
Yes, through something that converts the Responses API to Anthropic's Messages API and back, such as kRouter with model = "cc/claude-sonnet-4-6" or gh/claude-sonnet-4.6. Expect four differences from GPT models: you set model_context_window yourself, Codex's reasoning-effort setting does not reach the model, its built-in web search is not available, and as of kRouter 0.5.163 tools from MCP servers do not work through the translation.
Where does the API key go for a custom provider?
In the provider table, as an Authorization header under http_headers or as env_key naming an environment variable. A custom provider does not use auth.json or codex login by default, and there is no api_key field. For kRouter on the same machine the key is optional unless you require one; from any other machine it is required.
Why does Codex warn about fallback model metadata?
It does not recognise the model name, so it uses generic settings, including a 272,000-token context window. If the model has less room, set model_context_window and model_auto_compact_token_limit so Codex compacts before the provider refuses the request.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- How kRouter worksXcode with any model: chat, Claude Agent and CodexXcode 26 and 27 can run chat, Claude Agent and Codex on other models, but each reads its own config. How to point all three at one local router.
- How kRouter worksCopilot BYOK custom endpoint: any model in VS Code chatVS Code's Custom Endpoint lets Copilot Chat use any Chat Completions, Responses or Messages API. Using it with a local router, and what still needs GitHub.
- How kRouter worksJetBrains AI Assistant BYOK: any model through one endpointAI Assistant's OpenAI-compatible provider takes one URL. Put kRouter behind it for chat on any model, and see why Junie and completion stay out.
Relevant docs
- Core conceptsProviders, combos, account routing, token savers, MITM mode, quota tracking, the response cache and remote access: the eight ideas behind kRouter.
- The Zenith Routing EngineHow kRouter picks which of your accounts serves a request: Zenith scoring, conversation stickiness, cooldowns, and the round-robin, P2C and random options.
- Token savers: RTK, Caveman, Ponytail, Headroom, PXPIPEFive ways kRouter can shrink a request before it reaches the provider. Only RTK is on by default. What each saver does, what it needs, and its limits.