Codex "stream disconnected before completion": every cause
Codex fails on a custom provider with "stream disconnected before completion". What the error means, its five causes, and how to tell them apart.
You point Codex at your own endpoint -- a gateway, a proxy, a local router -- and a task that should take a minute ends like this:
stream disconnected before completion: stream closed before response.completedBy then Codex has already retried on its own -- the TUI counts Reconnecting... 1/5 up to 5/5 -- and each retry sent the whole request again.
The usual reaction is to raise the retry count or the timeout. That rarely helps. The error does not name a cause; it says one thing never arrived, and at least five different problems produce that.
What the error actually means
Codex talks to a custom provider only over the OpenAI Responses API (config reference). It sends POST {base_url}/responses and reads the reply as a stream of events. The turn is finished only when one of three terminal events arrives: response.completed, response.incomplete or response.failed.
Anything that ends the stream before that becomes "stream disconnected before completion". Codex treats it as transient and retries, by default up to 5 times (stream_max_retries). The text after the colon is the useful part:
| After the colon | What happened |
|---|---|
stream closed before response.completed | The connection ended and no terminal event had arrived |
idle timeout waiting for SSE | Nothing arrived for stream_idle_timeout_ms (default 300,000 ms, five minutes) |
failed to parse ResponseCompleted: ... | A response.completed arrived, but Codex could not read it |
Incomplete response returned, reason: ... | The backend sent response.incomplete, for example after hitting its output token cap |
websocket closed by server before response.completed | The built-in OpenAI provider's WebSocket transport (openai/codex#18960) |
Two parser details matter later. Lines that are not Responses events are skipped, so Chat Completions chunks and data: [DONE] count for nothing. And response.completed only counts if it carries a response object with an id.
1. The endpoint does not speak the Responses API
Since February 2026, wire_api = "chat" stops Codex at startup (discussion #7782), so every provider block says responses. That sets what Codex sends, not what the server understands.
If {base_url}/responses reaches something that answers 200 without Responses events -- a web page, a JSON body, a Chat Completions stream -- Codex reads no events and reports the stream closed. (A missing route usually gives an HTTP status error instead.)
How to tell: the curl test further down returns HTML, a single JSON object, or chunks containing "choices" instead of lines like event: response.output_text.delta.
The fix: check base_url. For OpenAI-style servers it normally ends in /v1, because Codex appends only /responses. A server that only implements Chat Completions needs a translator in front of it; Responses API vs Chat Completions covers what that involves.
2. The proxy streams, but never finishes properly
Typical of a hand-rolled or partial translator. Text arrives, so the setup looks like it works, then the stream ends without a usable terminal event:
- The proxy sends the last text delta and closes, or ends with
data: [DONE]. That is the Chat Completions sentinel, and Codex ignores it. - It sends
response.completedwithout aresponseobject. Codex does not count that. - Its
response.completedcarries Chat Completions usage (prompt_tokens,completion_tokens) instead ofinput_tokens,output_tokensandtotal_tokens. Codex cannot parse it and says so after the colon.
How to tell: short prompts fail exactly like long ones, and every retry fails in the same way.
The fix: in the proxy, not in Codex. No retry setting can make a terminal event appear that the server never sends.
3. Something in the middle cuts long streams
Short tasks work and long ones die, often after about the same elapsed time:
- A reverse proxy. By default nginx closes the upstream connection after 60 seconds without data (
proxy_read_timeout) and buffers responses (proxy_buffering on) (nginx docs). A reasoning model can be silent for longer than that. - Corporate TLS inspection. In one OpenAI community thread, turning off Zscaler fixed it.
- Load balancers, tunnels and VPNs with their own idle or connection-length limits.
How to tell: the curl test passes with a short prompt, and failures cluster at one duration.
The fix: raise the middlebox's limits or take it out of the path. For nginx in front of a streaming API:
location / {
proxy_pass http://127.0.0.1:20128;
proxy_http_version 1.1;
proxy_buffering off;
proxy_read_timeout 600s;
}4. The backend went quiet or died mid-reply
Sometimes the provider itself stalls, gets overloaded or drops the connection partway through a reply. The retry may succeed. Five minutes of silence shows up as idle timeout waiting for SSE instead.
How to tell: failures are intermittent, they do not track prompt length or a fixed duration, and the curl test passes most of the time.
The fix: make sure the next attempt can go somewhere healthier -- another account, or another model.
5. Codex itself changed
If the error appeared on the day you updated Codex, suspect Codex first. Version 0.147.0 (August 7, 2026) started grouping the default tools under a functions namespace for some models, and some backends reject that request shape.
In openai/codex#37425, a LiteLLM setup in front of Bedrock failed on every request after the upgrade. Bedrock rejected the functions namespace and LiteLLM passed that on as an HTTP 500. Codex logged stream disconnected - retrying sampling request until the retries ran out, then showed "We're currently experiencing high demand". Going back to 0.146.0 fixed it; the issue was still open in early October 2026, with Codex at 0.160.1. A commenter there kept the new version working by exporting the model catalog (codex debug models --bundled), setting use_responses_lite to false for the affected models, and pointing model_catalog_json at the edited file.
That log line misleads in general: Codex writes it for every retry, including retries after an HTTP error, and turns a final 500 into the "high demand" message.
How to tell: the curl test passes, but Codex fails on every request.
The fix: go back to the last version that worked until the issue is fixed:
npm install -g @openai/codex@<last-version-that-worked>What does not work
Raising stream_max_retries. Every retry sends the full request again. With causes 1, 2 and 5 they all fail the same way, and whenever the backend actually ran the request, as in cause 2, each attempt is billed again.
Raising stream_idle_timeout_ms when the message says "stream closed". That timeout only covers silence. A closed connection is not silence, so a longer timeout never comes into play.
Running codex login again. A custom provider normally takes its key from its own block (env_key, http_headers or env_http_headers), not from your ChatGPT login. A bad key shows up as a 401, not as this error.
Working out which one you hit
Send a minimal version of the request Codex makes. Replace <base_url>, the key and the model with the values from your provider block in ~/.codex/config.toml (for a local kRouter, <base_url> is http://localhost:20128/v1):
curl -sN "<base_url>/responses" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <your-key>" \
-d '{"model": "<your-model>", "input": "Reply with the word ok", "stream": true}' \
| tail -n 6A healthy endpoint ends with a data: line carrying "type":"response.completed" and a response object that has an id. Codex reads that type field, not the event: line, and a trailing data: [DONE] is harmless. To see Codex's side, run codex exec "reply with the word ok", or start the TUI with codex -c log_dir=./.codex-log to get a plaintext codex-tui.log.
| You see | Likely cause | Do this |
|---|---|---|
HTML, one JSON object, or "choices" chunks | Wrong base_url, or not a Responses endpoint (1) | Fix the URL or put a translator in front |
Text events, then the stream ends with no response.completed | Proxy never finishes the stream (2) | Fix or replace the proxy |
| curl passes; long Codex tasks fail at a similar duration | A middlebox cutting the stream (3) | Raise its timeouts, turn off buffering |
| curl passes; failures are random | Unstable backend (4) | Add fallback accounts or models |
| curl passes; every Codex request fails since an upgrade | Codex regression (5) | Pin the last working version |
If kRouter is the proxy
kRouter accepts Responses requests on /v1/responses and converts them for whatever the model behind the name speaks: Claude Messages, Chat Completions or Gemini. On those translated routes it writes the Responses events itself, and a reply that finishes ends with a response.completed that carries a response id. If a Chat Completions backend closes its stream without a finish reason, kRouter still ends the reply with response.completed. That covers causes 1 and 2.
The OpenAI Codex CLI / App card under CLI Tools writes the provider block for you: base_url ending in /v1, wire_api = "responses" and the key as an Authorization header. The Codex setup guide walks through it.
If the error still appears with kRouter in the path, look at the backend behind it:
- kRouter has its own watchdog. If an upstream stream sends no bytes for 60 seconds, kRouter aborts it; some subscription backends, including
cx/(Codex),gh/(GitHub Copilot) andkr/(Kiro), get three minutes. Runkrouter -land these show up asstream stall timeoutin the log. - A dead upstream looks like a dead connection. If the backend drops or stalls partway through a reply, no
response.completedreaches Codex. On a translated route the stream just ends. Where the backend speaks Responses itself, such ascx/, kRouter passes events through and, if they stop before a terminal event, adds aresponse.failedwith the codestream_disconnectedand the message "stream closed before response.completed". Either way Codex shows the same wording as for a dropped connection and retries, so check the kRouter log to see which side dropped.
kRouter cannot continue a half-sent reply from another account. It can answer the retry: if a provider refuses it with a rate limit or server error, kRouter moves on to another account, or to the next model in a combo.
Behind nginx or a tunnel, cause 3 applies to kRouter too; the deploy guide has a streaming-safe nginx block.
If you call the OpenAI API or Azure directly, a router only adds a hop. If the error started with a Codex update, pin the version instead.
Common questions
What does "stream disconnected before completion" mean in Codex?
The response stream ended before a terminal event such as response.completed arrived. The text after the colon says how: "stream closed" means the connection ended; "idle timeout waiting for SSE" means no data for stream_idle_timeout_ms, five minutes by default.
Why is it more common with a custom provider?
OpenAI's own backend produces it too, when a connection drops. But Codex speaks only the Responses API to custom providers, and many proxies implement that stream only partly. If they skip the closing response.completed, or send one Codex cannot parse, every turn ends with this error even though text arrived.
Should I increase stream_max_retries or stream_idle_timeout_ms?
Only after you know the cause. Retries send the whole request again, so a structural problem fails every time, and each attempt the backend runs is billed. The idle timeout only covers silence, not a closed connection.
Why does the Codex log say "stream disconnected" when the backend returned an error?
Codex logs stream disconnected - retrying sampling request for every retry, whatever the cause. An HTTP 500 is retried the same way, and when the retries run out Codex shows "We're currently experiencing high demand". Read the sampling_error field on that log line, or run the curl test.
Does kRouter fix it?
It removes the translation causes: kRouter speaks Responses to Codex and ends every finished reply with a response.completed Codex can parse. It does not fix a reverse proxy cutting long streams, a provider dying mid-reply, or a Codex regression. When a provider refuses the retry, it can send it to another account, or to the next model if you use a combo.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- Fix an errorClaude Code "rate limit exceeded": every cause and the fixClaude Code stops with a 429 or a usage limit. Four different limits cause it, each needs a different fix, and the one everyone tries first makes it worse.
- Fix an errorClaude Code 529 Overloaded: who sent it and how to fail overA 529 is not your quota: the model is out of capacity, and Claude Code has already retried. How to tell who sent it, and how to keep working.
- Fix an errorClaude Code context window with non-Claude models: the fixClaude Code assumes 200K for a model it doesn't recognize, so it compacts too early or overruns the real limit. How to declare the real window.
Relevant docs
- TroubleshootingFixes for common kRouter problems: banned or rate-limited accounts, OAuth sign-ins, MITM certificates, localhost and Cursor, Docker logins and the CLI.
- Error referenceWhat the common AI provider errors mean, why they happen, and how to get past them.
- Combos & fallbackPut several models behind one name. kRouter falls back from one to the next, rotates them, or asks a panel of models and merges the answers.