Skip to main content
kRouter
All posts
Fix an error

Codex "stream disconnected before completion": every cause

Codex fails on a custom provider with "stream disconnected before completion". What the error means, its five causes, and how to tell them apart.

Kodelyth · The team behind kRouter
· Updated
9 min read

You point Codex at your own endpoint -- a gateway, a proxy, a local router -- and a task that should take a minute ends like this:

stream disconnected before completion: stream closed before response.completed

By then Codex has already retried on its own -- the TUI counts Reconnecting... 1/5 up to 5/5 -- and each retry sent the whole request again.

The usual reaction is to raise the retry count or the timeout. That rarely helps. The error does not name a cause; it says one thing never arrived, and at least five different problems produce that.

What the error actually means

Codex talks to a custom provider only over the OpenAI Responses API (config reference). It sends POST {base_url}/responses and reads the reply as a stream of events. The turn is finished only when one of three terminal events arrives: response.completed, response.incomplete or response.failed.

Anything that ends the stream before that becomes "stream disconnected before completion". Codex treats it as transient and retries, by default up to 5 times (stream_max_retries). The text after the colon is the useful part:

After the colonWhat happened
stream closed before response.completedThe connection ended and no terminal event had arrived
idle timeout waiting for SSENothing arrived for stream_idle_timeout_ms (default 300,000 ms, five minutes)
failed to parse ResponseCompleted: ...A response.completed arrived, but Codex could not read it
Incomplete response returned, reason: ...The backend sent response.incomplete, for example after hitting its output token cap
websocket closed by server before response.completedThe built-in OpenAI provider's WebSocket transport (openai/codex#18960)

Two parser details matter later. Lines that are not Responses events are skipped, so Chat Completions chunks and data: [DONE] count for nothing. And response.completed only counts if it carries a response object with an id.

1. The endpoint does not speak the Responses API

Since February 2026, wire_api = "chat" stops Codex at startup (discussion #7782), so every provider block says responses. That sets what Codex sends, not what the server understands.

If {base_url}/responses reaches something that answers 200 without Responses events -- a web page, a JSON body, a Chat Completions stream -- Codex reads no events and reports the stream closed. (A missing route usually gives an HTTP status error instead.)

How to tell: the curl test further down returns HTML, a single JSON object, or chunks containing "choices" instead of lines like event: response.output_text.delta.

The fix: check base_url. For OpenAI-style servers it normally ends in /v1, because Codex appends only /responses. A server that only implements Chat Completions needs a translator in front of it; Responses API vs Chat Completions covers what that involves.

2. The proxy streams, but never finishes properly

Typical of a hand-rolled or partial translator. Text arrives, so the setup looks like it works, then the stream ends without a usable terminal event:

  • The proxy sends the last text delta and closes, or ends with data: [DONE]. That is the Chat Completions sentinel, and Codex ignores it.
  • It sends response.completed without a response object. Codex does not count that.
  • Its response.completed carries Chat Completions usage (prompt_tokens, completion_tokens) instead of input_tokens, output_tokens and total_tokens. Codex cannot parse it and says so after the colon.

How to tell: short prompts fail exactly like long ones, and every retry fails in the same way.

The fix: in the proxy, not in Codex. No retry setting can make a terminal event appear that the server never sends.

3. Something in the middle cuts long streams

Short tasks work and long ones die, often after about the same elapsed time:

  • A reverse proxy. By default nginx closes the upstream connection after 60 seconds without data (proxy_read_timeout) and buffers responses (proxy_buffering on) (nginx docs). A reasoning model can be silent for longer than that.
  • Corporate TLS inspection. In one OpenAI community thread, turning off Zscaler fixed it.
  • Load balancers, tunnels and VPNs with their own idle or connection-length limits.

How to tell: the curl test passes with a short prompt, and failures cluster at one duration.

The fix: raise the middlebox's limits or take it out of the path. For nginx in front of a streaming API:

location / {
    proxy_pass         http://127.0.0.1:20128;
    proxy_http_version 1.1;
    proxy_buffering    off;
    proxy_read_timeout 600s;
}

4. The backend went quiet or died mid-reply

Sometimes the provider itself stalls, gets overloaded or drops the connection partway through a reply. The retry may succeed. Five minutes of silence shows up as idle timeout waiting for SSE instead.

How to tell: failures are intermittent, they do not track prompt length or a fixed duration, and the curl test passes most of the time.

The fix: make sure the next attempt can go somewhere healthier -- another account, or another model.

5. Codex itself changed

If the error appeared on the day you updated Codex, suspect Codex first. Version 0.147.0 (August 7, 2026) started grouping the default tools under a functions namespace for some models, and some backends reject that request shape.

In openai/codex#37425, a LiteLLM setup in front of Bedrock failed on every request after the upgrade. Bedrock rejected the functions namespace and LiteLLM passed that on as an HTTP 500. Codex logged stream disconnected - retrying sampling request until the retries ran out, then showed "We're currently experiencing high demand". Going back to 0.146.0 fixed it; the issue was still open in early October 2026, with Codex at 0.160.1. A commenter there kept the new version working by exporting the model catalog (codex debug models --bundled), setting use_responses_lite to false for the affected models, and pointing model_catalog_json at the edited file.

That log line misleads in general: Codex writes it for every retry, including retries after an HTTP error, and turns a final 500 into the "high demand" message.

How to tell: the curl test passes, but Codex fails on every request.

The fix: go back to the last version that worked until the issue is fixed:

npm install -g @openai/codex@<last-version-that-worked>

What does not work

Raising stream_max_retries. Every retry sends the full request again. With causes 1, 2 and 5 they all fail the same way, and whenever the backend actually ran the request, as in cause 2, each attempt is billed again.

Raising stream_idle_timeout_ms when the message says "stream closed". That timeout only covers silence. A closed connection is not silence, so a longer timeout never comes into play.

Running codex login again. A custom provider normally takes its key from its own block (env_key, http_headers or env_http_headers), not from your ChatGPT login. A bad key shows up as a 401, not as this error.

Working out which one you hit

Send a minimal version of the request Codex makes. Replace <base_url>, the key and the model with the values from your provider block in ~/.codex/config.toml (for a local kRouter, <base_url> is http://localhost:20128/v1):

curl -sN "<base_url>/responses" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <your-key>" \
  -d '{"model": "<your-model>", "input": "Reply with the word ok", "stream": true}' \
  | tail -n 6

A healthy endpoint ends with a data: line carrying "type":"response.completed" and a response object that has an id. Codex reads that type field, not the event: line, and a trailing data: [DONE] is harmless. To see Codex's side, run codex exec "reply with the word ok", or start the TUI with codex -c log_dir=./.codex-log to get a plaintext codex-tui.log.

You seeLikely causeDo this
HTML, one JSON object, or "choices" chunksWrong base_url, or not a Responses endpoint (1)Fix the URL or put a translator in front
Text events, then the stream ends with no response.completedProxy never finishes the stream (2)Fix or replace the proxy
curl passes; long Codex tasks fail at a similar durationA middlebox cutting the stream (3)Raise its timeouts, turn off buffering
curl passes; failures are randomUnstable backend (4)Add fallback accounts or models
curl passes; every Codex request fails since an upgradeCodex regression (5)Pin the last working version

If kRouter is the proxy

kRouter accepts Responses requests on /v1/responses and converts them for whatever the model behind the name speaks: Claude Messages, Chat Completions or Gemini. On those translated routes it writes the Responses events itself, and a reply that finishes ends with a response.completed that carries a response id. If a Chat Completions backend closes its stream without a finish reason, kRouter still ends the reply with response.completed. That covers causes 1 and 2.

The OpenAI Codex CLI / App card under CLI Tools writes the provider block for you: base_url ending in /v1, wire_api = "responses" and the key as an Authorization header. The Codex setup guide walks through it.

If the error still appears with kRouter in the path, look at the backend behind it:

  • kRouter has its own watchdog. If an upstream stream sends no bytes for 60 seconds, kRouter aborts it; some subscription backends, including cx/ (Codex), gh/ (GitHub Copilot) and kr/ (Kiro), get three minutes. Run krouter -l and these show up as stream stall timeout in the log.
  • A dead upstream looks like a dead connection. If the backend drops or stalls partway through a reply, no response.completed reaches Codex. On a translated route the stream just ends. Where the backend speaks Responses itself, such as cx/, kRouter passes events through and, if they stop before a terminal event, adds a response.failed with the code stream_disconnected and the message "stream closed before response.completed". Either way Codex shows the same wording as for a dropped connection and retries, so check the kRouter log to see which side dropped.

kRouter cannot continue a half-sent reply from another account. It can answer the retry: if a provider refuses it with a rate limit or server error, kRouter moves on to another account, or to the next model in a combo.

Behind nginx or a tunnel, cause 3 applies to kRouter too; the deploy guide has a streaming-safe nginx block.

If you call the OpenAI API or Azure directly, a router only adds a hop. If the error started with a Codex update, pin the version instead.

Common questions

What does "stream disconnected before completion" mean in Codex?

The response stream ended before a terminal event such as response.completed arrived. The text after the colon says how: "stream closed" means the connection ended; "idle timeout waiting for SSE" means no data for stream_idle_timeout_ms, five minutes by default.

Why is it more common with a custom provider?

OpenAI's own backend produces it too, when a connection drops. But Codex speaks only the Responses API to custom providers, and many proxies implement that stream only partly. If they skip the closing response.completed, or send one Codex cannot parse, every turn ends with this error even though text arrived.

Should I increase stream_max_retries or stream_idle_timeout_ms?

Only after you know the cause. Retries send the whole request again, so a structural problem fails every time, and each attempt the backend runs is billed. The idle timeout only covers silence, not a closed connection.

Why does the Codex log say "stream disconnected" when the backend returned an error?

Codex logs stream disconnected - retrying sampling request for every retry, whatever the cause. An HTTP 500 is retried the same way, and when the retries run out Codex shows "We're currently experiencing high demand". Read the sampling_error field on that log line, or run the curl test.

Does kRouter fix it?

It removes the translation causes: kRouter speaks Responses to Codex and ends every finished reply with a response.completed Codex can parse. It does not fix a reverse proxy cutting long streams, a provider dying mid-reply, or a Codex regression. When a provider refuses the retry, it can send it to another account, or to the next model if you use a combo.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter