Skip to main content
kRouter
All posts
How kRouter works

Responses API vs Chat Completions: why proxies break

Codex, some Copilot models and OpenCode Go's Grok speak only OpenAI's Responses API. Five ways it differs from Chat Completions, and where proxies break.

Kodelyth · The team behind kRouter
· Updated
10 min read

You put a gateway in front of your models and every Chat Completions client works. Then you point Codex at it, and every turn ends like this:

stream disconnected before completion: stream closed before response.completed

Or you send a Grok model to OpenCode Go's Chat Completions endpoint, and it refuses before it even checks your key:

HTTP 401
{"type":"error","error":{"type":"ModelError","message":"Model grok-4.7 is not supported for format oa-compat"}}

Both come from the same place. OpenAI has two APIs for chat models, Chat Completions and Responses, and they look close enough that a proxy can treat one as a renamed copy of the other. It cannot. Past the renamed fields, they differ in five places: how a stream ends, where instructions go, how a reply reports being cut off or failing, how tool calls are shaped, and who keeps the history.

Two formats for the same conversation

Chat Completions (POST /v1/chat/completions) takes messages and streams back deltas. Responses (POST /v1/responses) takes instructions plus typed input items, and streams back typed events. OpenAI's migration guide says Chat Completions "remains supported" and recommends Responses for new projects.

Chat CompletionsResponses
Conversationmessagesinput: a string, or a list of items
Instructionssystem or developer messagesinstructions, plus system or developer items in input
Tool definition{type, function: {name, parameters}}{type, name, parameters}
Tool call and resulttool_calls on an assistant message, then a tool messagefunction_call and function_call_output items, matched by call_id
Output capmax_tokens or max_completion_tokensmax_output_tokens
JSON moderesponse_formattext.format
Reasoning effortreasoning_effortreasoning.effort
Streamdelta chunks, then data: [DONE]typed events, ending in a terminal event
How it endedfinish_reasonstatus, plus incomplete_details.reason
Token countsprompt_tokens, completion_tokensinput_tokens, output_tokens
Historyresent with every requestresent, or referenced with previous_response_id

Here is one request in both shapes: what an OpenAI client sends kRouter, then what kRouter 0.5.163 sends OpenCode Go, which serves Grok only on its Responses endpoint:

{
  "model": "ocg/grok-4.7",
  "messages": [
    { "role": "system", "content": "You are a coding agent." },
    { "role": "developer", "content": "Never edit files outside src/." },
    { "role": "user", "content": "What is in src/?" }
  ],
  "tools": [{ "type": "function", "function": { "name": "list_dir",
    "parameters": { "type": "object", "properties": { "path": { "type": "string" } } } } }],
  "max_tokens": 2000,
  "stream": true
}
{
  "model": "grok-4.7",
  "instructions": "You are a coding agent.\n\nNever edit files outside src/.",
  "input": [
    { "type": "message", "role": "user",
      "content": [{ "type": "input_text", "text": "What is in src/?" }] }
  ],
  "tools": [{ "type": "function", "name": "list_dir", "description": "",
    "parameters": { "type": "object", "properties": { "path": { "type": "string" } } } }],
  "max_output_tokens": 2000,
  "store": false,
  "stream": true
}

1. How the stream ends

A Chat Completions stream is a run of chat.completion.chunk objects. The last one sets finish_reason, and data: [DONE] closes the stream. Usage comes in one extra chunk, only if the client asked with stream_options.include_usage.

A Responses stream is a run of typed events (response.created, response.output_text.delta, response.function_call_arguments.delta and so on). The reply is over only when a terminal event arrives: response.completed, response.incomplete, response.failed, or a top-level error event. Codex does not treat [DONE] as the end of a turn.

So a proxy that answers Codex with chat chunks gives it nothing it can read, since a chunk without a type is skipped. One that sends Responses deltas and then [DONE] gets the text on screen and still no finished turn. Either way Codex reports "stream disconnected before completion" (how to confirm it with curl).

The reverse direction breaks too. Until 0.5.163, when kRouter served a chat client from a Responses backend, it ignored response.incomplete, which is how a reply cut off by the output cap ends, so the client got no finish reason and no usage, and Claude Code's stream never received its message_stop. A cut-off Grok reply now converts like this (trimmed):

event: response.output_text.delta
data: {"type":"response.output_text.delta","delta":"The src/ folder holds"}
 
event: response.incomplete
data: {"type":"response.incomplete","response":{"status":"incomplete",
  "incomplete_details":{"reason":"max_output_tokens"},
  "usage":{"input_tokens":412,"output_tokens":16,"total_tokens":428}}}
data: {"object":"chat.completion.chunk","choices":[{"delta":{"content":"The src/ folder holds"},"finish_reason":null}]}
 
data: {"object":"chat.completion.chunk","choices":[{"delta":{},"finish_reason":"length"}],
  "usage":{"prompt_tokens":412,"completion_tokens":16,"total_tokens":428}}

A stream that stops with no terminal event at all, because the connection dropped, now gets a closing chunk too.

2. System and developer instructions

In Chat Completions, instructions are messages: system, or developer, which OpenAI's SDK documents as replacing system for o1 and newer models.

Responses has a top-level instructions field, and system and developer messages can also sit inside input. Codex sends developer messages there alongside its instructions.

A chat-to-Responses converter has to place every one of them. Before 0.5.163, kRouter's converter kept the first system message as instructions and dropped the rest, so Codex's developer messages never reached models kRouter calls through a Responses API, such as some of GitHub Copilot's GPT models. Now all of them are joined into instructions, in order, as in the example above. Toward a Chat Completions backend, kRouter rewrites developer to system, because not every OpenAI-compatible server accepts the newer role.

3. Endings, usage and failures

Chat Completions reports why a reply ended in finish_reason: stop, length, tool_calls or content_filter. Responses reports a status such as completed, incomplete or failed, and an incomplete reply names the reason in incomplete_details.reason, for example max_output_tokens. kRouter maps the common endings like this:

The Responses stream ends withAn OpenAI client getsA Claude client gets
response.completed after textfinish_reason: "stop"stop_reason: "end_turn"
response.completed after a tool call"tool_calls""tool_use"
response.incomplete, reason max_output_tokens"length""max_tokens"

Failures are the dangerous part. A Responses backend that fails mid-reply has usually sent HTTP 200 already, so it reports the failure inside the stream, as response.failed or an error event. A proxy that turns that into ordinary text hands the client a successful reply made of an error message, and an agent may carry on as if it were an answer. Before 0.5.163, kRouter did that to non-streaming clients. They now get a real error status: 429 for a rate limit, 502 otherwise. A streaming client still sees the error as text, because its 200 went out before the failure.

4. Tool calls as separate items

In Chat Completions, tool calls sit in a tool_calls array on an assistant message, and each streamed fragment names its call by index. Results go back as tool messages carrying the tool_call_id.

In Responses, a call and its result are separate items linked by a call_id. Each call is its own function_call item, its argument fragments arrive tagged with that item's id, and the result goes back as a function_call_output item. Two kRouter bugs show what goes wrong in between:

  • Parallel calls merged into one (fixed in 0.5.150). A backend often opens all its parallel calls before finishing any. A converter that moves to the next chat index only when a call finishes puts every fragment on index 0, so the client gets one tool call whose input is several JSON objects stuck together. kRouter now keys each call by its item id.
  • Text cut off by an empty tool list (fixed in 0.5.138). A chat backend can send tool_calls: [] next to ordinary text. An empty array counts as true in JavaScript, so the converter took it for a tool call and closed the text early, cutting off what Codex was reading.

5. Who keeps the history

With Chat Completions, the client sends the whole conversation every time. Responses stores each response on OpenAI's side by default, so the next request can send only previous_response_id plus the new input.

A proxy that does not store responses cannot honour that, and kRouter does not store them. It also removes previous_response_id before a request reaches Codex or Grok CLI, which could not resolve it either. The client has to send the full conversation every turn, or the model sees only the newest part. VS Code's Copilot Chat chains turns this way on its Responses type; the custom endpoint guide shows the setting that makes it resend everything.

Why Codex, Copilot and Grok are where it shows

Most tools are fine with a chat-only gateway. Three places force the Responses API on you:

  • Codex as a client. OpenAI announced in December 2025 that Codex would drop Chat Completions (discussion #7782), and current Codex refuses to load a provider set to wire_api = "chat". Every Codex request is a Responses request, whichever model is behind it.
  • GitHub Copilot as a backend. Copilot serves some OpenAI models only on /responses. When its chat endpoint answers 400 that a model is "not accessible via the /chat/completions endpoint", kRouter resends the request to /responses and remembers that for the model until kRouter restarts. Claude models go to Copilot's Messages endpoint instead.
  • Grok through OpenCode Go or Grok CLI. OpenCode Go serves its Grok, GPT Luna and Muse Spark models only on /responses (OpenCode docs). On /chat/completions they fail: OpenCode answers with the 401 ModelError from the top of this page before it looks at the key, and with a valid key a 400 ModelProtocolUnsupported has been reported. Until 0.5.163, kRouter sent them there and read that 401 as a bad key, so it treated a request it could never serve as an account problem and put the model on a cooldown. Since 0.5.163 they go to /responses, and the ModelError is reported as a 400. Grok CLI's backend is a Responses API as well.

What does not work

Renaming the endpoint. Routing /responses to a chat handler gives Codex chunks it cannot parse and no terminal event, [DONE] or not. A chat-only server without that route just answers 404.

Retrying. Every failure above is structural: a retry goes through the same conversion and fails the same way, and if the backend ran it, you pay again.

How kRouter handles both

kRouter accepts each format on its own endpoint: /v1/chat/completions, /v1/responses (with /codex/... as an alias) and Anthropic's /v1/messages, all listed in the API reference. For most pairs, Chat Completions is the pivot: a request is converted to chat shape, then to whatever the provider speaks, and the reply makes the same trip back.

Client speaksBackend speaksWhat kRouter does
Chat CompletionsChat CompletionsNo format change; developer becomes system
Responses (Codex)Chat Completions (most providers)Converts the request to chat, and the stream back to Responses events
Chat Completions or ClaudeResponses only (Codex, grok-cli, some Copilot GPT models, OpenCode Go's Grok)Converts to Responses at the last step, and maps the terminal events back
ResponsesResponses (cx/, gcli/)Leaves the format as it is

For Codex, the OpenAI Codex CLI / App card under CLI Tools writes a [model_providers.krouter] block to ~/.codex/config.toml with wire_api = "responses" and a base_url ending in /v1; Codex CLI with any model covers which models behave well there. For OpenCode Go, each model on the provider page has a Protocol control (Auto, Chat Completions, Messages or Responses), so you can follow OpenCode the day it moves a model (details). For your own server that only speaks Responses, use Add OpenAI Compatible on the Providers page with API Type set to Responses API.

The fixes in sections 1 to 3 shipped in 0.5.163. Check with krouter -v, and update with:

npm install -g @sifxprime/krouter@latest

On Docker, pull sifxprime/krouter:latest and recreate the container with the same data volume, so your settings and password carry over. To see the conversion work, send a non-streaming Responses request with a tight output cap (no key is needed from the machine kRouter runs on). Set "stream": false explicitly: kRouter streams when the field is missing.

curl -s http://localhost:20128/v1/responses \
  -H "Content-Type: application/json" \
  -d '{"model": "ocg/grok-4.7", "input": "Count from 1 to 200.", "max_output_tokens": 16, "stream": false}'

Without OpenCode Go, use a model from any OpenAI-compatible provider you have connected. A current kRouter answers with "object": "response", "status": "incomplete", an incomplete_details.reason of max_output_tokens, and usage including total_tokens. Before 0.5.163, this request to an OpenAI-compatible provider came back as a chat.completion, which a Responses client cannot read.

When kRouter is not the answer

  • You call OpenAI's models from OpenAI's SDK. Use the Responses API directly; a proxy only adds a hop.
  • Your app relies on stored responses. previous_response_id and other server-side state need a server that stores responses, such as OpenAI's.
  • You depend on OpenAI's hosted tools, such as file search, code interpreter or remote MCP. Chat Completions has no equivalent, so a route that converts through it drops them.

Common questions

What is the difference between the Responses API and Chat Completions?

Chat Completions takes messages and streams deltas ending in data: [DONE]. Responses takes instructions plus typed input items, streams typed events that end with a terminal event such as response.completed, keeps tool calls and results as separate items linked by call_id, and can store conversation state on the server.

Is Chat Completions deprecated?

No. OpenAI's migration guide says it remains supported, while recommending Responses for new projects. Codex is the exception that matters for proxies: it no longer accepts wire_api = "chat" for custom providers, so anything Codex talks to has to speak Responses.

Can a proxy convert between them without losing anything?

For text, images, function tools, streaming and usage, yes, if it maps every difference above, terminal events and failures included. Stored conversation state and OpenAI's hosted tools, such as file search or code interpreter, cannot be converted, because they run on OpenAI's servers.

Why does my client get a successful reply whose text is an error?

The backend failed after the stream had started, and the proxy passed the failure on as text. kRouter 0.5.163 and later give non-streaming clients a real error status instead: 429 for rate limits, 502 otherwise. A streaming client still sees text, because its 200 was already sent.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter