Skip to main content
kRouter
All posts
How kRouter works

OpenCode Go in Claude Code and Codex: three protocols

OpenCode Go serves its models on three APIs. Which model needs which, why the wrong one fails with ModelProtocolUnsupported, and how to fix it.

Kodelyth · The team behind kRouter
· Updated
9 min read

You add your OpenCode Go key to a gateway or an OpenAI-compatible client, pick one of the Muse Spark or GPT Luna models, and the first request comes back with this:

400 {"type":"error","error":{"type":"ModelProtocolUnsupported","message":"Model does not support this protocol."}}

GLM and MiniMax on the same key answer normally. The key is valid, the subscription is active, and waiting does not help.

OpenCode Go is three APIs behind one key, and each model is served on only some of them. Claude Code speaks one of the three, Codex another, and a model that lives elsewhere is out of reach until something translates.

What Go costs and what it caps

From OpenCode's Go documentation at the time of writing: Go is $10 a month and Go Plus is $40 a month, with the same models and token prices on both. Go Plus raises the limits.

The limits are dollar amounts, not request counts: at most 20% of the monthly limit in a 5-hour window and 50% in a week. The docs give every model its own monthly figure, which sets how fast that model's usage counts against those windows, so the same plan goes much further on some models than on others. OpenCode's own estimate for the Go plan is about 110 requests per 5 hours on Kimi K3, 169 on Grok 4.7 and 13,000 on DeepSeek V4 Flash.

Three more details from the same page:

  • Over the limit, you can keep using the free models, or turn on Use balance in the OpenCode console so Go draws on your Zen balance instead of blocking.
  • Privacy is per model. The Muse Spark Contributor models are discounted in exchange for using your prompts and completions for training, and are only offered in the regions Meta allows. Grok and GPT Luna keep data for 30 days, and most of the rest keep none.
  • Go expects coding-agent traffic. OpenCode monitors for abuse and asks clients to send a stable x-opencode-session header per conversation. kRouter sends it on all three endpoints; the MissingSessionID post explains why it matters.

Which model lives on which endpoint

All three endpoints share the base https://opencode.ai/zen/go/v1:

EndpointWire formatModels, per OpenCode's docs
/chat/completionsOpenAI Chat CompletionsGLM, Kimi, DeepSeek, MiMo, LongCat, Hy, Space Bunny
/messagesAnthropic MessagesMiniMax, Qwen
/responsesOpenAI ResponsesGrok, GPT Luna, Muse Spark

The third row is the strict one, in both directions: Grok, GPT Luna and Muse Spark fail on Chat Completions, and /responses turns away Kimi and MiniMax. kRouter treats seven models as Responses-only: grok-4.5, grok-4.6, grok-4.7, gpt-5.6-luna, gpt-6-luna, muse-spark-1.2-contributor and muse-spark-1.3-contributor. The line between the first two rows is looser in practice: in a test with a real key, reported on kRouter's issue tracker, MiniMax M3 also answered on /chat/completions. Responses API vs Chat Completions covers how the formats differ.

The live roster at https://opencode.ai/zen/go/v1/models returns 43 model ids, more than the docs table, including older ones such as kimi-k2.5 and glm-5. kRouter lists the same 43. OpenCode renamed Space Bunny to space-bunny when its free promotion ended; kRouter 0.5.163 still listed the old space-bunny-free, which OpenCode no longer accepts, and v0.5.164 uses the new id.

What each client reaches on its own

Claude Code only speaks Anthropic Messages. Codex only speaks the Responses API (the Codex guide explains why). OpenCode lists both as validated clients for Go, so connecting either one directly works, for the models on its endpoint:

Client, connected directlyModels on its endpointModels on the other two
Claude CodeMiniMax, QwenKimi, GLM, DeepSeek, MiMo and the rest of Chat Completions; Grok, GPT Luna, Muse Spark
CodexGrok, GPT Luna, Muse SparkKimi, GLM, DeepSeek, MiMo and the rest of Chat Completions; MiniMax, Qwen
OpenCode itselfEvery Go model--

Kimi, GLM and DeepSeek are documented for Chat Completions only, so do not count on reaching them directly from either client.

If every model you want is on your client's endpoint, connect directly: a router only adds a hop. The rest of this post is for using the whole roster from one client, or mixing Go with other providers.

Two ways OpenCode says no

OpenCode rejects a model on the wrong endpoint in one of two ways. Some requests pass a first check and fail after authentication with the 400 above; GPT Luna and Muse Spark on /chat/completions behave this way. Others are refused before OpenCode looks at the key at all. Sent without a key, a Grok request to the Chat Completions endpoint comes back as:

HTTP 401
{"type":"error","error":{"type":"ModelError","message":"Model grok-4.7 is not supported for format oa-compat"}}

The format name in that message tells you which endpoint refused: oa-compat is /chat/completions, anthropic is /messages, and openai is /responses. A ModelError with no format, such as Model space-bunny-free is not supported, means OpenCode does not know the id at all.

Everywhere else a 401 means bad credentials, and clients and gateways act on it that way. Before 0.5.163, kRouter did too: it handled OpenCode's ModelError as an authentication failure and passed the 401 on to your client.

Because that first check comes before the key check, a request without a key costs nothing and still tells you something:

curl -s https://opencode.ai/zen/go/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "grok-4.7", "messages": [{"role": "user", "content": "hi"}]}'

A ModelError means that endpoint will not serve the model. "Missing API key." only means the request got past the first check; GPT Luna gets exactly that on /chat/completions, then fails there with a real key.

kRouter 0.5.163 changed three things:

  • The seven Responses-only models go to /responses. kRouter converts to and from the Responses API itself, so Claude Code, Codex and plain OpenAI clients can all use them, streaming or not, with tools.
  • OpenCode's ModelError 401 is reported as a 400, a request error, so it is no longer handled as a bad key.
  • Every OpenCode Go model gets a Protocol setting, covered below.

What does not work

A new API key. The key was never the problem; a second key gets the same answer.

Waiting or retrying. A protocol mismatch is not a rate limit. The same request fails the same way every time.

Moving every model to one protocol. No single endpoint serves the whole roster. Sent to /responses, Kimi K3 gets a ModelError for format openai.

Claude Code on every Go model

Install kRouter:

npm install -g @sifxprime/krouter
krouter -t

Open http://localhost:20128/dashboard, go to Providers, choose OpenCode Go and add the API key from the OpenCode console. Then open CLI Tools, pick the Claude Code card, map the Opus, Sonnet and Haiku slots to ocg/ models and click Apply. kRouter writes the settings into ~/.claude/settings.json. The shell equivalent:

export ANTHROPIC_BASE_URL=http://localhost:20128/v1
export ANTHROPIC_AUTH_TOKEN=<your-krouter-key>
export ANTHROPIC_DEFAULT_OPUS_MODEL=ocg/grok-4.7
export ANTHROPIC_DEFAULT_SONNET_MODEL=ocg/minimax-m3
export ANTHROPIC_DEFAULT_HAIKU_MODEL=ocg/deepseek-v4-flash

Those three slots land on three different OpenCode endpoints -- Responses, Messages and Chat Completions -- and Claude Code never sees the difference. It shows the plumbing, not a ranking; which open model holds up best as the main agent is covered in Claude Code on Kimi, GLM and DeepSeek.

Two settings deserve thought:

  • The Haiku slot. Claude Code also uses that model for background work. Give it a cheap model such as DeepSeek V4 Flash, and keep Kimi K3, Grok 4.7 and Qwen 3.8 Max out of it, or background calls will eat into your 5-hour window.
  • The context window. Claude Code picks a window from the model name, and ocg/minimax-m3 is not a name it knows. Set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the real figure for your model. The Claude Code card's Max context dropdown writes it with presets from 200K to 1M; for any other figure, set the variable yourself. See context windows with other models.

On the same machine kRouter does not require a key, and the card writes the placeholder sk_krouter if you have none. From another machine every request needs a real key from the Endpoint page.

Codex on every Go model

On the same CLI Tools page, open the OpenAI Codex CLI / App card, choose an ocg/ model and click Apply. kRouter merges this into ~/.codex/config.toml:

model = "ocg/kimi-k3"
model_provider = "krouter"
 
[model_providers.krouter]
name = "kRouter"
base_url = "http://localhost:20128/v1"
wire_api = "responses"
 
[model_providers.krouter.http_headers]
Authorization = "Bearer <your-krouter-key>"
 
[agents]
default_subagent_model = "ocg/kimi-k3"

Codex always sends Responses requests. kRouter turns them into Chat Completions for Kimi, GLM and DeepSeek, into Messages for MiniMax, and sends the Responses-only models to OpenCode's /responses. For a single run, codex -m ocg/glm-5.3 overrides the model.

Picking a protocol per model

In the dashboard, open Providers, then OpenCode Go. Once you have added a connection, every model row has a Protocol select: Auto (with kRouter's choice in parentheses), Chat Completions, Messages or Responses. The choice applies to all your OpenCode Go connections, including ones you add later.

It only picks the upstream endpoint. Claude Code still sends Messages and Codex still sends Responses; kRouter converts in between.

Leave it on Auto unless a model fails with a protocol error, then read the error:

Error mentionsNot served onSet Protocol to
format oa-compat/chat/completionsResponses for Grok, GPT Luna, Muse Spark; Messages for MiniMax, Qwen
format anthropic/messagesChat Completions, or Responses for the Responses-only models
format openai/responsesChat Completions, or Messages for MiniMax, Qwen
ModelProtocolUnsupported onlyThe endpoint kRouter choseThe endpoint OpenCode's model table lists
is not supported, no formatAny endpoint: the id is unknownNothing; use the current id (see below)

Three cases to watch at the time of writing:

  • Qwen. OpenCode's docs list Qwen 3.8 Max, Qwen 3.8 Flash and Qwen 3.7 Plus on /messages. From v0.5.164 kRouter's Auto sends every Qwen model there; on 0.5.163 Auto sent them to Chat Completions, so if you are on an older version and a Qwen model fails with a protocol error, set it to Messages.
  • Muse Spark. On Go, OpenCode requires you to opt in to the training terms on your workspace's Go page in the console. Without that, a request on the right endpoint gets a 403 saying the model collects data and requires explicit opt-in. That is not a protocol problem, so the Protocol select will not fix it.
  • New or renamed models. For an id kRouter does not list yet (on 0.5.163 that includes space-bunny), use Add Model on the same page. Custom models get the Protocol select too, and Auto sends them to Chat Completions.

To confirm a route through kRouter, send one real request (it counts against your allowance):

curl -s http://localhost:20128/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "ocg/grok-4.7", "messages": [{"role": "user", "content": "Reply with the word ok"}], "max_tokens": 32}'

When Go hits its limit

Once a window is spent, OpenCode answers 429 with a message such as 5-hour usage limit reached. Resets in ... (or the weekly or monthly version), unless Use balance is on. The window runs down fastest on the priciest models.

Put your Go models in a combo with an entry from another provider after them. When a Go entry returns the limit error, the next entry answers, and your client keeps using the combo's name.

kRouter does not show Go's 5-hour and weekly windows on its Quota Tracker page, so watch your usage in the OpenCode console.

Common questions

Can I use OpenCode Go with Claude Code?

Yes. Connected directly, Claude Code reaches the models on OpenCode's Anthropic Messages endpoint, which the docs list as MiniMax and Qwen. For Kimi, GLM, DeepSeek, Grok, GPT Luna and Muse Spark it needs a translator such as kRouter, which sends each request to the endpoint the model lives on.

Why does one OpenCode Go model fail with ModelProtocolUnsupported while others work?

That model is not served on the endpoint the request went to. Grok, GPT Luna and Muse Spark are served only on the Responses API, and OpenCode's docs put every other model on Chat Completions or Messages. Your key and subscription are fine.

Does ModelProtocolUnsupported mean my API key is bad?

No, and neither does the related ModelError, even though OpenCode returns that one as HTTP 401. A new key, a retry or waiting changes nothing. The request has to go to a different endpoint, or, if the message names no format, use a model id OpenCode still serves.

How much does OpenCode Go cost and how do its limits work?

Go is $10 a month and Go Plus is $40. Limits are dollar amounts: at most 20% of the monthly limit per 5 hours and 50% per week, and each model's listed monthly figure sets how fast its usage counts. Kimi K3, Grok 4.7 and Qwen 3.8 Max get far fewer requests than the cheaper models.

Which Protocol should I pick in kRouter?

Auto. Change a model only when it fails with a protocol error, using the format named in the error. The setting changes which OpenCode endpoint kRouter calls, not what your client speaks.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter