OpenCode Go in Claude Code and Codex: three protocols
OpenCode Go serves its models on three APIs. Which model needs which, why the wrong one fails with ModelProtocolUnsupported, and how to fix it.
You add your OpenCode Go key to a gateway or an OpenAI-compatible client, pick one of the Muse Spark or GPT Luna models, and the first request comes back with this:
400 {"type":"error","error":{"type":"ModelProtocolUnsupported","message":"Model does not support this protocol."}}GLM and MiniMax on the same key answer normally. The key is valid, the subscription is active, and waiting does not help.
OpenCode Go is three APIs behind one key, and each model is served on only some of them. Claude Code speaks one of the three, Codex another, and a model that lives elsewhere is out of reach until something translates.
What Go costs and what it caps
From OpenCode's Go documentation at the time of writing: Go is $10 a month and Go Plus is $40 a month, with the same models and token prices on both. Go Plus raises the limits.
The limits are dollar amounts, not request counts: at most 20% of the monthly limit in a 5-hour window and 50% in a week. The docs give every model its own monthly figure, which sets how fast that model's usage counts against those windows, so the same plan goes much further on some models than on others. OpenCode's own estimate for the Go plan is about 110 requests per 5 hours on Kimi K3, 169 on Grok 4.7 and 13,000 on DeepSeek V4 Flash.
Three more details from the same page:
- Over the limit, you can keep using the free models, or turn on Use balance in the OpenCode console so Go draws on your Zen balance instead of blocking.
- Privacy is per model. The Muse Spark Contributor models are discounted in exchange for using your prompts and completions for training, and are only offered in the regions Meta allows. Grok and GPT Luna keep data for 30 days, and most of the rest keep none.
- Go expects coding-agent traffic. OpenCode monitors for abuse and asks clients to send a stable
x-opencode-sessionheader per conversation. kRouter sends it on all three endpoints; the MissingSessionID post explains why it matters.
Which model lives on which endpoint
All three endpoints share the base https://opencode.ai/zen/go/v1:
| Endpoint | Wire format | Models, per OpenCode's docs |
|---|---|---|
/chat/completions | OpenAI Chat Completions | GLM, Kimi, DeepSeek, MiMo, LongCat, Hy, Space Bunny |
/messages | Anthropic Messages | MiniMax, Qwen |
/responses | OpenAI Responses | Grok, GPT Luna, Muse Spark |
The third row is the strict one, in both directions: Grok, GPT Luna and Muse Spark fail on Chat Completions, and /responses turns away Kimi and MiniMax. kRouter treats seven models as Responses-only: grok-4.5, grok-4.6, grok-4.7, gpt-5.6-luna, gpt-6-luna, muse-spark-1.2-contributor and muse-spark-1.3-contributor. The line between the first two rows is looser in practice: in a test with a real key, reported on kRouter's issue tracker, MiniMax M3 also answered on /chat/completions. Responses API vs Chat Completions covers how the formats differ.
The live roster at https://opencode.ai/zen/go/v1/models returns 43 model ids, more than the docs table, including older ones such as kimi-k2.5 and glm-5. kRouter lists the same 43. OpenCode renamed Space Bunny to space-bunny when its free promotion ended; kRouter 0.5.163 still listed the old space-bunny-free, which OpenCode no longer accepts, and v0.5.164 uses the new id.
What each client reaches on its own
Claude Code only speaks Anthropic Messages. Codex only speaks the Responses API (the Codex guide explains why). OpenCode lists both as validated clients for Go, so connecting either one directly works, for the models on its endpoint:
| Client, connected directly | Models on its endpoint | Models on the other two |
|---|---|---|
| Claude Code | MiniMax, Qwen | Kimi, GLM, DeepSeek, MiMo and the rest of Chat Completions; Grok, GPT Luna, Muse Spark |
| Codex | Grok, GPT Luna, Muse Spark | Kimi, GLM, DeepSeek, MiMo and the rest of Chat Completions; MiniMax, Qwen |
| OpenCode itself | Every Go model | -- |
Kimi, GLM and DeepSeek are documented for Chat Completions only, so do not count on reaching them directly from either client.
If every model you want is on your client's endpoint, connect directly: a router only adds a hop. The rest of this post is for using the whole roster from one client, or mixing Go with other providers.
Two ways OpenCode says no
OpenCode rejects a model on the wrong endpoint in one of two ways. Some requests pass a first check and fail after authentication with the 400 above; GPT Luna and Muse Spark on /chat/completions behave this way. Others are refused before OpenCode looks at the key at all. Sent without a key, a Grok request to the Chat Completions endpoint comes back as:
HTTP 401
{"type":"error","error":{"type":"ModelError","message":"Model grok-4.7 is not supported for format oa-compat"}}The format name in that message tells you which endpoint refused: oa-compat is /chat/completions, anthropic is /messages, and openai is /responses. A ModelError with no format, such as Model space-bunny-free is not supported, means OpenCode does not know the id at all.
Everywhere else a 401 means bad credentials, and clients and gateways act on it that way. Before 0.5.163, kRouter did too: it handled OpenCode's ModelError as an authentication failure and passed the 401 on to your client.
Because that first check comes before the key check, a request without a key costs nothing and still tells you something:
curl -s https://opencode.ai/zen/go/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "grok-4.7", "messages": [{"role": "user", "content": "hi"}]}'A ModelError means that endpoint will not serve the model. "Missing API key." only means the request got past the first check; GPT Luna gets exactly that on /chat/completions, then fails there with a real key.
kRouter 0.5.163 changed three things:
- The seven Responses-only models go to
/responses. kRouter converts to and from the Responses API itself, so Claude Code, Codex and plain OpenAI clients can all use them, streaming or not, with tools. - OpenCode's
ModelError401 is reported as a 400, a request error, so it is no longer handled as a bad key. - Every OpenCode Go model gets a Protocol setting, covered below.
What does not work
A new API key. The key was never the problem; a second key gets the same answer.
Waiting or retrying. A protocol mismatch is not a rate limit. The same request fails the same way every time.
Moving every model to one protocol. No single endpoint serves the whole roster. Sent to /responses, Kimi K3 gets a ModelError for format openai.
Claude Code on every Go model
Install kRouter:
npm install -g @sifxprime/krouter
krouter -tOpen http://localhost:20128/dashboard, go to Providers, choose OpenCode Go and add the API key from the OpenCode console. Then open CLI Tools, pick the Claude Code card, map the Opus, Sonnet and Haiku slots to ocg/ models and click Apply. kRouter writes the settings into ~/.claude/settings.json. The shell equivalent:
export ANTHROPIC_BASE_URL=http://localhost:20128/v1
export ANTHROPIC_AUTH_TOKEN=<your-krouter-key>
export ANTHROPIC_DEFAULT_OPUS_MODEL=ocg/grok-4.7
export ANTHROPIC_DEFAULT_SONNET_MODEL=ocg/minimax-m3
export ANTHROPIC_DEFAULT_HAIKU_MODEL=ocg/deepseek-v4-flashThose three slots land on three different OpenCode endpoints -- Responses, Messages and Chat Completions -- and Claude Code never sees the difference. It shows the plumbing, not a ranking; which open model holds up best as the main agent is covered in Claude Code on Kimi, GLM and DeepSeek.
Two settings deserve thought:
- The Haiku slot. Claude Code also uses that model for background work. Give it a cheap model such as DeepSeek V4 Flash, and keep Kimi K3, Grok 4.7 and Qwen 3.8 Max out of it, or background calls will eat into your 5-hour window.
- The context window. Claude Code picks a window from the model name, and
ocg/minimax-m3is not a name it knows. SetCLAUDE_CODE_MAX_CONTEXT_TOKENSto the real figure for your model. The Claude Code card's Max context dropdown writes it with presets from 200K to 1M; for any other figure, set the variable yourself. See context windows with other models.
On the same machine kRouter does not require a key, and the card writes the placeholder sk_krouter if you have none. From another machine every request needs a real key from the Endpoint page.
Codex on every Go model
On the same CLI Tools page, open the OpenAI Codex CLI / App card, choose an ocg/ model and click Apply. kRouter merges this into ~/.codex/config.toml:
model = "ocg/kimi-k3"
model_provider = "krouter"
[model_providers.krouter]
name = "kRouter"
base_url = "http://localhost:20128/v1"
wire_api = "responses"
[model_providers.krouter.http_headers]
Authorization = "Bearer <your-krouter-key>"
[agents]
default_subagent_model = "ocg/kimi-k3"Codex always sends Responses requests. kRouter turns them into Chat Completions for Kimi, GLM and DeepSeek, into Messages for MiniMax, and sends the Responses-only models to OpenCode's /responses. For a single run, codex -m ocg/glm-5.3 overrides the model.
Picking a protocol per model
In the dashboard, open Providers, then OpenCode Go. Once you have added a connection, every model row has a Protocol select: Auto (with kRouter's choice in parentheses), Chat Completions, Messages or Responses. The choice applies to all your OpenCode Go connections, including ones you add later.
It only picks the upstream endpoint. Claude Code still sends Messages and Codex still sends Responses; kRouter converts in between.
Leave it on Auto unless a model fails with a protocol error, then read the error:
| Error mentions | Not served on | Set Protocol to |
|---|---|---|
format oa-compat | /chat/completions | Responses for Grok, GPT Luna, Muse Spark; Messages for MiniMax, Qwen |
format anthropic | /messages | Chat Completions, or Responses for the Responses-only models |
format openai | /responses | Chat Completions, or Messages for MiniMax, Qwen |
ModelProtocolUnsupported only | The endpoint kRouter chose | The endpoint OpenCode's model table lists |
is not supported, no format | Any endpoint: the id is unknown | Nothing; use the current id (see below) |
Three cases to watch at the time of writing:
- Qwen. OpenCode's docs list Qwen 3.8 Max, Qwen 3.8 Flash and Qwen 3.7 Plus on
/messages. From v0.5.164 kRouter's Auto sends every Qwen model there; on 0.5.163 Auto sent them to Chat Completions, so if you are on an older version and a Qwen model fails with a protocol error, set it to Messages. - Muse Spark. On Go, OpenCode requires you to opt in to the training terms on your workspace's Go page in the console. Without that, a request on the right endpoint gets a 403 saying the model collects data and requires explicit opt-in. That is not a protocol problem, so the Protocol select will not fix it.
- New or renamed models. For an id kRouter does not list yet (on 0.5.163 that includes
space-bunny), use Add Model on the same page. Custom models get the Protocol select too, and Auto sends them to Chat Completions.
To confirm a route through kRouter, send one real request (it counts against your allowance):
curl -s http://localhost:20128/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ocg/grok-4.7", "messages": [{"role": "user", "content": "Reply with the word ok"}], "max_tokens": 32}'When Go hits its limit
Once a window is spent, OpenCode answers 429 with a message such as 5-hour usage limit reached. Resets in ... (or the weekly or monthly version), unless Use balance is on. The window runs down fastest on the priciest models.
Put your Go models in a combo with an entry from another provider after them. When a Go entry returns the limit error, the next entry answers, and your client keeps using the combo's name.
kRouter does not show Go's 5-hour and weekly windows on its Quota Tracker page, so watch your usage in the OpenCode console.
Common questions
Can I use OpenCode Go with Claude Code?
Yes. Connected directly, Claude Code reaches the models on OpenCode's Anthropic Messages endpoint, which the docs list as MiniMax and Qwen. For Kimi, GLM, DeepSeek, Grok, GPT Luna and Muse Spark it needs a translator such as kRouter, which sends each request to the endpoint the model lives on.
Why does one OpenCode Go model fail with ModelProtocolUnsupported while others work?
That model is not served on the endpoint the request went to. Grok, GPT Luna and Muse Spark are served only on the Responses API, and OpenCode's docs put every other model on Chat Completions or Messages. Your key and subscription are fine.
Does ModelProtocolUnsupported mean my API key is bad?
No, and neither does the related ModelError, even though OpenCode returns that one as HTTP 401. A new key, a retry or waiting changes nothing. The request has to go to a different endpoint, or, if the message names no format, use a model id OpenCode still serves.
How much does OpenCode Go cost and how do its limits work?
Go is $10 a month and Go Plus is $40. Limits are dollar amounts: at most 20% of the monthly limit per 5 hours and 50% per week, and each model's listed monthly figure sets how fast its usage counts. Kimi K3, Grok 4.7 and Qwen 3.8 Max get far fewer requests than the cheaper models.
Which Protocol should I pick in kRouter?
Auto. Change a model only when it fails with a protocol error, using the format named in the error. The setting changes which OpenCode endpoint kRouter calls, not what your client speaks.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- How kRouter worksXcode with any model: chat, Claude Agent and CodexXcode 26 and 27 can run chat, Claude Agent and Codex on other models, but each reads its own config. How to point all three at one local router.
- How kRouter worksClaude Cowork with other models: third-party inference setupClaude Desktop can run Cowork through any Messages API gateway. Setting it up with kRouter, what the app assumes about your models, and what you give up.
- How kRouter worksKilo Code codebase indexing: picking an embedding providerWhat Kilo's indexer sends, which embedding model to use for code, why changing it means a full re-index, and how to route it all through one endpoint.
Relevant docs
- Core conceptsProviders, combos, account routing, token savers, MITM mode, quota tracking, the response cache and remote access: the eight ideas behind kRouter.
- The Zenith Routing EngineHow kRouter picks which of your accounts serves a request: Zenith scoring, conversation stickiness, cooldowns, and the round-robin, P2C and random options.
- Token savers: RTK, Caveman, Ponytail, Headroom, PXPIPEFive ways kRouter can shrink a request before it reaches the provider. Only RTK is on by default. What each saver does, what it needs, and its limits.