Skip to main content
kRouter
All posts
How kRouter works

Copilot BYOK custom endpoint: any model in VS Code chat

VS Code's Custom Endpoint lets Copilot Chat use any Chat Completions, Responses or Messages API. Using it with a local router, and what still needs GitHub.

Kodelyth · The team behind kRouter
· Updated
9 min read

You want Copilot Chat in VS Code to answer with a model GitHub does not offer: Kimi K2.6 from an OpenCode Go plan, GLM, or a Claude subscription you already pay for. Under Manage Language Models there is a provider called Custom Endpoint that asks for an API key and an API type, then opens a JSON file.

Two things decide whether it works well: the API type you pick and the two token numbers in that file. Part of Copilot stays out of reach whatever you configure.

What changed in 2026

  • April 22, 2026. GitHub made bring-your-own-key available in VS Code for Copilot Business and Enterprise. Usage is billed by your provider and "does not count against GitHub Copilot request quotas."
  • VS Code 1.122, May 28, 2026. The release notes brought the Custom Endpoint provider, which speaks Chat Completions, Responses and Anthropic Messages, to Stable, and let BYOK work without a GitHub sign-in. VS Code's documentation says it replaces the deprecated OpenAI Compatible provider and its github.copilot.chat.customOAIModels setting.

The release notes also say: "Requests go directly to your provider." VS Code sends them itself, so with kRouter on the same machine http://localhost:20128 works with no tunnel. Cursor calls custom base URLs from its own servers and never reaches your localhost.

What a custom endpoint does not cover

  • Inline suggestions and next edit suggestions. GitHub says "BYOK does not apply to code completions", and VS Code's language model documentation says you "cannot connect to a local model for inline suggestions."
  • Semantic search and features built on embeddings. These still need a GitHub account.
  • Background jobs, until you assign a model. Chat titles, commit messages and pull request descriptions run on utility models that come with a GitHub sign-in. Signed out, those defaults are unreachable until you point chat.utilityModel and chat.utilitySmallModel at your own models.
  • Organizations that turned it off. On Business and Enterprise, an admin can disable the Bring Your Own Language Model Key in VS Code policy.

Setting up the endpoint with kRouter

kRouter serves all three APIs that Custom Endpoint speaks -- /v1/chat/completions, /v1/responses and /v1/messages -- on one port, and converts between them and whatever each provider speaks. One endpoint in VS Code then reaches every provider you have connected.

1. Install, start and connect providers.

npm install -g @sifxprime/krouter
krouter -t

Open http://localhost:20128/dashboard and connect providers on the Providers page. kRouter listens on every network interface by default; if only VS Code on this machine uses it, start it with krouter -t --host 127.0.0.1.

2. Make a key for VS Code. Use Create Key in the API Keys section of the Endpoint page. Local requests need no key unless you turn on Require API key in that section or set REQUIRE_API_KEY=true, but Custom Endpoint asks for one, and a separate key can be deactivated on its own. If kRouter runs on another machine, the key is mandatory: remote callers without one get a 401, "API key required for remote API access".

3. Find the model names and prove the path. kRouter names models provider/model, such as ocg/kimi-k2.6 (OpenCode Go) or cc/claude-sonnet-4-6 (Claude Code); a name without a slash is a combo or an alias. Then send one request per format you plan to use, with the header VS Code will send: Authorization: Bearer for Chat Completions and Responses, x-api-key for Messages. A failure here is kRouter or the provider, not VS Code.

KROUTER_KEY="<your-krouter-key>"   # from the Endpoint page
curl -s http://localhost:20128/v1/models -H "Authorization: Bearer $KROUTER_KEY"
 
curl -s http://localhost:20128/v1/chat/completions \
  -H "Authorization: Bearer $KROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"ocg/kimi-k2.6","messages":[{"role":"user","content":"Reply with OK"}],"stream":false}'
 
curl -s http://localhost:20128/v1/messages \
  -H "x-api-key: $KROUTER_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"cc/claude-sonnet-4-6","max_tokens":64,"messages":[{"role":"user","content":"Reply with OK"}]}'
 
curl -s http://localhost:20128/v1/responses \
  -H "Authorization: Bearer $KROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"cx/gpt-5.5","input":"Reply with OK","stream":false}'

Each reply should come back in the format you sent: a choices array for Chat Completions, a content array for Messages, an output array for Responses.

4. Add the endpoint in VS Code. Run Chat: Manage Language Models from the Command Palette, select Add Models, then Custom Endpoint. Enter a group name, a display name, your kRouter key and an API type. VS Code opens chatLanguageModels.json. kRouter's CLI Tools page does not write this file, so fill in the models yourself:

[
  {
    "name": "kRouter",
    "vendor": "customendpoint",
    "apiKey": "${input:krouterKey}",
    "apiType": "chat-completions",
    "models": [
      {
        "id": "ocg/kimi-k2.6",
        "name": "Kimi K2.6 (kRouter)",
        "url": "http://localhost:20128/v1/chat/completions",
        "toolCalling": true,
        "vision": false,
        "maxInputTokens": 229376,
        "maxOutputTokens": 32768
      },
      {
        "id": "cc/claude-sonnet-4-6",
        "name": "Claude Sonnet 4.6 (kRouter)",
        "url": "http://localhost:20128/v1/messages",
        "apiType": "messages",
        "toolCalling": true,
        "vision": true,
        "maxInputTokens": 136000,
        "maxOutputTokens": 64000
      },
      {
        "id": "coding",
        "name": "Coding combo (kRouter)",
        "url": "http://localhost:20128/v1/chat/completions",
        "toolCalling": true,
        "vision": false,
        "maxInputTokens": 168000,
        "maxOutputTokens": 32000
      }
    ]
  }
]

The parts that are easy to get wrong:

  • url is the full endpoint, path included, as VS Code's documentation recommends.
  • id is the kRouter model name exactly as /v1/models lists it; name is only the picker label. coding stands for a combo you built on the Combos page.
  • A model's apiType overrides the group's. That is how one group mixes Claude on Messages with the rest on Chat Completions. Change the url path with it.
  • toolCalling and vision are promises. Mark a combo as tool-calling only if every model in it handles tools; turn vision on only after an image request through that model works.
  • Keep the apiKey line VS Code wrote. It is an ${input:...} reference to the key in VS Code's secret storage; krouterKey above stands for its id.

Which API type to pick

Any of the three reaches any model. kRouter converts through the OpenAI chat format and skips conversion when the request already matches what the provider speaks, so for Chat Completions and Messages the choice only sets how many conversions happen.

Behind the model idapiTypeWhy
Claude through Claude Code (cc/...)messageskRouter talks to Claude in Messages format
GPT through OpenAI Codex (cx/...)responses, with the flag belowCodex is a Responses API
OpenCode Go's Kimi, GLM, DeepSeekchat-completionskRouter sends them to OpenCode on Chat Completions
OpenCode Go's MiniMax modelsmessageskRouter sends them to OpenCode on Messages
OpenCode Go's Grok, GPT Luna, Muse Sparkchat-completionsResponses-only at OpenCode, but kRouter handles them as chat models and converts at the last step
A combo, or anything you are unsure aboutchat-completionsThe pivot format, so the fewest conversions

Watch Qwen. OpenCode's documentation lists Qwen 3.8 Max, 3.8 Flash and 3.7 Plus on its Messages endpoint, while kRouter's Auto choice sends Qwen to Chat Completions. If a Qwen model fails with a protocol error, set its Protocol to Messages on the OpenCode Go provider page, as OpenCode Go's three protocols explains.

If you pick responses, add "zeroDataRetentionEnabled": true to that model's entry. On the Responses API, VS Code chains turns: it sends the last reply's id as previous_response_id plus only the messages after it, and expects the server to have stored the rest. kRouter's reply ids have the usual resp_ form, so VS Code chains them, but kRouter does not store responses. The model sees only the newest part of the conversation, and a tool result can arrive without the call it answers. With the flag set, VS Code sends the whole conversation each time.

Run kRouter 0.5.163 or later (krouter -v prints the version) if you pick responses or use any model kRouter reaches through a Responses API: Codex, GitHub Copilot, Grok CLI, or OpenCode Go's Grok, GPT Luna and Muse Spark. That release fixed truncated replies with no finish reason, failures returned as successes, dropped developer and later system messages, and non-streaming /v1/responses requests answered in chat format. Responses API vs Chat Completions explains why these formats trip up proxies.

Getting the token numbers right

VS Code uses the sum of maxInputTokens and maxOutputTokens as the model's context window, and its documentation says the sum must not exceed the real one. Too high and long sessions send more than the model accepts; too low and you waste context.

GET /v1/models reports context_length and max_completion_tokens for the chat models of your connected providers. Three caveats:

  • They can be a default. Without specific data for a model, kRouter reports 200,000 and 64,000. Check the provider's documentation for models you depend on.
  • The output limit can equal the whole window. kRouter lists 262,144 for both on Kimi K2.6. Pick an output limit and subtract: 262,144 minus 32,768 gives the 229,376 above. Custom Endpoint also accepts contextWindow in place of maxInputTokens; give it the full window and VS Code subtracts maxOutputTokens itself.
  • Combos list no limits. Use the smallest window among the combo's models; the example assumes 200,000.

kRouter lists Claude Sonnet 4.6 at 1,000,000 tokens; the example uses a conservative 200,000. Raise it once a long session works at the size you choose.

Using what you already pay for

VS Code's built-in providers cover API-key services such as Anthropic, Azure, Gemini, OpenAI and OpenRouter. Many models sit behind a login instead. kRouter puts 95+ providers behind one endpoint, logins such as Claude Code, OpenAI Codex, Kiro and Grok CLI included.

Know the risk first. kRouter's provider pages for Claude Code, Codex, Kiro, Gemini CLI, Qoder, Antigravity and GitHub Copilot carry a Risk Notice: those subscription sessions are not licensed for proxy use, and the account can be restricted or banned. API-key providers carry no such notice.

Three things help here:

  • A combo as one model. A combo's default strategy, fallback, tries its models in order, moving those with more remaining quota first where kRouter tracks quota. One picker entry keeps answering when its first model is rate limited.
  • Cheap utility models. Signed out, you have to assign chat.utilityModel and chat.utilitySmallModel anyway. Pick an inexpensive kRouter model, so titles and commit messages do not spend your best model's quota.
  • RTK. Agent mode sends terminal output, search results and file reads back as tool results. RTK, the only token saver on by default, compresses that output. kRouter's README claims 20-40% fewer input tokens, a project figure rather than a measurement of your sessions.

When you do not need kRouter

  • One provider that VS Code has built in. Use the built-in provider with your key.
  • Only Ollama. The built-in Ollama provider is deprecated; install the official Ollama extension.
  • One compatible service. Point Custom Endpoint straight at it.
  • Your own model for inline completions. Nothing here helps; BYOK does not reach completions.

A router is worth it when you mix sources -- logins and keys, several accounts for one provider -- or already run one for Claude Code, Codex or Cline and want VS Code on the same models.

The older route: intercepting Copilot

kRouter's MITM mode can also reroute Copilot's own chat requests: it points api.individual.githubcopilot.com at your machine through the hosts file, installs a local root certificate and swaps in the model you choose. Business and Enterprise seats use other hosts, which kRouter does not intercept, and MITM needs admin or sudo rights and does not work in Docker. Custom Endpoint needs none of that. If you still want interception, setup is on the Copilot integration page.

Common questions

Do I need a Copilot subscription to use a custom endpoint?

No. Since VS Code 1.122, BYOK models work without a GitHub sign-in or a Copilot plan. Signed out, you lose inline suggestions, next edit suggestions, semantic search and embedding-based features, and you assign your own utility models. Signed in or not, BYOK never powers inline completions.

Does localhost work, unlike in Cursor?

Yes. VS Code sends BYOK requests itself, straight to the URL you configure, so with kRouter on the same machine http://localhost:20128/v1/chat/completions works without a tunnel. Cursor sends its requests from its own servers, which cannot reach your localhost.

Why doesn't my model show up for agents?

Its entry needs "toolCalling": true. Also check that the id matches a name from kRouter's /v1/models exactly, and restart VS Code if a newly added model is missing.

Why does a model on the Responses type forget earlier turns?

VS Code chains Responses requests with previous_response_id and sends only the newest messages, expecting the server to remember the rest, and kRouter does not store responses. Add "zeroDataRetentionEnabled": true to that model's entry so VS Code sends the whole conversation, or use Chat Completions.

Does BYOK spend my Copilot AI Credits?

No. Copilot's monthly plans moved to GitHub AI Credits on June 1, 2026, but GitHub says BYOK usage is billed directly by your chosen provider and does not count against Copilot's request quotas. The requests go from VS Code to your endpoint, not through GitHub. You pay whatever the provider behind each model charges. Background tasks such as chat titles are separate: signed in, they run on GitHub's utility models unless you assign your own.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter