JetBrains AI Assistant BYOK: any model through one endpoint
AI Assistant's OpenAI-compatible provider takes one URL. Put kRouter behind it for chat on any model, and see why Junie and completion stay out.
You open AI Chat in IntelliJ IDEA or PyCharm, click OpenAI-compatible, LM Studio, Ollama, choose OpenAI-compatible as the provider, and the form asks for three things: a Base URL, a key and a model. You have more than one source in mind -- Kimi from an OpenCode Go plan, a Claude subscription for the hard questions, something cheap for commit messages -- and the form takes one URL.
That one URL can be a router with every provider behind it. What no URL can do is reach all of AI Assistant: chat and the editor features yes, the built-in agents and code completion no.
What BYOK reaches, and what it does not
JetBrains shipped BYOK in December 2025. With your own key, AI Assistant works without a JetBrains AI subscription, and JetBrains says the keys are stored on your machine and never shared with it. The 2026.2 documentation splits what an OpenAI-compatible endpoint can do:
- AI Chat. Yes. The endpoint's models appear in their own section of the chat model selector.
- Editor and VCS features. Yes, once you assign a model. Models from OpenAI-compatible endpoints are not matched to features automatically; you pick them under Models Assignment: Core features (in-editor code generation, commit messages, the default chat model) and Instant helpers (chat titles, context collection, name suggestions).
- Code completion and next edit suggestions. No. These run on JetBrains models by default, Mellum for completion. A separate AI Completion section takes an OpenAI-compatible model, but only one that supports fill-in-the-middle or edit prediction; JetBrains notes that general chat models usually don't.
- Junie, Claude Agent and Codex. No. The agent activation page says Junie takes keys issued directly by OpenAI or Anthropic, and keys from other providers "are not supported" even when they serve OpenAI or Anthropic models. Claude Agent takes an Anthropic key or Anthropic Console; Codex takes an OpenAI key or a ChatGPT login. The OpenAI and Anthropic provider entries those keys go into have no URL field.
- Managed organizations. If your company runs AI through JetBrains IDE Services or JetBrains Central, an administrator can stop you from connecting third-party models at all.
So an OpenAI-compatible endpoint, kRouter included, gives you chat, the editor features and the VCS helpers. Agents need another route, covered below.
Adding kRouter as the provider
kRouter is a self-hosted router: one OpenAI-compatible endpoint in front of 95+ providers, with logins, API keys and free tiers mixed. AI Assistant sends Chat Completions requests; kRouter converts them to whatever each provider speaks -- Claude Messages, the Responses API, Gemini, Kiro -- and converts the reply back.
1. Install, start and connect providers.
npm install -g @sifxprime/krouter
krouter -tOpen http://localhost:20128/dashboard and connect providers on the Providers page. kRouter listens on every network interface by default; if only this machine uses it, start it with krouter -t --host 127.0.0.1.
2. Make a key for the IDE. The Endpoint page creates a Default Key on your first visit; use Create Key in its API Keys section to give the IDE one of its own, which you can pause without touching your other tools. Local requests need no key unless Require API key is on. If kRouter runs on another machine, the key is mandatory: remote callers without one get a 401.
3. Check the model list before the IDE does.
export KROUTER_KEY="<your-krouter-key>" # from the Endpoint page
curl -s http://localhost:20128/v1/models -H "Authorization: Bearer $KROUTER_KEY"
curl -s http://localhost:20128/v1/chat/completions \
-H "Authorization: Bearer $KROUTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"ocg/kimi-k2.6","messages":[{"role":"user","content":"Reply with OK"}],"stream":false}'/v1/models lists the chat models of the providers you connected, plus your combos. Names follow provider/model: ocg/kimi-k2.6 is OpenCode Go, cc/claude-sonnet-4-6 is Claude Code. A name without a slash is a combo or an alias. If this request fails, the problem is kRouter or the provider, not the IDE.
4. Add the endpoint in the IDE. There are two doors to the same setting:
- First run. In AI Chat, click OpenAI-compatible, LM Studio, Ollama, set Provider to OpenAI-compatible, fill in Base URL, Key and Model, and click Continue. Enter one kRouter model name.
- Settings. Settings | Tools | AI Assistant | Providers & API keys, section Third-party AI providers, provider OpenAI-compatible: the URL, the key, and a Tool calling switch. Click Test Connection, then Apply.
The URL is http://localhost:20128/v1. Keep the /v1: kRouter's API lives under it, and kRouter also accepts a doubled /v1/v1, so the address works whether or not the IDE appends a /v1 of its own. Model names with a slash are fine.
localhost is fine here. JetBrains' settings reference says the provider URL can point to a local server such as http://localhost:port, which is also how the Ollama and LM Studio providers connect. Cursor is the counterexample: it calls custom base URLs from its own servers and never reaches your localhost.
If Test Connection fails, run the curl above first; if that fails too, fix kRouter or the provider before touching the IDE. If curl works on your machine but you use remote development or a dev container, run the same curl from that environment's terminal. If it fails there, give the IDE an address that environment can reach, start kRouter without --host 127.0.0.1, and use a key.
5. Assign models to features. Under Models Assignment, choose a kRouter model for Core features and another for Instant helpers. The same section has a Context window field, 64,000 tokens by default, which JetBrains describes for local models; AI Assistant also trims attachments that exceed a percentage of the model's window. If you change the value, keep it at or below the real window of the model behind the entry -- for a combo, its smallest model.
Choosing what goes where
| AI Assistant slot | What to give it | Why |
|---|---|---|
| Chat | A combo, or your best model | One picker entry with fallback behind it |
| Core features | A strong model | Code generation and commit messages are judged on quality |
| Instant helpers | A cheap fast model, such as ocg/deepseek-v4-flash | Titles and name suggestions run often and need little |
| Tool calling switch | On only if every model behind the entry handles tools | It tells AI Assistant the model can call MCP tools |
| AI Completion | Leave it on JetBrains, or a FIM model served directly | kRouter has no fill-in-the-middle endpoint |
A combo is the useful trick here. Its default strategy, fallback, tries its models in order, moving those with more remaining quota first where kRouter tracks quota. When the first model answers with a rate limit or another error worth retrying, kRouter sends the same request to the next one, and AI Assistant only ever sees one model name.
Using models you already pay for
The built-in providers in AI Assistant take API keys: Anthropic, OpenAI and Google (the Gemini API or Vertex AI). Plenty of models sit behind a login or a plan instead, and kRouter can sign in to several: Claude Code, OpenAI Codex, GitHub Copilot, Kiro, Grok CLI.
Know the risk first. kRouter's provider pages for Claude Code, Codex, Kiro, Gemini CLI, Qoder, Antigravity and GitHub Copilot carry a Risk Notice: those subscription sessions are not licensed for proxy use, and the account can be restricted or banned. API-key and plan providers such as OpenCode Go, GLM or MiniMax carry no such notice.
Several of these providers are reached through a Responses API behind the scenes: Codex, Grok CLI, the Copilot models GitHub serves only that way (its Codex models), and seven OpenCode Go models (its Grok, GPT Luna and Muse Spark entries). If you route to any of them, run kRouter 0.5.163 or later (krouter -v prints the version). Before that release, a reply cut off by the token limit ended without a finish reason, a failed reply could come back as a successful one with the error as its text, only the first system message reached the model, and the seven OpenCode Go models failed outright. Format translation explained covers why conversions like this go wrong.
RTK, the only token saver on by default, compresses tool results such as file reads and search output. That matters for an agent like OpenCode, below, and for AI Chat only when Tool calling is on. The README's figure of 20-40% fewer input tokens is a project claim, not a measurement of your sessions.
Agents on kRouter models: go through ACP
The integrated agents are closed to an OpenAI-compatible endpoint, but AI Chat also runs agents over the Agent Client Protocol, and those bring their own model configuration. JetBrains' ACP documentation says they need no JetBrains AI subscription.
OpenCode is the simplest pairing. kRouter's CLI Tools page has an OpenCode card that writes ~/.config/opencode/opencode.json: a krouter provider with your base URL, key and the models you choose. OpenCode is also in JetBrains' ACP registry, but to run the copy you just configured, add it as a custom agent: in AI Chat, click the button in the upper-right corner, choose Add Custom Agent, and the IDE creates and opens ~/.jetbrains/acp.json:
{
"agent_servers": {
"OpenCode": {
"command": "/absolute/path/to/opencode",
"args": ["acp"]
}
}
}which opencode prints the path. OpenCode's documentation says it works the same over ACP as in the terminal, so it uses the models kRouter wrote. JetBrains does not support ACP agents under WSL.
If what you want is Junie itself on other models, the standalone Junie CLI is the one that allows it. It reads custom model profiles from ~/.junie/models/, and baseUrl is the complete endpoint, path included:
{
"baseUrl": "http://localhost:20128/v1/chat/completions",
"id": "ocg/kimi-k2.6",
"apiType": "OpenAICompletion",
"apiKey": "${KROUTER_KEY}",
"fasterModel": { "id": "ocg/deepseek-v4-flash" }
}Saved as ~/.junie/models/krouter.json, it runs with junie --model custom:krouter. Export KROUTER_KEY in the shell that starts Junie: if the variable is not set, the profile fails to load.
GitHub Copilot in JetBrains
If you use the GitHub Copilot plugin rather than AI Assistant, GitHub added OpenAI-compatible custom endpoints with API keys to it on July 14, 2026. kRouter's /v1 endpoint and a key fit that form. The VS Code version of that setup covers model ids and token limits in more detail.
When kRouter is not the answer
- You have one Anthropic or OpenAI key. Use AI Assistant's built-in provider for it. An Anthropic key also activates Junie and Claude Agent, and an OpenAI key Junie and Codex, which an OpenAI-compatible endpoint never will.
- You only run local models. The built-in Ollama and LM Studio providers connect directly.
- You want your own model for inline completion. Point the AI Completion section straight at a server that serves a FIM model.
A router earns its place when you mix sources -- logins and keys, several accounts for one provider -- or already run one for Claude Code, Codex or OpenCode and want the IDE on the same models.
Common questions
Do I need a JetBrains AI subscription to use BYOK?
No. With your own key, AI Chat and the features your provider's models support work without one. JetBrains suggests activating JetBrains AI as well if you want every feature, because its models then fill in where yours do not.
Does http://localhost work in AI Assistant?
Yes. JetBrains' settings reference says the provider URL can point to a local server such as http://localhost:port, the way the Ollama and LM Studio providers connect, so http://localhost:20128/v1 works on the machine where kRouter runs. Cursor, by contrast, sends these requests from its own servers.
Can Junie use kRouter?
Not the Junie built into AI Chat: JetBrains accepts only keys issued directly by OpenAI or Anthropic for it. The standalone Junie CLI can, through a custom model profile whose baseUrl is kRouter's full /v1/chat/completions URL.
Why is a model missing from the chat model selector?
The models AI Assistant offers come from your endpoint, so start with curl http://localhost:20128/v1/models. kRouter lists only chat models from providers you have connected and active, plus your combos. A model you disabled on its provider page drops out as well.
Can kRouter power inline code completion?
No. Inline completion needs a fill-in-the-middle model and next edit suggestions need an edit-prediction model, and kRouter routes chat models: it has no fill-in-the-middle endpoint. Leave AI Completion on JetBrains, or point it directly at a server that serves such a model.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- How kRouter worksCopilot BYOK custom endpoint: any model in VS Code chatVS Code's Custom Endpoint lets Copilot Chat use any Chat Completions, Responses or Messages API. Using it with a local router, and what still needs GitHub.
- How kRouter worksKiro IDE with any model through a local MITM proxyKiro IDE has no custom endpoint setting. kRouter's MITM mode sends its chat turns to any model you map. How it works, the v0.5.137 fix, and setup.
- How kRouter worksCodex CLI with any model: model_providers and wire_apiCodex now speaks only the Responses API. How to run it on Claude, GLM, Kimi or DeepSeek through config.toml, and what breaks along the way.
Relevant docs
- Core conceptsProviders, combos, account routing, token savers, MITM mode, quota tracking, the response cache and remote access: the eight ideas behind kRouter.
- The Zenith Routing EngineHow kRouter picks which of your accounts serves a request: Zenith scoring, conversation stickiness, cooldowns, and the round-robin, P2C and random options.
- Token savers: RTK, Caveman, Ponytail, Headroom, PXPIPEFive ways kRouter can shrink a request before it reaches the provider. Only RTK is on by default. What each saver does, what it needs, and its limits.