Kilo Code codebase indexing: picking an embedding provider
What Kilo's indexer sends, which embedding model to use for code, why changing it means a full re-index, and how to route it all through one endpoint.
You turn on codebase indexing in Kilo Code and pick OpenAI-Compatible as the embedding provider, so you can use a model Kilo has no preset for. You save, and the status goes to Error:
Cannot determine vector dimension for model "gemini/gemini-embedding-001" with provider "openai-compatible". Please set the model dimension explicitly.That takes ten seconds to fix. Kilo knows the vector sizes of the preset models of its built-in providers, and of nothing behind an OpenAI-compatible endpoint, so you type the number in.
The harder questions sit around it: which model to index code with, what happens to the index when you change it, and how to get a large repository through its first index without being cut off halfway. This post answers them for Kilo Code, including when routing through kRouter helps and when it adds nothing.
What the indexer sends, and how often
Kilo parses files locally with Tree-sitter and cuts them into blocks (functions, classes, methods) of up to about 1,000 characters. It skips binaries, files over 1 MB, node_modules, .git, vendor, and anything matched by .gitignore or .kilocodeignore. Nothing happens until you turn on Enable globally or Enable for this project.
Each block goes to the embedding provider in batches of 60 (embeddingBatchSize, 10 to 200), and the vectors are stored on your machine, in LanceDB by default or in a Qdrant server. The traffic has two phases:
- The first index sends every block in the repository. On a large codebase that is a long burst, and it is where rate limits bite.
- After that, a file watcher re-embeds only changed files, and each
semantic_searchthe agent runs embeds one query with the same model.
So the first index decides whether a free tier is enough.
Choosing an embedding model for code
| Model | Native size | Cost (vendor pages, October 2026) | Notes |
|---|---|---|---|
gemini-embedding-2 | 3072, reducible to 128 | Free tier; $0.20 per 1M text tokens on paid | Generally available since April 22, 2026; 8,192 input tokens; Google's named replacement for 001 |
gemini-embedding-001 | 3072, reducible to 128 | No longer on Google's pricing page | Text only, 2,048 input tokens; earliest shutdown May 14, 2028 |
voyage-code-4 | 1024 (also 256, 512, 2048) | 200M free tokens per account, then $0.12 per 1M | Voyage's current code model; Kilo's preset, voyage-code-3, is on Voyage's older-models list at $0.18 per 1M |
text-embedding-3-small | 1536 | $0.02 per 1M tokens | Kilo's default for its OpenAI provider |
nomic-embed-text (Ollama) | 768 | Your hardware | Nothing leaves the machine |
For a new index on a free tier, gemini-embedding-2 is the sensible default, provided your code may go to a free tier at all; the privacy section below covers that. Google recommends 768, 1536 or 3072 dimensions for both Gemini models; 768 takes a quarter of the space of 3072. For gemini-embedding-2, Google also suggests a task prefix inside the text, such as task: code retrieval | query: ... for searches. Kilo sends block and query text unchanged, through its Gemini provider and OpenAI-Compatible alike, so that prefix is never added.
Avoid three Gemini entries kRouter's catalog still lists. text-embedding-004 and the Gemini Embedding 2 preview are past the shutdown dates Google gives for them (January 14 and August 10, 2026), and text-embedding-005 is a Vertex AI model that Google's Gemini API documentation does not list. kRouter passes the model id through as typed, so gemini/gemini-embedding-2 works even though its list shows only the preview.
Pointing Kilo at kRouter
Install kRouter and start it:
npm install -g @sifxprime/krouter
krouter -tConnect a Gemini API key in the dashboard (Media Providers → Embedding lists the providers that serve embeddings), then check the route before Kilo touches it:
curl http://localhost:20128/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model": "gemini/gemini-embedding-2", "input": ["func add(a, b int) int { return a + b }"], "dimensions": 768}'The reply has the OpenAI shape, "object": "list" with one data entry per input, and each embedding array should hold 768 numbers: kRouter turns the OpenAI dimensions field into Gemini's own output-size setting. GET /v1/models/embedding lists the embedding models of your connected providers.
In Kilo, open Settings → Indexing in VS Code, or run /indexing in the CLI. Choose OpenAI-Compatible, set the base URL to http://localhost:20128/v1, enter the model and a Vector dimension, and save. Every Kilo client reads the same kilo.jsonc (~/.config/kilo/kilo.jsonc, or kilo.jsonc or .kilo/kilo.jsonc in the project); this is the equivalent block:
{
"indexing": {
"enabled": true,
"provider": "openai-compatible",
"model": "gemini/gemini-embedding-2",
"dimension": 768,
"vectorStore": "lancedb",
"openai-compatible": { "baseUrl": "http://localhost:20128/v1" },
"lancedb": {}
}
}There is no API key in it. Callers on the same machine need no kRouter key unless Require API key is on in the Endpoint page, so the Gemini key stays in kRouter instead of in a file that might get committed. If kRouter runs on another machine, put that host in baseUrl and add "apiKey" with a key from the Endpoint page; remote callers always need one. On an Intel Mac, choose Qdrant, because LanceDB does not support it. The Kilo Code card on kRouter's CLI Tools page sets up chat, not indexing.
With OpenAI-Compatible, Kilo sends the dimension with every request, so the model must accept a dimensions field. Through kRouter the Gemini models do, and so does text-embedding-3; OpenAI documents the field only for that generation and later. The trouble is a model that sizes its output with a different field. Voyage's is output_dimension, and kRouter's route forwards only the OpenAI fields, so it also drops the input_type that Kilo's built-in Voyage provider sends. For Voyage, use Kilo's own provider, and set the dimension to 1024 for voyage-code-4: that is its default size, Kilo's Voyage provider does not ask for another, and Kilo's presets do not list the model yet.
Dimensions, and why changing the model means re-indexing
Every embedding model builds its own vector space. A query embedded with one model and compared against blocks embedded with another model of the same size returns results that look plausible and mean nothing, with no error anywhere. Google states that gemini-embedding-001 and gemini-embedding-2 are incompatible and that upgrading means re-embedding everything.
Kilo guards against this. It stores the provider, model id and dimension beside the index, and when any of them changes it drops the index and runs a full scan. Switching models costs a whole first index again, rate limits included.
kRouter's embeddings route stays inside one vector space too. A combo name sent to /v1/embeddings is refused with "Invalid model format", so there is no fallback to a different model; the only fallback is between accounts of the same provider and model, which return vectors in the same space. For the same reason, give Kilo the full provider/model id rather than a kRouter model alias: if you later point the alias elsewhere, Kilo cannot tell that the vectors changed.
Getting through the first index without 429s
When a batch comes back 429, Kilo waits and tries again, starting at five seconds and doubling, but it gives up quickly: three attempts inside the embedder, then up to scannerMaxBatchRetries attempts per batch (3 by default). It does not read the Retry-After header. If more than a tenth of the files fail this way, the first index ends in Error, with a message that starts "Failed during initial scan: Indexing partially failed: Only N of M files were indexed." Three things keep the first index alive.
Smaller batches. Kilo's own advice for rate-limit errors is to lower embeddingBatchSize. That helps against a tokens-per-minute limit. Against a cap on requests, smaller batches make it worse, because the same blocks take more requests.
A second account behind kRouter. Connect two keys for the same provider and kRouter handles the 429 inside the same request: it puts the limited account on a cooldown and sends the batch to the other one, so Kilo never sees the error. Only when every account is cooling down does Kilo get the provider's error; kRouter adds a Retry-After header, which Kilo's backoff ignores. With Gemini, limits apply per project, not per key, so two keys from one project gain nothing; a free key from one project and a paid key from another do. The account rotation post covers how kRouter picks between them.
Not building the index on OpenRouter's free models. Free embedding models such as nvidia/llama-nemotron-embed-vl-1b-v2:free fall under the limits OpenRouter sets for every :free model: 20 requests a minute and 50 a day (1,000 once you have bought 10 credits). At Kilo's default batch of 60, 50 requests is at most 3,000 blocks a day, searches included: enough to keep an index current, not to build one. See the OpenRouter limits post.
What leaves your machine
Kilo sends blocks rather than files, but on the first index every block of every file that is not ignored reaches the provider as written, and so does every search query.
- kRouter does not redact embeddings. PII redaction covers the chat endpoints only, even when switched on, so a key hard-coded in a config file is sent as written. Use
.kilocodeignorefor anything that must not leave the machine. - Free tiers have a price. Google's terms say content sent to unpaid Gemini API quota is used to improve its products, may be read by human reviewers, and should not include sensitive or confidential information. For proprietary code, use a paid key or a local model.
- kRouter keeps no copy. It logs the path and model of each request and a token count when the provider reports an exact one (none for Gemini). It does not store the text.
If the code must stay on the machine, skip kRouter: use Kilo's Ollama provider with LanceDB.
If you are still on Roo Code
Roo Code's repository is archived; its last release, 3.54.0, shipped on May 15, 2026. Its indexer has the same OpenAI Compatible option, with three differences:
- API Key cannot be empty. Any string works against a local kRouter, unless Require API key is on.
- Roo does not send the dimension, so Embedding Dimension must be the model's native size: 3072 for the Gemini models, 1536 for
text-embedding-3-small. - Roo uses Qdrant only, and rebuilds the collection only when the dimension changes. Switch between two models of the same size and unchanged files keep the old model's vectors, so press Clear Index Data after any model change.
Choosing a setup
| Your situation | Use |
|---|---|
| One provider, one key | Kilo's built-in provider, or OpenAI-Compatible pointed straight at the provider. kRouter adds a hop and little else |
| Code that must not leave the machine | Ollama and LanceDB, directly in Kilo |
| First index keeps failing on 429, and you have a second account with that provider | kRouter, so the batch moves to the other account |
Provider keys kept out of kilo.jsonc, or shared by several machines | kRouter, locally or on one host, with a key from its Endpoint page per remote machine |
Common questions
Can kRouter fall back to a different embedding model when a provider fails?
No. Vectors from different models cannot be compared, so a silent switch would corrupt the index without an error. kRouter refuses combo names on /v1/embeddings and falls back only between accounts of the same provider and model.
Do I have to re-index after changing the embedding model?
Yes. Kilo records the provider, model id and dimension, and runs a full scan when any of them changes. Roo Code rebuilds only when the dimension changes, so press Clear Index Data yourself after switching models there.
What dimension should I enter for Gemini through kRouter?
768, 1536 or 3072, Google's recommended sizes for both Gemini embedding models. kRouter passes the value on as Gemini's output size, so vectors come back at exactly that size, and smaller ones take less space.
Is my code redacted before it is embedded?
No. kRouter's PII redaction does not cover /v1/embeddings, even when enabled. Keep secrets out of the index with .kilocodeignore, and keep sensitive code off free tiers whose terms let the provider use what you send.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- How kRouter worksCopilot BYOK custom endpoint: any model in VS Code chatVS Code's Custom Endpoint lets Copilot Chat use any Chat Completions, Responses or Messages API. Using it with a local router, and what still needs GitHub.
- How kRouter worksOpenCode Go in Claude Code and Codex: three protocolsOpenCode Go serves its models on three APIs. Which model needs which, why the wrong one fails with ModelProtocolUnsupported, and how to fix it.
- How kRouter worksXcode with any model: chat, Claude Agent and CodexXcode 26 and 27 can run chat, Claude Agent and Codex on other models, but each reads its own config. How to point all three at one local router.
Relevant docs
- Core conceptsProviders, combos, account routing, token savers, MITM mode, quota tracking, the response cache and remote access: the eight ideas behind kRouter.
- The Zenith Routing EngineHow kRouter picks which of your accounts serves a request: Zenith scoring, conversation stickiness, cooldowns, and the round-robin, P2C and random options.
- Token savers: RTK, Caveman, Ponytail, Headroom, PXPIPEFive ways kRouter can shrink a request before it reaches the provider. Only RTK is on by default. What each saver does, what it needs, and its limits.