Skip to main content
kRouter
All posts
How kRouter works

Kilo Code codebase indexing: picking an embedding provider

What Kilo's indexer sends, which embedding model to use for code, why changing it means a full re-index, and how to route it all through one endpoint.

Kodelyth · The team behind kRouter
· Updated
9 min read

You turn on codebase indexing in Kilo Code and pick OpenAI-Compatible as the embedding provider, so you can use a model Kilo has no preset for. You save, and the status goes to Error:

Cannot determine vector dimension for model "gemini/gemini-embedding-001" with provider "openai-compatible". Please set the model dimension explicitly.

That takes ten seconds to fix. Kilo knows the vector sizes of the preset models of its built-in providers, and of nothing behind an OpenAI-compatible endpoint, so you type the number in.

The harder questions sit around it: which model to index code with, what happens to the index when you change it, and how to get a large repository through its first index without being cut off halfway. This post answers them for Kilo Code, including when routing through kRouter helps and when it adds nothing.

What the indexer sends, and how often

Kilo parses files locally with Tree-sitter and cuts them into blocks (functions, classes, methods) of up to about 1,000 characters. It skips binaries, files over 1 MB, node_modules, .git, vendor, and anything matched by .gitignore or .kilocodeignore. Nothing happens until you turn on Enable globally or Enable for this project.

Each block goes to the embedding provider in batches of 60 (embeddingBatchSize, 10 to 200), and the vectors are stored on your machine, in LanceDB by default or in a Qdrant server. The traffic has two phases:

  • The first index sends every block in the repository. On a large codebase that is a long burst, and it is where rate limits bite.
  • After that, a file watcher re-embeds only changed files, and each semantic_search the agent runs embeds one query with the same model.

So the first index decides whether a free tier is enough.

Choosing an embedding model for code

ModelNative sizeCost (vendor pages, October 2026)Notes
gemini-embedding-23072, reducible to 128Free tier; $0.20 per 1M text tokens on paidGenerally available since April 22, 2026; 8,192 input tokens; Google's named replacement for 001
gemini-embedding-0013072, reducible to 128No longer on Google's pricing pageText only, 2,048 input tokens; earliest shutdown May 14, 2028
voyage-code-41024 (also 256, 512, 2048)200M free tokens per account, then $0.12 per 1MVoyage's current code model; Kilo's preset, voyage-code-3, is on Voyage's older-models list at $0.18 per 1M
text-embedding-3-small1536$0.02 per 1M tokensKilo's default for its OpenAI provider
nomic-embed-text (Ollama)768Your hardwareNothing leaves the machine

For a new index on a free tier, gemini-embedding-2 is the sensible default, provided your code may go to a free tier at all; the privacy section below covers that. Google recommends 768, 1536 or 3072 dimensions for both Gemini models; 768 takes a quarter of the space of 3072. For gemini-embedding-2, Google also suggests a task prefix inside the text, such as task: code retrieval | query: ... for searches. Kilo sends block and query text unchanged, through its Gemini provider and OpenAI-Compatible alike, so that prefix is never added.

Avoid three Gemini entries kRouter's catalog still lists. text-embedding-004 and the Gemini Embedding 2 preview are past the shutdown dates Google gives for them (January 14 and August 10, 2026), and text-embedding-005 is a Vertex AI model that Google's Gemini API documentation does not list. kRouter passes the model id through as typed, so gemini/gemini-embedding-2 works even though its list shows only the preview.

Pointing Kilo at kRouter

Install kRouter and start it:

npm install -g @sifxprime/krouter
krouter -t

Connect a Gemini API key in the dashboard (Media Providers → Embedding lists the providers that serve embeddings), then check the route before Kilo touches it:

curl http://localhost:20128/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model": "gemini/gemini-embedding-2", "input": ["func add(a, b int) int { return a + b }"], "dimensions": 768}'

The reply has the OpenAI shape, "object": "list" with one data entry per input, and each embedding array should hold 768 numbers: kRouter turns the OpenAI dimensions field into Gemini's own output-size setting. GET /v1/models/embedding lists the embedding models of your connected providers.

In Kilo, open Settings → Indexing in VS Code, or run /indexing in the CLI. Choose OpenAI-Compatible, set the base URL to http://localhost:20128/v1, enter the model and a Vector dimension, and save. Every Kilo client reads the same kilo.jsonc (~/.config/kilo/kilo.jsonc, or kilo.jsonc or .kilo/kilo.jsonc in the project); this is the equivalent block:

{
  "indexing": {
    "enabled": true,
    "provider": "openai-compatible",
    "model": "gemini/gemini-embedding-2",
    "dimension": 768,
    "vectorStore": "lancedb",
    "openai-compatible": { "baseUrl": "http://localhost:20128/v1" },
    "lancedb": {}
  }
}

There is no API key in it. Callers on the same machine need no kRouter key unless Require API key is on in the Endpoint page, so the Gemini key stays in kRouter instead of in a file that might get committed. If kRouter runs on another machine, put that host in baseUrl and add "apiKey" with a key from the Endpoint page; remote callers always need one. On an Intel Mac, choose Qdrant, because LanceDB does not support it. The Kilo Code card on kRouter's CLI Tools page sets up chat, not indexing.

With OpenAI-Compatible, Kilo sends the dimension with every request, so the model must accept a dimensions field. Through kRouter the Gemini models do, and so does text-embedding-3; OpenAI documents the field only for that generation and later. The trouble is a model that sizes its output with a different field. Voyage's is output_dimension, and kRouter's route forwards only the OpenAI fields, so it also drops the input_type that Kilo's built-in Voyage provider sends. For Voyage, use Kilo's own provider, and set the dimension to 1024 for voyage-code-4: that is its default size, Kilo's Voyage provider does not ask for another, and Kilo's presets do not list the model yet.

Dimensions, and why changing the model means re-indexing

Every embedding model builds its own vector space. A query embedded with one model and compared against blocks embedded with another model of the same size returns results that look plausible and mean nothing, with no error anywhere. Google states that gemini-embedding-001 and gemini-embedding-2 are incompatible and that upgrading means re-embedding everything.

Kilo guards against this. It stores the provider, model id and dimension beside the index, and when any of them changes it drops the index and runs a full scan. Switching models costs a whole first index again, rate limits included.

kRouter's embeddings route stays inside one vector space too. A combo name sent to /v1/embeddings is refused with "Invalid model format", so there is no fallback to a different model; the only fallback is between accounts of the same provider and model, which return vectors in the same space. For the same reason, give Kilo the full provider/model id rather than a kRouter model alias: if you later point the alias elsewhere, Kilo cannot tell that the vectors changed.

Getting through the first index without 429s

When a batch comes back 429, Kilo waits and tries again, starting at five seconds and doubling, but it gives up quickly: three attempts inside the embedder, then up to scannerMaxBatchRetries attempts per batch (3 by default). It does not read the Retry-After header. If more than a tenth of the files fail this way, the first index ends in Error, with a message that starts "Failed during initial scan: Indexing partially failed: Only N of M files were indexed." Three things keep the first index alive.

Smaller batches. Kilo's own advice for rate-limit errors is to lower embeddingBatchSize. That helps against a tokens-per-minute limit. Against a cap on requests, smaller batches make it worse, because the same blocks take more requests.

A second account behind kRouter. Connect two keys for the same provider and kRouter handles the 429 inside the same request: it puts the limited account on a cooldown and sends the batch to the other one, so Kilo never sees the error. Only when every account is cooling down does Kilo get the provider's error; kRouter adds a Retry-After header, which Kilo's backoff ignores. With Gemini, limits apply per project, not per key, so two keys from one project gain nothing; a free key from one project and a paid key from another do. The account rotation post covers how kRouter picks between them.

Not building the index on OpenRouter's free models. Free embedding models such as nvidia/llama-nemotron-embed-vl-1b-v2:free fall under the limits OpenRouter sets for every :free model: 20 requests a minute and 50 a day (1,000 once you have bought 10 credits). At Kilo's default batch of 60, 50 requests is at most 3,000 blocks a day, searches included: enough to keep an index current, not to build one. See the OpenRouter limits post.

What leaves your machine

Kilo sends blocks rather than files, but on the first index every block of every file that is not ignored reaches the provider as written, and so does every search query.

  • kRouter does not redact embeddings. PII redaction covers the chat endpoints only, even when switched on, so a key hard-coded in a config file is sent as written. Use .kilocodeignore for anything that must not leave the machine.
  • Free tiers have a price. Google's terms say content sent to unpaid Gemini API quota is used to improve its products, may be read by human reviewers, and should not include sensitive or confidential information. For proprietary code, use a paid key or a local model.
  • kRouter keeps no copy. It logs the path and model of each request and a token count when the provider reports an exact one (none for Gemini). It does not store the text.

If the code must stay on the machine, skip kRouter: use Kilo's Ollama provider with LanceDB.

If you are still on Roo Code

Roo Code's repository is archived; its last release, 3.54.0, shipped on May 15, 2026. Its indexer has the same OpenAI Compatible option, with three differences:

  • API Key cannot be empty. Any string works against a local kRouter, unless Require API key is on.
  • Roo does not send the dimension, so Embedding Dimension must be the model's native size: 3072 for the Gemini models, 1536 for text-embedding-3-small.
  • Roo uses Qdrant only, and rebuilds the collection only when the dimension changes. Switch between two models of the same size and unchanged files keep the old model's vectors, so press Clear Index Data after any model change.

Choosing a setup

Your situationUse
One provider, one keyKilo's built-in provider, or OpenAI-Compatible pointed straight at the provider. kRouter adds a hop and little else
Code that must not leave the machineOllama and LanceDB, directly in Kilo
First index keeps failing on 429, and you have a second account with that providerkRouter, so the batch moves to the other account
Provider keys kept out of kilo.jsonc, or shared by several machineskRouter, locally or on one host, with a key from its Endpoint page per remote machine

Common questions

Can kRouter fall back to a different embedding model when a provider fails?

No. Vectors from different models cannot be compared, so a silent switch would corrupt the index without an error. kRouter refuses combo names on /v1/embeddings and falls back only between accounts of the same provider and model.

Do I have to re-index after changing the embedding model?

Yes. Kilo records the provider, model id and dimension, and runs a full scan when any of them changes. Roo Code rebuilds only when the dimension changes, so press Clear Index Data yourself after switching models there.

What dimension should I enter for Gemini through kRouter?

768, 1536 or 3072, Google's recommended sizes for both Gemini embedding models. kRouter passes the value on as Gemini's output size, so vectors come back at exactly that size, and smaller ones take less space.

Is my code redacted before it is embedded?

No. kRouter's PII redaction does not cover /v1/embeddings, even when enabled. Keep secrets out of the index with .kilocodeignore, and keep sensitive code off free tiers whose terms let the provider use what you send.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter