Skip to main content
kRouter
All posts
Save money

Caveman vs Ponytail: two ways to cut output tokens

Caveman cuts the words, Ponytail cuts the code. What each prompt does, what their benchmarks measured, and why you should run only one of them.

Kodelyth · The team behind kRouter
· Updated
9 min read

kRouter's Token Saver page has two output toggles, Compress LLM output (Caveman) and Lazy senior dev (Ponytail). Switch both on, expecting two savers to save twice, and every routed chat request logs a warning ending like this (start kRouter with krouter -l to see server logs, or open the dashboard's Console Log page):

[PERSONA] Caveman + Ponytail both enabled - skipping Caveman; pick one on the dashboard

That refusal is deliberate. The two prompts give the model contradictory orders, so kRouter keeps one. And the one you keep will save less than its headline number on agent work, for a reason that has nothing to do with either prompt.

Output is the small part of an agent bill

Input is everything the client sends: system prompt, tool definitions, the conversation so far, every file and tool result. Output is what the model writes back. In a coding agent input dominates, because every turn resends the accumulated context. One user's count across 662 Claude Code sessions, listed on caveman's own Honest Numbers page, put output at 0.3% of all tokens.

RTK, the only saver kRouter turns on by default, compresses tool output inside the request, so it works on input. Caveman and Ponytail are system prompts that change how the model answers, so they work on output only. The outside measurements line up with that:

  • JetBrains ran caveman in Claude Code on SkillsBench coding tasks, with the skill forced on for every reply. Over 82 paired tasks, output tokens fell 8.5% (592k to 542k) and quality was statistically unchanged (p = 0.82). The write-up calls a high-single-digit saving the realistic ceiling for agentic work and says the advertised 65% belongs to chat-style Q&A.
  • LemonCrow's maintainer ran 20 engineering prompts, 300 runs per arm, on claude-opus-4-8. Caveman cut billed output 44.1% and moved total cost by +0.06%. Shorter answers, same bill; the prompt-cache section below explains why.
  • Ponytail's own benchmark reports 45% fewer output tokens but only 26% lower cost, so the rest of the bill fell by much less than the output did.

The number to watch is your provider's bill for the same task, not the output count.

Caveman: fewer words, same code

Caveman makes the model talk tersely. kRouter's version tells it to drop articles, filler, pleasantries and hedging, allows fragments, and asks for the pattern [thing] [action] [reason]. [next step]. It also says what not to touch: code blocks, file paths, commands, errors and URLs stay exact, and security warnings, irreversible confirmations and ordered multi-step instructions are written normally before the terse style resumes.

  • lite keeps grammar and full sentences, and drops filler, hedging and pleasantries
  • full, the default level, drops articles and allows fragments
  • ultra abbreviates and uses arrows for cause and effect
  • With the dashboard in Simplified or Traditional Chinese, three classical-Chinese levels appear too: wenyan-lite, wenyan and wenyan-ultra

The dashboard says "~65% fewer output tokens (up to 87%)". Treat that as a claim kRouter has not measured. Caveman's own Honest Numbers page says earlier releases applied a fixed 65% ratio without a committed, reviewed result. Its current test compares against a plain Answer concisely. instruction, which modern models already follow: on ten dev questions, /caveman cut 3% more output than that at the median and /ultracave 35% more.

Caveman pays off where answers are mostly prose: explanations, reviews, Q&A. Where answers are mostly code it does little, because code is exactly what it is told to leave alone.

Ponytail: less code, via a decision ladder

Ponytail goes after the other part of an answer. kRouter's version casts the model as a "lazy senior developer" who climbs a ladder before writing anything and stops at the first rung that holds:

  1. Does this need to exist at all?
  2. Does the standard library do it?
  3. Does a native platform feature cover it (CSS over JS, a database constraint over app code)?
  4. Does an installed dependency solve it? Never add a new one for a few lines.
  5. Can it be one line?
  6. Only then, the minimum code that works.

It bans unrequested abstractions and scaffolding "for later", and wants the code first with at most three short lines on what was skipped. It must never simplify away validation at trust boundaries, error handling that prevents data loss, security, accessibility, or anything you asked for. Non-trivial logic must leave one runnable check behind, an assert or one small test file, so the persona can add lines as well as cut them.

Levels are lite (build what was asked and name the lazier option), full (the default level, ladder enforced) and ultra (ship the one-liner and challenge the rest of the requirement).

Ponytail's benchmark ran Ponytail 5 in headless Claude Code (Opus 5.5, 39 tasks, five runs each) and reports 53% fewer lines of code, 45% fewer output tokens and 26% lower cost than the same agent without it. On the 18 tasks with hidden correctness checks the arms were level: 87 of 90 runs passed against 86 of 90. Bash was disabled, so agents could write tests but not run them; it measures what the agent writes, not how it debugs. It is careful, and it is the author's.

kRouter ships adaptations, not the upstream files

Both prompts in kRouter are adapted from the upstream skills, and neither matches the current release. The full-level Caveman text is about 960 characters, where caveman's README puts its current /caveman rules at about 1,160 tokens, a much longer text. kRouter's lite, full and ultra levels follow caveman's older naming; caveman 3.2 ships /caveman, /ultracave and /megacave. kRouter's Ponytail ladder has six rungs; Ponytail 5's has seven, adding a check for code already in your repository. kRouter has published no benchmark of its own versions, so the upstream numbers describe the upstream prompts. The only figure that holds for your setup is one you measure.

Why stacking them garbles output

Both prompts close by declaring themselves "ACTIVE EVERY RESPONSE", then disagree on what a response looks like. Caveman wants [thing] [action] [reason]. [next step]. in fragments. Ponytail wants [code] → skipped: [X], add when [Y]. with the code first. A model told to obey both on every turn has to pick each time, and kRouter's changelog notes that Gemini in particular got confused when it received both.

So kRouter enforces one:

  • On the Token Saver card of the Endpoint page, switching one on switches the other off.
  • The standalone Token Saver page lets you set both, but the router then runs Ponytail, skips Caveman and logs the warning above. It prefers Ponytail as the more code-aware of the two.

The same goes for skills installed in your agent. Caveman loaded by the agent plus Caveman from kRouter is two versions of the same instruction; caveman in the agent plus Ponytail from kRouter is the conflict again. One persona, in one place.

Choosing

Your traffic is mostlyPickWhy
Questions, explanations, code review, chatCavemanMost of the answer is prose, which is what it removes
Feature work where the model over-buildsPonytailIt cuts the code, not just the words around it
Long agent sessions, mostly tool callsNeither matters muchInput dominates; RTK is the lever, and it is already on
Requests counted per message or per dayNeitherA shorter answer is still one request
Docs, release notes, text for other peopleNeither, or bypass per requestTerse output is the wrong product

Turning one on for every tool

kRouter injects the persona on the server, just before the request goes to the provider, so one switch covers every client routed through it: Claude Code pointed at another provider, Codex, Cline, OpenCode, Kiro, or a script with no skill system at all.

npm install -g @sifxprime/krouter
krouter -t

Open http://localhost:20128/dashboard, go to Token Saver, switch on Compress LLM output (Caveman) or Lazy senior dev (Ponytail), and pick a level. Both start off. The setting is global, not per combo.

kRouter appends the persona to the system prompt your client already sends (adding one if there is none), or to instructions on a Responses API request. Kiro-format requests have no system slot, so there it goes at the front of the current user message.

Bypassing it for one request

When one call needs the model's normal voice, send X-9Router-Token-Saver: off:

curl http://localhost:20128/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "X-9Router-Token-Saver: off" \
  -d '{"model": "ocg/kimi-k2.6", "messages": [{"role": "user", "content": "Draft the release notes for 2.1"}]}'

Use any provider/model you have connected. The header switches off every token saver for that request (RTK, Caveman, Ponytail, Headroom and PXPIPE) and leaves the dashboard settings alone. By default local callers need no API key; from another machine, add -H "Authorization: Bearer <your-krouter-key>" with a key from the Endpoint page.

The prompt-cache caveat

A persona is input you pay for on every request: about 960 characters for Caveman at the full level and about 1,570 for Ponytail, a few hundred tokens per turn. Where the provider caches prompts, that gets cheaper after the first turn; on Claude-format requests whose system prompt carries a cache breakpoint, kRouter inserts the persona before the last one so it lands inside the cached prefix. Where nothing is cached you pay full price every turn, and on short answers the persona can cost more than it saves. Caveman's Honest Numbers page lists terse coding Q&A as a case where a user measured a net loss, and LemonCrow's maintainer traced their flat bill to the upstream skill adding about 1,500 tokens of cached context per run.

Caching has two more consequences:

  • Claude Code talking to the Claude Code provider gets no persona. Anthropic's cache is keyed on the exact bytes of the request, so in the default Auto mode kRouter leaves that request untouched and skips every saver. The control is under Settings, Cache Control: Always preserve also skips the savers for Anthropic-compatible providers, and Never preserve (legacy) applies them everywhere.
  • Switching a persona on mid-session changes the prompt prefix, so the next turn on a caching provider misses the cache. Decide before you start.

When kRouter is not the answer

If you use one agent and want the whole toolkit, install the upstream project; kRouter ships only the persona prompt. Caveman also has commit and review commands, a memory-file compressor and a local proxy that shrinks what the agent reads. Ponytail has /ponytail-review and /ponytail-audit. Caveman's proxy overlaps with RTK: both compress tool output before the model sees it, so sending the same traffic through both does that job twice. Keep one input compressor per path, for the same reason as one persona.

kRouter's version earns its place when you use several clients, or clients that cannot load skills, and want one switch plus a per-request off button.

Common questions

Can I run Caveman and Ponytail at the same time?

No. Both prompts claim every response and prescribe different output shapes, so the model gets contradictory instructions. On kRouter's Endpoint page, turning one on turns the other off; if both are set anyway, kRouter runs Ponytail, skips Caveman and logs a warning.

How much will Caveman save in a coding agent?

Less than the headline. JetBrains measured 8.5% fewer output tokens over 82 paired coding tasks in Claude Code, because agent output is mostly code and tool calls, which Caveman leaves alone. Compare your provider's billed totals for the same task with it on and off.

Does Ponytail make code less safe?

Its prompt forbids simplifying away validation at trust boundaries, data-loss error handling, security and accessibility. In the author's benchmark of Ponytail 5, hidden correctness checks passed at about the same rate with and without it (87 of 90 runs against 86 of 90), and every arm passed all 30 security-task runs. That tests the upstream prompt, not kRouter's shorter adaptation.

Does the persona apply when Claude Code uses my Claude subscription?

Not by default. When Claude Code talks to the Claude Code provider, kRouter skips all token savers so Anthropic's prompt cache keeps hitting. Routed to any other provider, Claude Code gets the persona like every other client.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter