Caveman vs Ponytail: two ways to cut output tokens
Caveman cuts the words, Ponytail cuts the code. What each prompt does, what their benchmarks measured, and why you should run only one of them.
kRouter's Token Saver page has two output toggles, Compress LLM output (Caveman) and Lazy senior dev (Ponytail). Switch both on, expecting two savers to save twice, and every routed chat request logs a warning ending like this (start kRouter with krouter -l to see server logs, or open the dashboard's Console Log page):
[PERSONA] Caveman + Ponytail both enabled - skipping Caveman; pick one on the dashboardThat refusal is deliberate. The two prompts give the model contradictory orders, so kRouter keeps one. And the one you keep will save less than its headline number on agent work, for a reason that has nothing to do with either prompt.
Output is the small part of an agent bill
Input is everything the client sends: system prompt, tool definitions, the conversation so far, every file and tool result. Output is what the model writes back. In a coding agent input dominates, because every turn resends the accumulated context. One user's count across 662 Claude Code sessions, listed on caveman's own Honest Numbers page, put output at 0.3% of all tokens.
RTK, the only saver kRouter turns on by default, compresses tool output inside the request, so it works on input. Caveman and Ponytail are system prompts that change how the model answers, so they work on output only. The outside measurements line up with that:
- JetBrains ran caveman in Claude Code on SkillsBench coding tasks, with the skill forced on for every reply. Over 82 paired tasks, output tokens fell 8.5% (592k to 542k) and quality was statistically unchanged (p = 0.82). The write-up calls a high-single-digit saving the realistic ceiling for agentic work and says the advertised 65% belongs to chat-style Q&A.
- LemonCrow's maintainer ran 20 engineering prompts, 300 runs per arm, on claude-opus-4-8. Caveman cut billed output 44.1% and moved total cost by +0.06%. Shorter answers, same bill; the prompt-cache section below explains why.
- Ponytail's own benchmark reports 45% fewer output tokens but only 26% lower cost, so the rest of the bill fell by much less than the output did.
The number to watch is your provider's bill for the same task, not the output count.
Caveman: fewer words, same code
Caveman makes the model talk tersely. kRouter's version tells it to drop articles, filler, pleasantries and hedging, allows fragments, and asks for the pattern [thing] [action] [reason]. [next step]. It also says what not to touch: code blocks, file paths, commands, errors and URLs stay exact, and security warnings, irreversible confirmations and ordered multi-step instructions are written normally before the terse style resumes.
- lite keeps grammar and full sentences, and drops filler, hedging and pleasantries
- full, the default level, drops articles and allows fragments
- ultra abbreviates and uses arrows for cause and effect
- With the dashboard in Simplified or Traditional Chinese, three classical-Chinese levels appear too: wenyan-lite, wenyan and wenyan-ultra
The dashboard says "~65% fewer output tokens (up to 87%)". Treat that as a claim kRouter has not measured. Caveman's own Honest Numbers page says earlier releases applied a fixed 65% ratio without a committed, reviewed result. Its current test compares against a plain Answer concisely. instruction, which modern models already follow: on ten dev questions, /caveman cut 3% more output than that at the median and /ultracave 35% more.
Caveman pays off where answers are mostly prose: explanations, reviews, Q&A. Where answers are mostly code it does little, because code is exactly what it is told to leave alone.
Ponytail: less code, via a decision ladder
Ponytail goes after the other part of an answer. kRouter's version casts the model as a "lazy senior developer" who climbs a ladder before writing anything and stops at the first rung that holds:
- Does this need to exist at all?
- Does the standard library do it?
- Does a native platform feature cover it (CSS over JS, a database constraint over app code)?
- Does an installed dependency solve it? Never add a new one for a few lines.
- Can it be one line?
- Only then, the minimum code that works.
It bans unrequested abstractions and scaffolding "for later", and wants the code first with at most three short lines on what was skipped. It must never simplify away validation at trust boundaries, error handling that prevents data loss, security, accessibility, or anything you asked for. Non-trivial logic must leave one runnable check behind, an assert or one small test file, so the persona can add lines as well as cut them.
Levels are lite (build what was asked and name the lazier option), full (the default level, ladder enforced) and ultra (ship the one-liner and challenge the rest of the requirement).
Ponytail's benchmark ran Ponytail 5 in headless Claude Code (Opus 5.5, 39 tasks, five runs each) and reports 53% fewer lines of code, 45% fewer output tokens and 26% lower cost than the same agent without it. On the 18 tasks with hidden correctness checks the arms were level: 87 of 90 runs passed against 86 of 90. Bash was disabled, so agents could write tests but not run them; it measures what the agent writes, not how it debugs. It is careful, and it is the author's.
kRouter ships adaptations, not the upstream files
Both prompts in kRouter are adapted from the upstream skills, and neither matches the current release. The full-level Caveman text is about 960 characters, where caveman's README puts its current /caveman rules at about 1,160 tokens, a much longer text. kRouter's lite, full and ultra levels follow caveman's older naming; caveman 3.2 ships /caveman, /ultracave and /megacave. kRouter's Ponytail ladder has six rungs; Ponytail 5's has seven, adding a check for code already in your repository. kRouter has published no benchmark of its own versions, so the upstream numbers describe the upstream prompts. The only figure that holds for your setup is one you measure.
Why stacking them garbles output
Both prompts close by declaring themselves "ACTIVE EVERY RESPONSE", then disagree on what a response looks like. Caveman wants [thing] [action] [reason]. [next step]. in fragments. Ponytail wants [code] → skipped: [X], add when [Y]. with the code first. A model told to obey both on every turn has to pick each time, and kRouter's changelog notes that Gemini in particular got confused when it received both.
So kRouter enforces one:
- On the Token Saver card of the Endpoint page, switching one on switches the other off.
- The standalone Token Saver page lets you set both, but the router then runs Ponytail, skips Caveman and logs the warning above. It prefers Ponytail as the more code-aware of the two.
The same goes for skills installed in your agent. Caveman loaded by the agent plus Caveman from kRouter is two versions of the same instruction; caveman in the agent plus Ponytail from kRouter is the conflict again. One persona, in one place.
Choosing
| Your traffic is mostly | Pick | Why |
|---|---|---|
| Questions, explanations, code review, chat | Caveman | Most of the answer is prose, which is what it removes |
| Feature work where the model over-builds | Ponytail | It cuts the code, not just the words around it |
| Long agent sessions, mostly tool calls | Neither matters much | Input dominates; RTK is the lever, and it is already on |
| Requests counted per message or per day | Neither | A shorter answer is still one request |
| Docs, release notes, text for other people | Neither, or bypass per request | Terse output is the wrong product |
Turning one on for every tool
kRouter injects the persona on the server, just before the request goes to the provider, so one switch covers every client routed through it: Claude Code pointed at another provider, Codex, Cline, OpenCode, Kiro, or a script with no skill system at all.
npm install -g @sifxprime/krouter
krouter -tOpen http://localhost:20128/dashboard, go to Token Saver, switch on Compress LLM output (Caveman) or Lazy senior dev (Ponytail), and pick a level. Both start off. The setting is global, not per combo.
kRouter appends the persona to the system prompt your client already sends (adding one if there is none), or to instructions on a Responses API request. Kiro-format requests have no system slot, so there it goes at the front of the current user message.
Bypassing it for one request
When one call needs the model's normal voice, send X-9Router-Token-Saver: off:
curl http://localhost:20128/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-9Router-Token-Saver: off" \
-d '{"model": "ocg/kimi-k2.6", "messages": [{"role": "user", "content": "Draft the release notes for 2.1"}]}'Use any provider/model you have connected. The header switches off every token saver for that request (RTK, Caveman, Ponytail, Headroom and PXPIPE) and leaves the dashboard settings alone. By default local callers need no API key; from another machine, add -H "Authorization: Bearer <your-krouter-key>" with a key from the Endpoint page.
The prompt-cache caveat
A persona is input you pay for on every request: about 960 characters for Caveman at the full level and about 1,570 for Ponytail, a few hundred tokens per turn. Where the provider caches prompts, that gets cheaper after the first turn; on Claude-format requests whose system prompt carries a cache breakpoint, kRouter inserts the persona before the last one so it lands inside the cached prefix. Where nothing is cached you pay full price every turn, and on short answers the persona can cost more than it saves. Caveman's Honest Numbers page lists terse coding Q&A as a case where a user measured a net loss, and LemonCrow's maintainer traced their flat bill to the upstream skill adding about 1,500 tokens of cached context per run.
Caching has two more consequences:
- Claude Code talking to the Claude Code provider gets no persona. Anthropic's cache is keyed on the exact bytes of the request, so in the default Auto mode kRouter leaves that request untouched and skips every saver. The control is under Settings, Cache Control: Always preserve also skips the savers for Anthropic-compatible providers, and Never preserve (legacy) applies them everywhere.
- Switching a persona on mid-session changes the prompt prefix, so the next turn on a caching provider misses the cache. Decide before you start.
When kRouter is not the answer
If you use one agent and want the whole toolkit, install the upstream project; kRouter ships only the persona prompt. Caveman also has commit and review commands, a memory-file compressor and a local proxy that shrinks what the agent reads. Ponytail has /ponytail-review and /ponytail-audit. Caveman's proxy overlaps with RTK: both compress tool output before the model sees it, so sending the same traffic through both does that job twice. Keep one input compressor per path, for the same reason as one persona.
kRouter's version earns its place when you use several clients, or clients that cannot load skills, and want one switch plus a per-request off button.
Common questions
Can I run Caveman and Ponytail at the same time?
No. Both prompts claim every response and prescribe different output shapes, so the model gets contradictory instructions. On kRouter's Endpoint page, turning one on turns the other off; if both are set anyway, kRouter runs Ponytail, skips Caveman and logs a warning.
How much will Caveman save in a coding agent?
Less than the headline. JetBrains measured 8.5% fewer output tokens over 82 paired coding tasks in Claude Code, because agent output is mostly code and tool calls, which Caveman leaves alone. Compare your provider's billed totals for the same task with it on and off.
Does Ponytail make code less safe?
Its prompt forbids simplifying away validation at trust boundaries, data-loss error handling, security and accessibility. In the author's benchmark of Ponytail 5, hidden correctness checks passed at about the same rate with and without it (87 of 90 runs against 86 of 90), and every arm passed all 30 security-task runs. That tests the upstream prompt, not kRouter's shorter adaptation.
Does the persona apply when Claude Code uses my Claude subscription?
Not by default. When Claude Code talks to the Claude Code provider, kRouter skips all token savers so Anthropic's prompt cache keeps hitting. Routed to any other provider, Claude Code gets the persona like every other client.
Related
Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.
Install kRouterRelated posts
- Save moneyGLM, MiniMax and Alibaba coding plans: stack two of themWhat the GLM, MiniMax and Alibaba coding plans meter, where their 5-hour and weekly windows bite, and how to fall back from one plan to the next.
- Save moneyWhy your Cline bill hit $200 a month, and how to cut itA Cline bill grows with every file it reads and every command it runs, because each step resends them. Where the tokens go and what really cuts them.
- Save moneyGemini CLI's free tier is gone: what still works for freeGoogle stopped serving Gemini CLI to free, AI Pro and Ultra users on June 18, 2026. Who kept access, what is still free, and how to keep using Gemini.
Relevant docs
- Token savers: RTK, Caveman, Ponytail, Headroom, PXPIPEFive ways kRouter can shrink a request before it reaches the provider. Only RTK is on by default. What each saver does, what it needs, and its limits.
- Combos & fallbackPut several models behind one name. kRouter falls back from one to the next, rotates them, or asks a panel of models and merges the answers.
- ProvidersConnect 95+ AI providers to kRouter: OAuth subscriptions, free services, free-tier and pay-per-token API keys, and browser-cookie logins.