AI GlossaryㅋWords you meet while using AI
Caveman
An open-source tool that compresses the data exchanged with AI coding agents to cut token costs
In plain words
Caveman is an open-source tool that makes AI coding agents talk short, like cavemen, to save money. Developers get billed based on how much text goes back and forth every time they use a coding agent, so Caveman trims the agent's explanations the way a long email gets compressed into a telegram. Parts that need to stay exact, like code or commands, are left untouched—only the wordy explanations get cut down.
While the early version only shortened what the agent 'said,' a later proxy also touches what the agent 'reads.' Chunks of text that get sent back and forth unchanged every time—things like tool lists, logs, or previous conversation history—get converted into a single image instead. Because text and images are billed differently, turning dense blocks of text into a picture exploits a gap that actually lowers the cost.
But as the developer himself admits, the savings vary by situation. Just shortening the response doesn't touch the cost of other data exchanged every turn, and converting to images can actually backfire on code that's already short. That's why the repository description includes the line: 'we don't make the agent dumb just to make it cheap.'
How it shows up in the news
The article cites it with a specific figure, saying something like "'Caveman 2' and the new local proxy... cut input tokens by 33.2% according to provider reporting." This number shouldn't be mistaken for an official performance improvement announced by Anthropic. Caveman isn't a feature built by Anthropic—it's a separate open-source tool created and released by developer Julius Brussee, meant to be layered on top of tools like Claude Code.
Try it yourself
- Follow the install instructions in the Caveman repository to set it up as a skill or proxy in your coding agent environment.
- Ask the agent the exact same question or task you would have asked before installing it.
- Compare how the response length and token usage (visible via provider billing or logs) change.
- Since tasks with lots of short code snippets might actually use more tokens, it's more accurate to test across several different kinds of tasks.
See also
Stories using this term
- Caveman Cuts Claude Code Token Usage by 33% Using Caveman-SpeakAI · 2026.08.20
- Claude Code auto-inserts session links into every commit, sparking backlashAI · 2026.08.31
- Claude Code hooks block rule-skipping with codeAI · 2026.09.04
- Claude Code adds /design skill for UI drafts as research previewAI · 2026.08.18
- Claude Code Gets Faster Startup, Clearer Token Tracking, Remote Control FixesAI · 2026.08.30
- Claude Code weekly limits: September 14 permanent 25% hike is a real-world cutAI · 2026.08.30
