AI Harness Tax
Same model. Same task. Same success rate.
Half the bill.
A new study from UC Berkeley and Arena put a name to something many of us building with AI agents have felt but rarely measured: the harness tax.
The harness is the agent wrapper around the model: Claude Code, Codex CLI, Pi, and others. It turns out the harness can change your inference costs as much as the model you pick.
A few findings stood out:
→ Claude Fable 5 solved ~97% of tasks in Claude Code, Codex CLI, and Pi. But Claude Code cost $1.33 per rollout, while Pi cost $0.67.
→ On SWE-bench Lite, Claude Code cost about 2x more than Pi and 1.6x more than Codex. Success rates were within 2 percentage points of each other.
→ Much of the gap appears before the agent does anything. Claude Code starts with 27,000+ tokens of context, compared with ~2,000 for Pi.
→ Models don't always perform best inside their own vendor's harness. Simple setups are often surprisingly competitive.
The most interesting part for me: even when two setups post identical benchmark numbers, they often fail on completely different tasks. A leaderboard score won't tell you which one fits your codebase.
The researchers' advice is practical:
1- Test a few model and harness combinations on your actual engineering workload
2- Measure cost per solved task, not raw cost per run
3- Pick the cheapest option that clears your reliability bar
4- Retest whenever the model or harness changes
If you're running agents at scale, whether for your own product or for clients, this is the difference between a margin and a leak. The best agent setup isn't the most elaborate one. It's the one that solves your problems at a cost you can sustain.
15
3 comments
Waleed Ijaz
2
AI Harness Tax
Brendan's AI Community
skool.com/brendan
A free community for AI Voice Agents, Claude Code & n8n.
Join to learn, share ideas, and build real systems for the future.
Leaderboard (30-day)
Powered by