Same model. Same task. Same success rate. Half the bill. A new study from UC Berkeley and Arena put a name to something many of us building with AI agents have felt but rarely measured: the harness tax. The harness is the agent wrapper around the model: Claude Code, Codex CLI, Pi, and others. It turns out the harness can change your inference costs as much as the model you pick. A few findings stood out: → Claude Fable 5 solved ~97% of tasks in Claude Code, Codex CLI, and Pi. But Claude Code cost $1.33 per rollout, while Pi cost $0.67. → On SWE-bench Lite, Claude Code cost about 2x more than Pi and 1.6x more than Codex. Success rates were within 2 percentage points of each other. → Much of the gap appears before the agent does anything. Claude Code starts with 27,000+ tokens of context, compared with ~2,000 for Pi. → Models don't always perform best inside their own vendor's harness. Simple setups are often surprisingly competitive. The most interesting part for me: even when two setups post identical benchmark numbers, they often fail on completely different tasks. A leaderboard score won't tell you which one fits your codebase. The researchers' advice is practical: 1- Test a few model and harness combinations on your actual engineering workload 2- Measure cost per solved task, not raw cost per run 3- Pick the cheapest option that clears your reliability bar 4- Retest whenever the model or harness changes If you're running agents at scale, whether for your own product or for clients, this is the difference between a margin and a leak. The best agent setup isn't the most elaborate one. It's the one that solves your problems at a cost you can sustain.