Activity
Mon
Wed
Fri
Sun
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
What is this?
Less
More
2 contributions to Brendan's AI Community
AI Harness Tax
Same model. Same task. Same success rate. Half the bill. A new study from UC Berkeley and Arena put a name to something many of us building with AI agents have felt but rarely measured: the harness tax. The harness is the agent wrapper around the model: Claude Code, Codex CLI, Pi, and others. It turns out the harness can change your inference costs as much as the model you pick. A few findings stood out: → Claude Fable 5 solved ~97% of tasks in Claude Code, Codex CLI, and Pi. But Claude Code cost $1.33 per rollout, while Pi cost $0.67. → On SWE-bench Lite, Claude Code cost about 2x more than Pi and 1.6x more than Codex. Success rates were within 2 percentage points of each other. → Much of the gap appears before the agent does anything. Claude Code starts with 27,000+ tokens of context, compared with ~2,000 for Pi. → Models don't always perform best inside their own vendor's harness. Simple setups are often surprisingly competitive. The most interesting part for me: even when two setups post identical benchmark numbers, they often fail on completely different tasks. A leaderboard score won't tell you which one fits your codebase. The researchers' advice is practical: 1- Test a few model and harness combinations on your actual engineering workload 2- Measure cost per solved task, not raw cost per run 3- Pick the cheapest option that clears your reliability bar 4- Retest whenever the model or harness changes If you're running agents at scale, whether for your own product or for clients, this is the difference between a margin and a leak. The best agent setup isn't the most elaborate one. It's the one that solves your problems at a cost you can sustain.
AI Harness Tax
💰 $5000 Voice AI Hackathon! (Feb 25th)
We're launching a $5000 Voice AI Hackathon, in partnership with the amazing team at Retell AI! This is a 7-day sprint to build voice agents that solve real business problems. What’s the point? → Build and demo a working AI voice agent using Retell's cutting-edge tech. → Compete for a $5,000 prize pool → Get hands-on experience building the next generation of voice AI. → Showcase your build to a community of agency owners, founders, and automators. This is for AI agency owners, automation builders, SaaS founders, and anyone excited about the future of voice AI. 🥇 1st Place: $1,000 cash + $1,000 in Retell credits 🥈 2nd Place: $500 cash + $750 in Retell credits 🥉 3rd Place: $250 cash + $500 in Retell credits 🏅 2 Honorable Mentions: $500 Retell credits each Want extra build support? Launchpad members get access to our dev team, reviews, and troubleshooting during the hackathon. The official kickoff is on February 25th @ 2PM EST. Launchpad members receive a 48-hour early build window starting February 23rd, with access to setup resources and dev support. To participate: Comment 'HACKATHON' I will send you the registration link via a direct message here.
💰 $5000 Voice AI Hackathon! (Feb 25th)
0 likes • Feb 14
HACKATHON
1-2 of 2
Waleed Ijaz
2
6 points to level up
@waleed-ijaz-6248
Building AI Agents and Automations for Businesses to increase revenue and reduce time

Active 3h ago
Joined Jun 30, 2025
Powered by