Codex vs. Claude...Decisions, Decisions, Decisions
I read the posts in this community almost daily, and one question keeps circling back: Codex or Claude, which one should you build with. I run both, in separate seats. Codex runs my research dispatches, a Claude line drafts and builds. Until this week my honest answer was a preference dressed up as analysis. Then a piece of real planning work handed me a way to make the agents answer it themselves.
I needed a new room in my workspace: an executive desk that separates my strategy, marketing, and revenue thinking, so a business decision gets read through all three lenses before I rule on it. Real need, no obvious shape. No build on the line yet either; this is the step before, the planning. So I wrote the brief once and gave it to both agents independently. Neither saw the other's work.
Claude came back with nine files. Three seats, each holding a one-page charter. One inbox file where a passing thought gets captured as a single tagged line. A two-file deliberation template that gets copied whenever a thought earns the full treatment. A record library.
Codex came back with an eight-stage pipeline: capture, route, one stage per seat, synthesis, operator decision, handoff. A contract at every stage, a config folder, durable IDs joining each thought's artifacts across stages. Somewhere between eighteen and twenty files, over twenty-five folders.
Same brief. Two different species of system.
I still didn't pick. Instead I had each agent analyze both plans, independently, against one question: which form fits how this work actually behaves. Then I read the two verdicts side by side.
They converged. Both analyses ruled the record library the right form, and both kept the pipeline's best discipline as written rules instead of folders. Codex's own report called the other agent's plan the stronger operating model for a solo operator and recommended rejecting its own eight stages. The reconciled design came out at seven files, smaller than either original plan. The reasoning both landed on is the part worth keeping: thought-work accumulates and varies, it does not repeat one production sequence. A pipeline is the right shape when the work repeats. A library is the right shape when the work accrues.
The lazy read is that Codex overbuilds and Claude keeps things lean. I don't trust that read, because the contest wasn't level. My Claude session carries months of memory. It watched me cut a seventeen-role agent roster down to two in June because described behavior that nothing enforces collapses into generic model output. It has my simplification rulings on file. Codex walked in and read the repo cold, no scars attached. Each agent designed to what it knew about me, not only to what it is. I can't prove whether the model or the memory decided this contest. I know which one I'd bet on.
Use this on your next design decision that matters:
  1. Write the brief once. Hand it to two different agents independently. Neither sees the other's work.
  2. Take a complete plan from each. Don't steer either one mid-flight.
  3. Swap: each agent analyzes both plans against one question, which form fits how the work actually behaves.
  4. Read the verdicts side by side. Where they diverge, that is a real decision and it belongs to you. Where they converge, confidence is high. An agent ruling against its own plan is the strongest signal on the board.
  5. Rule yourself, in writing. Nothing gets built until you sign. Mine is still a plan, deliberately. Both agents agree and it stays planning until I rule, because building is its own decision. That is the point of being the operator.
The frame under all of it is Jake's ICM. Both plans arrived as folder architectures because in this methodology the folders are the agent architecture, and the winning form wasn't invented for the contest. It was proven by a deliberation container I already run every week.
One more thing, in the interest of the disclosure I always make: this post was drafted by a line that runs on Claude. That is exactly why the swap step exists. The verdict I trust comes from two systems with two different makers reading the same evidence and landing in the same place, not from one agent's opinion of itself.
If you run more than one agent, I want to know: have you ever made them judge each other's work, and did they converge?
Drafted by the line, edited and signed by me. Eighth time.
5
0 comments
Jordan Shaw
6
Codex vs. Claude...Decisions, Decisions, Decisions
Clief Notes
skool.com/cliefnotes
What we give away free beats most paid courses. Build durable AI systems with a Marine vet and Edinburgh researcher. 40+ lessons, growing.
Leaderboard (30-day)
Powered by