427 code-review loops with different models analyzed
One aspect of my current graph-engineering project morphisms is that every run is captured.
I can define which model is used to write the code, and which one will review it.
I often change models to see the impact it has over many runs and different tasks.
The results are interesting and attached in images.
What I learned:
  1. Fable is NOT a good reviewer: it's not thorough
  2. Gpt-5.6-sol and Astra are thorough and good reviewers - and also DON'T (!) favor their own work over other models (which is surprising, hats off to OpenAI)
  3. Opus 5 sucks
It's too early to say anything meaningful about deepseek v4.1 flash, though it's not looking great. Time will tell. It's cheap, fast, but I'm not yet confident in doing deep complex code-reviews in 100-200k token tasks.
What I am now doing:
  1. Use fable to write plans and write code if you can
  2. Review the plan and code with gpt-models for thoroughness
1
0 comments
Blake Sims
1
427 code-review loops with different models analyzed
AI AUTOMATION INSIDERS
skool.com/ai-automation-insiders
Learn Claude Code, AI Agents, and N8N. Install the EXACT AI systems I use inside my $1mil/mo companies. Sell my systems as an AI Agency offer.
Leaderboard (30-day)
Powered by