Which model is best at reading and understanding construction drawings? @Josh Turner and I just launched an AI benchmark for construction tasks. If you've not come across a benchmark, it's a way to score AI models. There's benchmarks for software engineering, legal, maths, etc. For our benchmark, we went through multiple sets of large construction drawings and created a database of questions and answers (Very time-consuming). Questions like "how many footings are in the slab" or "what are the cable tray specifications. We then scored each model based on what percentage of questions they answered correctly and the cost of the analysis. So far, we've tested 9 recent mods. GPT 5.6 won. Surprisingly, Opus 5 and 4.8 both beat Fable 5. One caveat: we tested the models via the API, so you would likely get different performance (almost certainly better results) if you used the harness (Claude Code, Codex, Cursor, etc.) We'll be running and updating this for every new model that comes out. We will also be adding more domains to our analysis - take-offs, estimating, scheduling, etc. We've got come up with some pretty interesting ways to test them. Check out the results here: https://contractoros.build/benchmarks/