34 entries. Jake, Don, David on isolated piles, then a panel. This one was hard. A few of you built things that would have won a different competition, and you should hear which one.
Every entry got feedback, link at bottom of post!
๐ What we actually judged
Four questions, from the post:
- Does it diagnose? One cause. Not a list. Not a prescription.
- Is the domain specific enough to be useful?
- Does each file do one job?
- Can a stranger figure it out?
That is an outcome. Can a person who is not already in your head use this to see why something in their world broke.
A lot of the field went further. Refusal paths. Verifiers. Blind runs. Preserved failures. Those are real, and they are extra. We did not pick a winner on the extra. We picked a winner on the ask.
this does not mean extra is not better, in many cases I LOVED the solutions created.
If the outcome had been different, the name on this post changes honestly
A stranger holding a bounced invoice can drop the folder in and see why it bounced.
That was the assignment.
๐ Where someone else wins
๐ง If the outcome was "the instrument that can prove itself wrong" โ Sergey Manevitch, Radix https://github.com/sergeymanevitch/Radix Why a machine failed in a plant that already has notifications, rounds, and historian trends. Fourteen cold runs. He left the breaks in his own claims standing. A plant engineer is holding something serious. ๐งช If the outcome was "every claim reproduced, including the one that failed" โ Pemmy Broke, Visual Momentum https://github.com/hoodwanders/visual-momentum-diagnostician Why an AI explainer lost momentum. She published the run that broke her own doctrine, changed the rules, and the re-run abstained. That is the standard the rest of the field should steal from. ๐ซ If the outcome was "could reality contradict you" โ Marcelo Michelsohn, Why this conversation drained me https://github.com/marcelomichelsohn/why-this-conversation-drained-me Nine real accounts. Prediction committed before the run. Two rounds where the machine read better than he did. What he protected the person in pain from, he did not protect himself from. That is the right way round. โ๏ธ If the outcome was "the constraint lives in code, not in a README" โ Jodi Paige-Lee, Instruction-set autopsy https://github.com/jmarielee/instruction-set-autopsy The one defect in follow-up paperwork most likely to make a caregiver fail at home. Eleven gates. A judge planted a fake quote. It died on a named gate. Those four win those competitions. This competition was the folder a stranger can use. Colm is the pile-winner you can hand to someone who does not already know the domain. The interview assumes you have never heard of an NDR. That is the product.
๐ How they connect
Pemmy's letter points back at Colm: a stranger who does not know the domain can still use the folder. His interview is the shape her stranger test is reaching for. She should finish that test. Everything else already holds.
Sergey's receipts are the best in the round. If you want to see what those marker counts become when they actually run, look at Jodi's gates, and at Craig Howard's historian โ a script computes every number, the model only labels. https://github.com/craig-atr/js-ts-regression-historian Jodi's letter points at Marcelo for the other half of enforcement: an independent key that is a real person, not another pass of the same tool.
That is the field talking to itself. Steal from it.
๐ A few more I want named
Adam James, Greenlight โ four dated months of a live behavioral-health intake. "Green doesn't always mean go." Closest match to the five-file spec we posted. The operator still decides. https://github.com/awjames6875/Greenlight ๐ฌ Your feedback is in the vault
Every entrant gets a letter. What landed, what to close, and one or two people in this field who already did the specific thing you are missing.
Read yours. Then go read the person we pointed you at.
๐๏ธ Colm takes the Lyceum seat. First entry. The bicycle already works.