📦 EVERY ENTRANT GETS A FEEDBACK FILE 📦
🔍 HOW WE READ THESE
Every repo was cloned and pinned to the last commit that was public at the deadline, so nobody was judged on work that landed after the clock. Six repos had later commits. We read the earlier ones.
Then we read file by file. Identity, rules, examples, the reference layer, the code. Every self-test in the field was executed on our machine, not taken on trust.
And for five folders we did the thing the brief describes: dropped the folder in, wrote a draft that appears nowhere in the repo, worked it through as the specialist, then ran the entrant's own checker on what came back. All five passed their own gate.
The landing pages were the doorway. The judging happened inside the folders.
📚 WHAT THE FIELD TAUGHT
Three lines split forty-two builds:
✅ Enforcement moved into code. Comp #8's lesson landed hard. Nine entries ship a checker you can run without an API key. Across forty-two rules files, the phrase "use good judgment" appears zero times.
✅ The disguised ask is the real test. Almost everyone refuses "just rewrite it." The builds that went furthest anticipated the request wearing a disguise: ask me questions and assemble it, give me two options, tell me what it should say instead.
✅ The examples file is where methodology broke. Six entries shipped an examples.md larger than their rules.md, holding voice, philosophy and calibration that belonged in identity or reference. Last cycle it was the empty memory. This cycle it was the overloaded examples file.
🥇 THE WINNER
. The FICC editor, a proposal editor for one municipal culture fund in Campinas, Brazil. Here is why. Three real proponents ran it on real proposals, with consent, on 22 July, inside the fund's live submission window with money on the line. The method was written down before the rounds ran, so the receipts could not be shaped afterward. Inputs are preserved byte for byte, outputs pasted verbatim, and the errors are still in the transcripts because the method said to leave them there.
Then the re-review. The author's own 2023 proposal to this same fund ships in two versions, the draft and the one that won. Round 2 walks its own eighteen findings one at a time, resolved, partial, not addressed, IDs intact, against that real before-and-after pair. And the whole thing runs offline: verify.py checks call anchors, draft excerpts, ID discipline, twelve receipt hashes and five no-rewrite scans. It came back green.
The domain is the narrowest anyone picked. One call, one cycle, one city.
Marcelo takes the Lyceum seat. 🎟️
🥈 THE SHORTLIST
🔹 , Claimline. A claims editor for supplement and wellness copy. We handed his hardest case to a fresh instance cold, blocked from the answer key, and asked it three escalating times to rewrite a health claim. It refused all three and still delivered the full review, and his verifier confirmed every quote was grounded in the draft. He got there by handing the folder to nine strangers who found nineteen defects in it, including a citation he had propagated through seven files to a section number that does not exist. He fixed fourteen and published the rest in OPEN-DEFECTS.md. 🔹 , HOD-Review. Reviews AQA A-level Biology decks the way a Head of Department reads them before Monday. It ingests a real .pptx through a deterministic extractor and quote-checks every finding against the extracted slides, so a fabricated quote fails mechanically no matter how convincing it reads. Nine negative fixtures each get blocked on their exact named check. The same deck read twice in separate chats returned the same high-severity findings on the same slides. He also states plainly that no real teacher has used it, and refuses to dress the constructed run as real. 🔹 , desk-reject. A desk editor for colorectal surgery congress abstracts that computes what your design carries and names the rung your verb climbed without permission. We fed it a fresh single-arm abstract claiming "effective and safe." It read the phrase "No control group was used" through the negation, set the ceiling at describe, caught the causal verb, and blocked. Then read section 4 of his examples file, where he documents his own editor slipping a phrase to an author, explains why his gate is structurally blind to it, and logs it instead of patching a detector against one observed miss. 🔹 , Chalk. A second pair of eyes on Work Health and Safety assessment drafts. The finding schema has no field a fix could live in, so a rewrite has nowhere to ship. He calls it the closest a plain-text system gets to a compile error. We wrote a fresh assessment task, ran it through, and his validator passed the result on all three lenses. TEARDOWN.md records two strangers finding a hole in his own checker, an over claimed line, and an example that did not match its source. All three fixed. 🔹 , Second Pass. An editor for Meta ad concepts at laser tattoo removal clinics. The industry rules live in one swappable file, and he proved the swap by running a real ad from a working real estate agent, someone else's work in a domain he did not build for. It came back clean on Fair Housing, which is a real result rather than a rubber stamp. Four receipts ship the exact input, the verbatim pre-pass and the critique, all re-runnable. 🏅 HONORABLE MENTIONS
🔹 , Taper Editor. Two of her six specimens are plans from athletes who are not her, and one keeps an argument she lost, unresolved, on the record. 🔹 , Vera. Ships a plan with 23 seeded flaws, the answer key, and a cold run that caught all 23. Two minutes to falsify. 🔹 , Author Voice Checker. Shipped the same editor twice, once with a script holding the mechanical rules and once with the model holding them, and published what changes. 🔹 , Reid. The only build that names the request most editors fall for: ask me questions and assemble it. Refused, and a script scans his own examples for violations. 🔹 , Federal Proposal Editor. Refuses to review a proposal without the solicitation it will be scored against. 🔹 , KB Answer Auditor. One inherited SQL Server system, one maintainer, a swappable reader profile, and a paste-ready test case that takes thirty seconds. 🔹 , GCARS Summary Editor. One club, one weekly script, one repeater, and an audience described down to the fact that they cannot rewind. If your name is not on the podium, it does not mean you did not build something worth having. Read your file!
LFG 🚀