Upgraded a 284B local model, measured it 30% slower than the build it replaced, wrote that down as an acceptable trade-off. It wasn't — native multi-token prediction shipped switched off in the new checkpoint. Turning it on: 23.1 → 34.5 tok/s, +49%, and it reverses the regression entirely.
The curriculum lesson here is "never compare across versions without an audit" — two of our own acceptance tests were quietly lying at the same time. Full method map: https://hiddenstatedrift.com/method