At the ILOAI 2026 machine-translation olympiad, I fine-tuned a 7B model on chain-of-thought examples and got 0.075, below the untouched 14B model's 0.123. The eventual breakthrough wasn't another training run. It was a parser bug: a heuristic meant to strip leaked reasoning was also deleting correct answers that happened to start with the same words. One bug fix mattered more than every fine-tuning attempt combined. I finished second out of 48, a few hundredths behind first, on a T4, solo. ra...