Note
Critique's Harness Comparison
(56 words, 1 minute)
A frozen 20-task FeatBench study asking a simple question: when the model and verifier stay fixed, how much does the coding harness change?
It keeps astounding me how big of a difference the harness makes. It’s cool of them to publish these results even though their own harness doesn’t look super great in them.