Note

Critique's Harness Comparison

A frozen 20-task FeatBench study asking a simple question: when the model and verifier stay fixed, how much does the coding harness change?

It keeps astounding me how big of a difference the harness makes. It’s cool of them to publish these results even though their own harness doesn’t look super great in them.

Critique’s article.