Effect sizes, not leaderboards
Report the size of the effect and the number of trials, not a rank.
A leaderboard hides the size and the variance. A 0.2-point lead and a 30-point lead look the same on a board. We report the effect size so a reader can judge whether to care, and the trial count so they can judge whether to believe.
The test: Does the claim carry an effect size and an N, or just a ranking?
Why
Leaderboards optimize for being first, not for being true or useful. They reward overfitting to the benchmark and hide whether a difference would survive a re-run. An effect size with an N is replicable; a rank is not.
How we apply it
- O4 reports
0.893 vs 0.558, d = 1.01, N = 10, not "best on SysML". - The voice standard flags rank-and-superlative framing as a warning, and points the writer back here.
- Results ship with open data so the effect can be recomputed, not trusted.