← Principles

Effect sizes, not leaderboards

Report the size of the effect and the number of trials, not a rank.

A leaderboard hides the size and the variance. A 0.2-point lead and a 30-point lead look the same on a board. We report the effect size so a reader can judge whether to care, and the trial count so they can judge whether to believe.

The test: Does the claim carry an effect size and an N, or just a ranking?

Why

Leaderboards optimize for being first, not for being true or useful. They reward overfitting to the benchmark and hide whether a difference would survive a re-run. An effect size with an N is replicable; a rank is not.

How we apply it

  • O4 reports 0.893 vs 0.558, d = 1.01, N = 10, not "best on SysML".
  • The voice standard flags rank-and-superlative framing as a warning, and points the writer back here.
  • Results ship with open data so the effect can be recomputed, not trusted.