Bayesian credibility, and the influencers we refuse to rank

Search for a command to run...

The leaderboard has a rule that most leaderboards would rather not have. Below fifty scored claims, we do not rank. The influencer stays on the page, their data is visible, but the rank and the score read as "not enough data yet." We would rather show a hole than fabricate a verdict.
This is a post about the statistics behind that rule — why small samples lie, what Bayesian shrinkage actually does to fix them, and why even the fix has a floor below which honesty requires silence.
Start with the直觉 failure. An influencer makes five claims we can score. Four of them align with the research. Raw accuracy: eighty percent. Put that on a leaderboard and they look excellent.
But four out of five is not eighty percent in any meaningful sense. It is a single observation from a process with high variance. The true rate at which that influencer aligns with the research could easily be anywhere from forty to ninety-five percent and still produce a four-out-of-five result by chance. The point estimate is eighty. The honest range is enormous. Reporting the point estimate without the uncertainty is not just imprecise — it actively misleads, because eighty percent feels like knowledge and "somewhere between forty and ninety-five" feels like ignorance, and they are the same data.
This is the small-sample trap, and it is the thing that makes most leaderboards unreliable. Anyone near the top with few data points is probably there partly by luck. Anyone near the bottom is probably there partly by bad luck. The ranking reflects noise as much as signal.
The standard fix is to pull extreme small-sample estimates back toward the average, in proportion to how little evidence supports them. This is called shrinkage, and it is the difference between a naive percentage and a credible score.
Our score works like this. Every influencer has a weighted accuracy — the share of their claims that align with the research, with each claim weighted by how strong the underlying evidence is, so that being right about a heavily studied outcome counts for more. That weighted accuracy is the raw signal. Then we mix it with a prior: the average accuracy across everyone on the board. The mix is controlled by how many claims the influencer has.
The formula, in words: take the influencer's weighted accuracy, contribute it once for every claim they have, and contribute the global average a fixed number of times. Then divide by the total. The fixed number — fifty, in our case — is how many claims it takes before the influencer's own record outweighs the prior. With a handful of claims, the prior dominates and the score sits near the average, because the average is a better guess than five noisy observations. With hundreds of claims, the influencer's own record dominates and the score reflects their actual performance.
The effect is exactly what you want. A new influencer with four-out-of-five does not jump to the top. They sit near the middle of the board, because the math correctly says "I do not know enough about this person yet to move them far from average." An influencer with five hundred claims and a genuinely high record rises to the top, because the math now trusts the record. Shrinkage converts a noisy percentage into a credible estimate that knows its own uncertainty.
The constant — fifty claims before the record outweighs the prior — is the load-bearing choice, and it deserves to be defended.
Set it too low and you are back to crowning people on small samples. Set it too high and you never let a real record speak, flattening genuinely different influencers into the average. Fifty is the point at which a weighted accuracy estimate has enough backing that we are willing to let it pull someone meaningfully away from the crowd. It is not magic. It is a judgment about how much evidence a number needs before it is worth acting on, and we would rather err on the side of not ranking than ranking on a guess.
Different surfaces warrant different constants. This one is for a public leaderboard where a wrong rank is a public accusation of credibility. The bar is high on purpose.
Here is the part that surprised me, and the reason for the hard rule.
Shrinkage solves the statistical problem. It does not solve the interpretive one. An influencer with eight claims, even after shrinkage, has still only made eight claims we can score. We can place their shrunken estimate on the board honestly. But a user looking at a rank still reads it as a verdict — "this person is the eleventh most credible influencer we track" — and that verdict rests on eight data points no matter how cleverly we shrink them.
So we do two things. We use shrinkage so that the numbers we do show are not naive. And we refuse to show a rank at all below fifty claims, because below that, even the shrunken number is not a verdict worth delivering. The shrinkage makes the math honest. The floor makes the presentation honest. You need both.
There is a general principle in here that applies far past leaderboards.
Any time you rank entities by a rate computed from samples of wildly different sizes, you are ranking noise as much as signal, and the entities with the fewest samples will be the most extreme — both at the top and the bottom — purely by chance. Shrinkage fixes the math. A minimum-sample floor fixes the communication. Neither is enough without the other, because a correctly-shrunk estimate built on almost nothing is still almost nothing, and almost nothing should not come with a rank attached.
We would rather have a leaderboard of forty people we trust than four hundred people we ranked on guesses. The empty slots are not a failure of data collection. They are a feature of honesty.