How do you score whether a health influencer is right?

Search for a command to run...

The leaderboard is the part of Panacea Index that gets the most attention, and it is also the part where it is easiest to cheat. If you are not careful, you can produce a number for every influencer that looks scientific and means almost nothing. So before the leaderboard, here is the score: how it is built, what it weighs, and the threshold we set that most leaderboards would rather skip.
Everything starts with a single structured claim: an influencer said that supplement X increases (or decreases, or does nothing to) outcome Y. That is the atom. "Huberman says creatine increases cognitive function" is a claim. "Attia says omega-3 decreases cardiovascular risk" is a claim.
We extract these from transcripts. The model reads what the influencer actually said and assigns a direction — increase, decrease, neutral, or unclear — to each supplement-and-outcome pair it can find. That gives us hundreds of thousands of directional claims across the influencers we track.
A directional claim is a strong, falsifiable thing to attribute to someone. It is also the only kind of claim you can score against research, because the research itself is directional. So that is what we use.
For each supplement-and-outcome pair, we need to know which way the published evidence points. That is the consensus side, built from claims extracted out of the papers themselves. For a given pair we count how many studies found an increase, how many found a decrease, and how many found no effect, and we derive a net direction: increase, decrease, neutral, or mixed when the evidence pulls both ways.
Now the score becomes possible. For every influencer claim, we ask one question: does the direction the influencer stated match the net direction of the research? If yes, the claim aligns. If no, it does not.
Here is the first place it is tempting to cheat, and the first place we refused to.
You could score an influencer against every claim they make. The number would be bigger and feel more complete. But a lot of supplement claims have almost no research behind them. If two papers exist and they disagree, the "consensus" is noise, and scoring an influencer against noise produces a number that is itself noise. You would be grading people against a ruler that bends.
So we set a gate. An influencer claim is only scored when the research behind it has at least three studies. Below that, there is not enough to define a direction with, and we mark the claim unstudied instead of wrong. This shrinks the number of claims we can score, but it means every claim we do score is graded against something real.
The cost is honesty. The benefit is that the score means something.
Not all consensus is equal. A pair with eighty papers behind it is not the same as a pair with four. So the accuracy percentage is weighted by consensus strength — how many studies back the underlying direction. Being right about a heavily researched outcome counts for more than being right about a thin one. Being wrong about a heavily researched outcome counts for more against you.
This pulls the score toward the questions that actually matter, the ones the field has invested in answering, and away from the long tail of barely-studied endpoints where being "right" is partly luck.
The last problem is the one that ruins most leaderboards, and it is purely statistical. If an influencer has five scored claims and happens to get four right, they are at 80%. Put them at the top of the board and you have crowned someone on the basis of five data points. That is not a credibility score. It is a coin flip with a confident headline.
We handle this two ways. First, the headline number uses Bayesian shrinkage: an influencer with few claims is mathematically pulled toward the average of everyone on the board, so a small-sample hot streak cannot rocket them to the top. Second — and more visibly — we simply refuse to rank anyone with fewer than fifty scored claims. They stay on the page, with their data visible, but the rank and the score read as "not enough data yet." We would rather show a gap than fabricate a verdict.
This is the choice that makes the leaderboard trustworthy, and it is the choice that makes it smaller. Both are fine. A leaderboard that ranks forty people honestly is worth more than one that ranks two hundred people on guesses.
The leaderboard is a measure of how often an influencer's directional supplement claims align with the direction the published research points. That is a specific and useful thing. It is not a measure of how smart they are, how good a clinician they are, or whether you should listen to them about anything other than the supplement claims we were able to match.
Someone can score well and still be misleading in ways this metric does not catch — by overclaiming certainty, by ignoring dose, by extrapolating from animal data to humans. Someone can also score lower than they deserve because their claims land on endpoints where the research is thin and the net direction is a coin flip. The score is a signal, not a verdict.
Read it as: of the claims we could check, against evidence strong enough to check against, here is how often this person pointed the same way the research does. That is a question worth answering. It is not the only question.
Every score on the leaderboard links to the claims behind it. You can see what an influencer said, what the research says, and where they diverge. The disagreements are the interesting part — the places where serious people point in different directions and the evidence has something to say about who is closer.