How we ranked every supplement by health goal

Search for a command to run...

The "best supplements for sleep" list — and the equivalent for muscle, cognition, longevity, and every other health goal we cover — took seven versions to get right. That is not because ranking is hard. Ranking is easy. The hard part, which we got wrong six times before we got it right, was deciding what even counts as a sleep supplement, and making sure the ranking measured focus rather than fame.
This is the methodology behind those lists, and the story of why it kept breaking until we understood what we were actually measuring.
The obvious way to rank supplements for a health goal is to count. For sleep: which supplements have the most studies touching a sleep outcome. Rank by that count. Done.
This fails immediately, and the failure is instructive. Rank by raw count and the top of every list looks identical. Vitamin D. Magnesium. Omega-3s. Not because they are the best supplements for sleep, but because they are the most-studied supplements for everything. They have so much research, across so many outcomes, that they dominate any list sorted by volume, regardless of the topic. A ranking that puts the same three supplements at the top of sleep, muscle, mood, and vision is not ranking supplements for those goals. It is ranking supplements for fame, and calling it something else.
The whole point of a "best for sleep" list is that it is different from "best for muscle." If your method produces the same list for every goal, your method is measuring the wrong thing.
The insight, when it finally landed, was that we did not want to measure volume. We wanted to measure focus — how concentrated a supplement's research is on the goal in question, relative to all the other things that supplement is studied for.
Melatonin is the canonical example of focus. Almost all of melatonin's research is about sleep. If you ask "how sleep-focused is melatonin," the answer is very. Vitamin D is the canonical example of the opposite. Vitamin D has sleep studies, but it also has bone studies, immune studies, mood studies, metabolic studies, on and on. Ask how sleep-focused vitamin D is and the answer is barely, even though it has more sleep studies in absolute terms than many genuinely sleep-focused supplements.
So the score we settled on is a purity score. For each supplement, within a health goal, we take the volume of its research on that goal and multiply it by the fraction of its total research that the goal represents. Volume times concentration. A supplement scores highly when it has both a meaningful body of research on the goal and that body represents a large share of what the supplement is studied for. Creatine dominates the muscle list not just because it is heavily studied, but because so much of its research is about muscle. Lutein dominates vision for the same reason.
This single change is what made the lists different from each other, and different from a fame ranking. Focus, not volume.
Purity scoring has a failure mode of its own, and handling it is the second key decision.
If you rank purely by concentration, you get a different problem. A supplement with three studies, all of them on sleep, is technically 100 percent sleep-focused. It would top the list on purity alone, despite having almost no evidence behind it. A supplement studied once, for sleep, perfectly, would outrank melatonin. That is absurd.
So we set a floor. A supplement only enters the ranking for a goal when it has at least a meaningful body of research on that goal — fifteen claims, in our case. Below that, there is not enough to trust the purity, and the supplement is excluded rather than ranked on a whisper. The floor cuts the single-study curiosities and the noise, and leaves the ranking to supplements with enough evidence that focus means something.
Floor plus purity. Volume has to clear a bar to enter, and then concentration determines where you sit. Together they produce lists that are neither fame rankings nor noise rankings, but honest representations of where the research concentrates for each goal.
There is a subtler problem that took us longer to see, and it is about membership rather than ranking.
To rank supplements for sleep, you first have to decide which outcomes count as sleep outcomes. "Sleep quality" obviously. "Sleep latency," sure. But "insomnia," "circadian rhythm," "restless legs," "melatonin secretion" — these all relate to sleep, and whether they get pulled into the sleep bucket changes who appears on the list. Decide membership badly and you rank the wrong set of supplements before the ranking math even runs.
We tried two methods for membership. The first was embeddings — measuring the semantic similarity between each outcome and the concept of sleep, and pulling in everything above a threshold. This mostly worked for health goals. It failed spectacularly for supplement-type categories like peptides, where it kept classifying melatonin as a peptide because melatonin is mentioned near peptide-related words. The embedding had confused co-occurrence with membership.
The fix was to use different methods for different kinds of category. For health goals, where the members are outcomes, embeddings work well — sleep, muscle, cognition are fuzzy concepts that similarity captures. For supplement types, where the members are supplements, we use an explicit, curated membership list. Melatonin is not a peptide because we say so, definitively, in a list, and no embedding can override that. The lesson was that the right method depends on what you are classifying, and that using one method everywhere is a mistake dressed up as consistency.
I want to be honest that this was not a clean design that worked the first time. It was seven iterations, each of which fixed a real failure and exposed the next one. The first version ranked by raw volume and produced identical lists for every goal. The second added gating but still ranked by volume. The third tried pure purity and let single-study supplements win. The fourth added the floor. The fifth fixed how supplements were matched to their names, because fuzzy matching was letting the wrong entries in. The sixth merged synonyms that were being counted separately. The seventh got the supplement-type categories onto curated membership instead of embeddings.
Each version is preserved. We never delete the old rankings. The newest is the cleanest, and the history is the record of how we learned what the ranking actually needed to measure. That is, I think, the only honest way to build a system that ranks things people will act on — you keep the versions, you explain what changed, and you let the latest be the best while admitting it is not the last.
The supplement-for-sleep list is a list of supplements whose research concentrates on sleep, above a floor of evidence, with membership decided by the method appropriate to the category. It is not a list of the most effective sleep supplements. Effectiveness is a question of direction and agreement, and we answer it per supplement, on the consensus page, with the strength of the evidence shown and its uncertainty intact.
Use the list to find where the field has focused its attention for a goal you care about. Then go to the consensus pages to see what that attention actually found. The ranking points you toward the right supplements to investigate. The investigation is still your job, and ours, and the point of the whole thing.