Resources - RESEARCH LITERACY
Why two studies on the same compound reach opposite conclusions
Two studies disagree because of sample size, chance, differing populations, different outcome measures and selective publication - not because one team lied. Conflict is normal, and knowing why tells you which result to trust.
Two studies on the same compound reach opposite conclusions mostly because they were not really the same study. They enrolled different people, used different sample sizes, measured different outcomes over different timeframes, and were then filtered through a publishing system that prefers exciting results to boring ones. Layer ordinary random variation on top of that and disagreement stops being a scandal and becomes the expected behaviour of science. The useful skill is not picking a side but working out which of these explanations is doing the work.
This is a different question from where a piece of evidence sits on the ladder from test tube to animal to human trial. Our lesson on how to read peptide research without getting fooled covers that ladder, and it is worth reading first. What follows assumes you already sort evidence by type, and deals with the harder problem: two studies on the same rung, in humans, pointing in opposite directions.
Why do small studies swing so wildly?
Small studies are unstable because a handful of unusual participants can dominate the result. In a trial of twelve people, one person who responds unusually well and one who drops out can shift the group average enough to flip the conclusion. Statisticians describe this as low precision: the study is compatible with a wide range of true effects, from harmful to strongly beneficial, and the single number reported is just one point inside that range. The confidence interval, usually printed next to the headline figure, is the honest description of that range - and in small studies it is almost always uncomfortably wide.
There is a nastier consequence. Because small studies need a large apparent effect to reach statistical significance at all, the small studies that do get published tend to report exaggerated effects. This is sometimes called the winners curse. The published record of any young research area is therefore populated by early, small, impressive results that later, larger work steadily shrinks.
What is the difference between significance and effect size?
Statistical significance answers a narrow question: how surprising would this data be if the intervention did nothing? Effect size answers the question you actually care about: how much did it change, and does that amount matter to a human being? These come apart constantly. A very large study can find a statistically significant improvement in a marker that corresponds to a change nobody would ever notice, and a small study can find a genuinely large effect that misses significance because there were not enough participants to rule out chance.
When you read a result, look for the size of the change and the units it is measured in, not just the word significant. A relative figure such as a forty per cent reduction sounds far more impressive than the same result expressed absolutely, which might be a fall from five cases in a thousand to three. Both descriptions are true. Only one tells you what to expect.
What are p-hacking and outcome switching?
Every study involves dozens of small analytical choices: which participants to exclude, which timepoint to treat as primary, whether to adjust for age or weight, which of several questionnaires to lead with. If a researcher tries many combinations and reports the one that came out best, the result is not a discovery, it is a search. This is p-hacking, and it usually happens without any intent to deceive - it feels like refining the analysis.
The related problem is multiple comparisons. Test twenty outcomes at the conventional threshold and, by chance alone, roughly one will look significant even if the compound does nothing. A study that measures thirty markers and reports the three that moved is telling you almost nothing, unless it says up front which markers it intended to test.
Outcome switching is when a study registers one primary outcome before it starts and then reports a different one in the published paper, typically because the original outcome did not cooperate. It is common enough that comparing a paper against its own pre-registration is one of the highest-yield checks a careful reader can do. A trial that declares its primary outcome in advance and reports it plainly, positive or negative, deserves considerably more trust than one that does not.
What is the file-drawer problem?
The file-drawer problem is that studies finding nothing often never get written up or never get accepted, so they sit in a drawer. Journals have historically preferred novel positive findings, and researchers know it. The consequence is structural rather than individual: even if every published paper is honestly conducted, the published literature as a whole is skewed towards positive results, because the negative ones were filtered out before you could see them.
This is why an area with ten small positive papers and no negative ones should make you more suspicious, not less. If a compound really produced reliable effects in small samples, some of those samples would still have come out null by chance. A literature with no failures in it is a literature that is not showing you everything.
What did the replication crisis actually show?
Over the past fifteen years, coordinated projects across psychology, cancer biology and economics took well known published findings and deliberately re-ran them with larger samples and pre-registered plans. A large share failed to replicate, and among those that did survive, the effects were typically much smaller than the original reports. The lesson was not that scientists are dishonest. It was that the ordinary machinery of small samples, flexible analysis and selective publication reliably produces a published record more confident than the underlying evidence justifies.
The practical takeaway is straightforward. Published, peer reviewed and statistically significant are three weak filters, not three guarantees. Independent replication in a different laboratory with a different population is the filter that matters most, and it is the one most compounds never face.
How much does funding matter?
Funding matters, but not in the cartoon way. Industry-funded trials are often larger and better run than academic ones. The influence shows up earlier and more subtly: in which questions get asked, which comparator is chosen, how long follow-up runs, and whether a disappointing trial gets published at all. Reviews across several fields have consistently found that industry-sponsored studies report favourable conclusions more often than independently funded ones on the same questions. Read the funding statement and the author conflict declarations, treat their absence as a flag, and weight accordingly rather than dismissing outright.
Why does the same intervention behave differently in different groups?
Heterogeneity is the plainest and most underrated explanation for conflicting results. A study in trained young men and a study in sedentary older women are not testing the same thing, even with an identical protocol. Baseline status matters enormously: people who start deficient or impaired in some respect often show large improvements, while people who start well within normal range show little, because there is less room to move. Genetics, diet, medication, sex, age and the endpoint chosen all shift the answer.
So when two trials disagree, the first question is not which is right but who was in them. Very often both are right about their own populations, and the conflict exists only in the headlines, which strip the population out.
What do systematic reviews and meta-analyses add?
A systematic review defines a question in advance, searches the literature by an explicit and repeatable strategy, and appraises every study it finds against stated quality criteria. A meta-analysis then combines the numerical results into a single pooled estimate, which is usually more stable than any individual trial and can reveal effects too small for one study to detect. Done well, this is the strongest form of evidence synthesis available.
The limitation is that pooling cannot create quality. If the input trials were small, unblinded and selectively published, the meta-analysis produces a precise-looking number built on those same weaknesses - garbage in, garbage out. Good reviews acknowledge this: they report a measure of how much the studies disagreed with each other, they test whether small studies drove the result, and they grade their own confidence in the conclusion. A pooled figure with no discussion of study quality or heterogeneity is a summary, not an appraisal.
A checklist for two conflicting headlines
- How many people were in each study, and how wide is the confidence interval? Wide intervals mean the study cannot settle the question either way.
- Who was enrolled? Different baseline health, age or training status can explain the whole disagreement without either study being wrong.
- What exactly was measured, and for how long? A blood marker at four weeks and a functional outcome at a year are different claims.
- How big was the effect in absolute terms, and would a person notice it?
- Was the primary outcome declared in advance, and is it the one being reported?
- How many outcomes were tested in total? A long list with a few winners is a search, not a finding.
- Who funded it, and are conflicts declared?
- Has anyone independent repeated it? A result that has never been replicated is a hypothesis.
- Do the null results exist? A field with only positive findings is a field with a full filing cabinet.
Applying that list will rarely give you a clean verdict, and it is not supposed to. The realistic outcome of reading carefully is a more accurate sense of uncertainty: this looks promising but rests on two small trials, or this is well established across large replicated studies. Being able to tell those two states apart is most of research literacy. Certainty is usually the thing being sold to you.
This article is educational content about how research is conducted and reported. It is not medical advice, and nothing here is a recommendation to use, avoid or change any compound, supplement or treatment. Speak to a qualified healthcare professional about your own health.
Frequently asked
- Why do two studies on the same compound disagree?
- Usually because they were not the same experiment. Different sample sizes, different populations, different outcome measures and different lengths of follow-up all move the answer. Add ordinary random variation and selective publication, and two honest teams can report opposite headlines from the same underlying reality.
- Does a p-value below 0.05 mean a result is important?
- No. A p-value speaks to how surprising the data would be if the compound did nothing at all. It says nothing about how big the effect is. A large study can produce a statistically significant result for a change so small that nobody would notice it in daily life.
- What is publication bias?
- Publication bias is the tendency for positive, striking findings to be published while null results stay unpublished in researchers files. Because of it, the visible literature on almost any topic is systematically more favourable than the full set of experiments that were actually run.
- What did the replication crisis actually show?
- Large, deliberate replication projects re-ran well known published findings and a substantial share failed to reproduce, with surviving effects typically much smaller than first reported. It showed that published, peer reviewed and statistically significant is not the same as reliably true.
- Is a meta-analysis always the best evidence?
- Not automatically. A meta-analysis pools studies to get a more stable estimate, but it inherits the flaws of what it pools. If the underlying trials were small, biased or measured different things, the pooled number looks precise while resting on weak foundations.
- How should I react to a single striking study?
- Treat it as a hypothesis rather than a finding. A single result, however dramatic, is one draw from a noisy process. It earns the status of a finding when independent groups repeat it in different populations and get a similar answer.
Related reading
Educational content only. Not medical advice, diagnosis or treatment. Always consult a qualified healthcare professional before changing your health regimen.