This article is for informational purposes only and does not constitute medical advice. Consult a qualified healthcare provider before making health decisions based on this content.
By HealthDataConsortium.org Research Team | Last verified: August 2026
In This Article
- The Question: What Do P-Values Actually Tell Us?
- The Mechanism: How P-Values Work in Statistical Testing
- Current Evidence: How P-Values Are Used and Misused
- Evidence Table: Key Research on P-Value Use and Misinterpretation
- Practical Implications: What P-Values Mean for Patients and Consumers
- Limitations and Gaps: What We Don't Know
- Related Topics on HealthDataConsortium.org
The Question: What Do P-Values Actually Tell Us?
When you read that a new treatment shows results “statistically significant at p = 0.03,” what does that actually mean? This article explains the real definition of p-values, how they're calculated, why they're essential in medical research, and—most importantly—why they're so often misunderstood. Understanding p-values helps you evaluate whether health claims are based on solid evidence or statistical artifacts.
The Mechanism: How P-Values Work in Statistical Testing
Defining the P-Value in Plain Language
A p-value is the probability of observing results as extreme as (or more extreme than) what you actually found, assuming the null hypothesis is true. The null hypothesis is the assumption that there is no difference between groups—for example, that a new drug works no better than a placebo. Think of it this way: if you flipped a coin 100 times and got 60 heads, the p-value would answer the question, “How likely is it to get 60 or more heads if this coin is actually fair?” A p-value of 0.03 means there's a 3% probability of seeing results this extreme if the null hypothesis were true.
Critically, a p-value does not tell you the probability that your hypothesis is true. This is the most common mistake. A p-value of 0.03 does not mean there's a 97% chance the treatment works. It means: “If the treatment has no effect, we'd see these results 3% of the time just by random chance.” This distinction is profound and often missed even in medical literature.
The 0.05 Threshold and Its Origins
The standard cutoff of p < 0.05 was established by statistician Ronald Fisher in the 1920s as a somewhat arbitrary convention. It means researchers are willing to accept a 5% false-positive rate—the risk of claiming an effect exists when it actually doesn't. This threshold became embedded in medical journals and regulatory agencies, but it's a convention, not a law of nature. Some fields use p < 0.01 for higher confidence; others now argue that p < 0.05 is too permissive and contributes to irreproducible results.
Statistical Power and Type I/Type II Errors
Two types of errors can occur in hypothesis testing. A Type I error (false positive) occurs when you conclude an effect exists when it doesn't—this is the p-value threshold's concern. A Type II error (false negative) occurs when you fail to detect a real effect. The power of a study is its ability to detect a true effect if one exists; adequate power typically requires sample sizes large enough to detect meaningful differences. A study can show p = 0.06 (not “significant”) and still represent real clinical benefit if the study was underpowered due to small sample size.
Current Evidence: How P-Values Are Used and Misused
Large-Scale Analysis of P-Value Reporting
A 2015 analysis published in PLOS Biology examined over 250,000 papers across multiple scientific disciplines and found that p-values cluster suspiciously just below 0.05—suggesting potential manipulation or selective reporting. This “p-hacking” occurs when researchers test multiple hypotheses and report only those reaching statistical significance, inflating false-positive rates.
In clinical medicine, a 2016 review in JAMA found that among published studies showing “statistically significant” results, approximately 25-50% fail to replicate in larger, more rigorous follow-up studies. This replication crisis highlights that p < 0.05 alone is insufficient evidence for clinical practice.
Effect Size Versus Statistical Significance
A study might show p = 0.001 (highly statistically significant) but an effect size so small it's clinically meaningless. Conversely, a study might show p = 0.08 (not statistically significant) but an effect size large enough to warrant clinical consideration if the sample was small. Modern evidence standards increasingly require reporting both the p-value and the effect size (e.g., relative risk reduction, Cohen's d, or absolute risk difference).
Multiple Testing Problems
When researchers test many hypotheses or outcomes without adjustment, the probability of finding at least one false positive increases dramatically. Testing 20 independent hypotheses at p < 0.05 gives a 64% chance of at least one false positive. Proper statistical adjustment (e.g., Bonferroni correction) reduces this risk but requires predefined primary outcomes before data analysis.
Evidence Table: Key Research on P-Value Use and Misinterpretation
| Study/Source | Year | Design & Size | Key Finding | Evidence Grade |
|---|---|---|---|---|
| Ioannidis, J.P.A., “Why Most Published Research Findings Are False” (PLOS Medicine) | 2005 | Meta-analytical review; examined bias in medical research design | Demonstrated that in many fields, the majority of published findings are false positives due to bias, small sample sizes, and flexible statistical practices | High (foundational) |
| Wasserstein & Lazar, “The ASA's Statement on p-Values” (American Statistician) | 2016 | Expert consensus statement from American Statistical Association | Official guidance stating p-values should not be used alone to determine significance and do not measure effect size or practical importance | High (authoritative) |
| Open Science Collaboration, “Estimating the Reproducibility of Psychological Science” | 2015 | Replication study of 100 published psychology studies | Only 36% of studies replicated with statistically significant results; large effect sizes in originals often reduced in replication | High (empirical) |
| Colquhoun, D., “An Investigation of the False Discovery Rate and the Misinterpretation of P-Values” (Royal Society Open Science) | 2014 | Mathematical and statistical analysis | If prior probability of hypothesis is 10%, p = 0.05 still yields 26% false-positive rate; if prior probability is 1%, false-positive rate is 76% | High (theoretical) |
| FDA Guidance on Clinical Trial Multiplicity Issues | 2017 | Regulatory guidance for pharmaceutical development | Recommends prespecified primary endpoints and correction for multiple comparisons to avoid false discoveries in drug approval processes | High (regulatory) |
Practical Implications: What P-Values Mean for Patients and Consumers
Evaluating Health News and Claims
When a headline announces, “New Study Proves Treatment X Works (p = 0.045),” you should ask: What is the actual effect size? How large was the study? Was this the primary outcome or a secondary finding? Was this registered before the study began (pre-registration prevents p-hacking)? A single study with p < 0.05 is preliminary evidence, not proof. Robust evidence requires replication, larger sample sizes, and consistency across multiple studies.
Distinguishing Statistical from Clinical Significance
A blood pressure medication might lower systolic pressure by 2 mmHg with p = 0.01 (statistically significant in a large trial), but a 2 mmHg reduction may have negligible clinical impact. Conversely, a weight-loss intervention reducing body weight by 8 kg (p = 0.08 in a small study) might be clinically meaningful even if not “statistically significant.” Always check the absolute difference, not just the p-value.
Risk Communication
Evidence-based medicine now emphasizes absolute risk reduction and number needed to treat (NNT) over relative risk and p-values. For example, “This drug reduces your risk of heart attack from 2% to 1%” (absolute reduction of 1%) is clearer than “40% relative risk reduction” and more useful than a p-value. Ask your healthcare provider for these metrics.
Limitations and Gaps: What We Don't Know
P-Value Alternatives Still Developing
Bayesian statistics and confidence intervals offer alternatives or complements to p-values but require more computational sophistication and have not yet replaced frequentist p-values in mainstream medical publishing. Pre-registration of studies and transparent reporting of all outcomes (not just significant ones) are emerging best practices but not yet universal in clinical research.
Heterogeneity Across Disciplines
Different medical fields apply p-value thresholds inconsistently. Rare disease research may accept p < 0.10; genomics now routinely applies p < 5 × 10-8 thresholds to account for multiple testing. Guidelines for when and how to adjust for multiple comparisons remain debated.
Industry and Publication Bias
Studies funded by pharmaceutical companies are more likely to show positive (statistically significant) results than independently funded studies—not necessarily because the drug doesn't work, but because failed trials are less likely to be published. This distorts the apparent strength of evidence in the published literature.
Related Topics on HealthDataConsortium.org
- Confidence Intervals and Effect Sizes: How to interpret ranges of possible effects beyond binary significant/not significant judgments
- Study Design and Evidence Hierarchy: Why randomized controlled trials provide stronger evidence than observational studies
- Replication Crisis in Medical Research: Why published findings often fail to reproduce and what this means for clinical practice
- How to Read a Clinical Trial: Step-by-step guide to evaluating primary outcomes, endpoints, and quality metrics
- Bayesian Statistics in Medicine: Alternative frameworks for interpreting probability and evidence in clinical research
This article is for general information purposes only and does not constitute medical advice. Consult your doctor or qualified healthcare provider before making changes to your health routine.

