A p-value is not the probability that your hypothesis is correct or incorrect. It describes how compatible the observed data are with a statistical model that includes the null hypothesis.
What Is a P-Value?
A p-value is a probability calculated as part of a statistical hypothesis test. It helps evaluate how unusual the observed result would be if the null hypothesis and the assumptions of the statistical model were true.
More specifically, it represents the probability of obtaining results at least as extreme as those observed, assuming the null hypothesis is true.
A smaller p-value indicates that the observed data are less compatible with the null hypothesis under the assumptions of the test. It does not automatically tell you how important the result is.
Start With the Null and Alternative Hypotheses
P-values make more sense when you understand the two hypotheses involved in a conventional significance test.
Null Hypothesis (H₀)
The null hypothesis usually represents no difference, no association, or no effect in the population being studied.
Alternative Hypothesis (H₁)
The alternative hypothesis represents the effect, difference, or association that the analysis is designed to investigate.
Suppose a study compares the average test scores of students using two teaching methods. A simplified set of hypotheses could be:
- H₀: the population mean scores do not differ between the two methods.
- H₁: the population mean scores differ between the two methods.
The statistical test then evaluates the observed data in relation to the null hypothesis.
What Does Statistical Significance Mean?
Before conducting a conventional hypothesis test, a significance level is often selected. It is usually represented by the Greek letter alpha (α).
A commonly used value is:
α = 0.05
If the calculated p-value is less than the chosen alpha level, the result is conventionally described as statistically significant.
For example:
| P-value | Alpha | Conventional Decision |
|---|---|---|
| 0.021 | 0.05 | Statistically significant |
| 0.049 | 0.05 | Statistically significant |
| 0.051 | 0.05 | Not statistically significant |
| 0.180 | 0.05 | Not statistically significant |
The 0.05 threshold is a convention, not a universal rule. Appropriate significance levels can depend on the research context and analytical plan.
How Should You Interpret a P-Value?
Suppose an analysis produces:
p = 0.03
If the null hypothesis and the assumptions of the model were true, the probability of obtaining a result at least as extreme as the observed result would be 3%.
If the study had chosen α = 0.05 in advance, then 0.03 is below that threshold. Under the conventional decision rule, the result would be considered statistically significant and the null hypothesis would be rejected.
We did not say there is a 3% probability that the null hypothesis is true. That is a different statement and is not what the p-value represents.
Three Practical Examples
Comparing Two Groups
An independent-samples t-test compares scores from two groups and produces p = 0.018. The study uses α = 0.05.
Because 0.018 is below 0.05, the difference is statistically significant under the selected threshold. You would then examine the group means, estimated difference, confidence interval, effect size, and research context to understand the result more fully.
A Non-Significant Result
Another analysis produces p = 0.12 with α = 0.05.
Because 0.12 is above 0.05, the result is not statistically significant under that decision rule. The analysis does not provide sufficient evidence to reject the null hypothesis at the chosen significance level.
This should not automatically be rewritten as “the null hypothesis is proven true.”
A Result Close to 0.05
Imagine one analysis produces p = 0.049 and another produces p = 0.051.
Using a strict α = 0.05 rule places the results on opposite sides of the significance threshold, but the numerical difference between the two p-values is very small.
This is one reason researchers should consider the broader evidence rather than treating 0.05 as a boundary between an important finding and an unimportant one.
What a P-Value Does Not Tell You
Many problems with p-values come from asking them to answer questions they were not designed to answer.
- A p-value does not tell you the probability that the null hypothesis is true.
- It does not tell you the probability that the alternative hypothesis is true.
- It does not directly tell you the size of an effect.
- It does not tell you whether a result is practically important.
- It does not guarantee that the study design or statistical model was appropriate.
The p-value is one part of the analysis. Interpretation should also consider the study design, estimates, uncertainty, assumptions, sample size, and subject-matter context.
Statistical Significance Is Not Practical Significance
A statistically significant result can still represent a very small effect. Conversely, an effect that could matter in practice may fail to reach a conventional significance threshold in a study with limited information or a small sample.
Consider a large dataset where two groups differ by only a tiny amount. With enough observations, that small difference might produce a very small p-value.
The correct next question is not simply:
“Is the result statistically significant?” You should also ask: “How large is the estimated effect, how uncertain is that estimate, and does the difference matter in this context?”
This is why effect sizes and confidence intervals are often valuable alongside p-values.
Common P-Value Mistakes
1. Saying p > 0.05 proves there is no effect
Failure to reject the null hypothesis is not the same as proving that the null hypothesis is true. A non-significant result can occur for several reasons, including limited statistical power or imprecise estimates.
2. Treating p = 0.049 and p = 0.051 as completely different
A threshold can be useful for a predefined decision rule, but values immediately above and below it do not suddenly represent completely different bodies of evidence.
3. Assuming a smaller p-value means a larger effect
The p-value is influenced by more than effect magnitude, including sample size and variability. Compare effect sizes when you want to understand the magnitude of a difference or relationship.
4. Reporting only whether p < 0.05
Reporting the actual p-value, when appropriate, usually provides more information than simply labeling a result significant or non-significant.
5. Testing repeatedly until something becomes significant
Trying multiple analyses and selectively reporting the ones that cross a significance threshold can produce misleading conclusions. The analytical approach should be guided by the research design and appropriate statistical reasoning.
How Should P-Values Be Reported?
Reporting requirements depend on the discipline, institution, journal, and style guide. In many contexts, the actual p-value is reported rather than writing only that the result was significant.
Report the result in context
“The groups differed significantly on the measured outcome, p = .032.”
A complete statistical report often includes more than the p-value. Depending on the analysis, you may also need the test statistic, degrees of freedom, descriptive statistics, confidence interval, and effect size.
Always follow the reporting requirements that apply to your particular project.
A Simple P-Value Interpretation Checklist
Before interpreting a significance test, check the following:
What are the null and alternative hypotheses?
Was the statistical test appropriate for the data?
Were the relevant assumptions considered?
What significance level was selected?
What is the actual p-value?
What is the estimated size of the effect?
What does the confidence interval show?
Does the result matter in the research context?
Thinking through all of these questions produces a much better interpretation than reducing the analysis to whether the p-value is above or below 0.05.
Need Help Interpreting Your Analysis?
Share your research question, variables, statistical output, software, and project requirements. The results can then be reviewed in the context of your actual analysis.
