What statistical significance tells us
Statistical significance is commonly judged using a threshold such as p<0.05. The result depends on the size of the observed difference, variability in the data, sample size and the analysis used.
With a very large sample, a small and unimportant difference may be statistically significant. With a small sample, a potentially valuable effect may remain uncertain. Statistical significance should therefore be treated as one part of the evidence.
Look at the size and precision of the effect
The estimated between-group difference answers a more useful question: how much additional change was associated with treatment? A confidence interval shows the range of effects reasonably compatible with the data and makes uncertainty visible.
Standardised effect sizes can help compare outcomes measured on different scales, although they still require clinical interpretation. Absolute differences are usually easier for clinicians, sponsors and consumers to understand.
Clinically meaningful change
Meaningfulness can be considered at both group and individual levels. A minimal important difference estimates the smallest change likely to matter, while a responder analysis identifies the proportion of participants reaching a prespecified threshold.
Responder thresholds should be justified before results are known. Post-hoc definitions chosen because they produce an attractive result can exaggerate certainty.
- Report the estimated effect with its confidence interval
- Compare the effect with an established or justified meaningful-change threshold
- Show responder rates and, where useful, the number needed to treat
- Consider benefits alongside adverse events, burden, cost and adherence
- Report all prespecified primary and key secondary outcomes
Build interpretation into the protocol
The outcome hierarchy, meaningful-change criteria and principal analysis should be specified before unblinding. This reduces selective interpretation and ensures the sample size is aligned with the effect the study is intended to detect.
The strongest conclusion integrates statistical evidence, clinical importance, consistency across outcomes, study quality and the totality of prior research.
The central question is not merely whether an effect exists, but how large it is, how certain we are and whether it matters.
