Two studies can report the same p-value and support very different decisions. One may estimate a large benefit with wide uncertainty; another may estimate a small benefit precisely. The p-value alone does not show the effect size, its clinical importance, or the range of values that remain reasonably compatible with the data and model.
A confidence interval places the estimate and its precision on the outcome scale. Used carefully, it helps readers ask better questions: What is the best estimate? How uncertain is it? Does the interval include effects large enough to matter, effects too small to matter, no effect, or possible harm?
01Begin with the estimate
Suppose a study estimates that an intervention lowers mean systolic blood pressure by 4 mm Hg, with a 95% confidence interval from 1 to 7 mm Hg. The point estimate is 4. The interval communicates the sampling uncertainty under the analysis assumptions. It does not mean there is a 95% probability that the fixed true effect lies inside this particular interval under the usual frequentist interpretation.
The interval also does not account automatically for bias, missing data, outcome switching, model misspecification, measurement error, or multiplicity. A narrow interval around a biased estimate remains misleading. Interpretation must stay connected to design and analysis quality.
02Compare the interval with meaningful thresholds
Statistical significance asks whether a result crosses a reference threshold under a model, often the null. Clinical interpretation asks whether plausible effects are important. Those questions are related but not identical.
Before analysis, define a smallest effect that would influence practice, policy, or further research when such a threshold can be justified. Then interpret the interval against both the null and that meaningful-effect threshold.
Four broad patterns are useful:
- The interval excludes no effect and contains only effects judged meaningful.
- It excludes no effect but includes effects too small to matter.
- It includes no effect and both meaningful benefit and negligible effects.
- It includes benefit, no effect, and meaningful harm.
These patterns lead to different conclusions even if a binary “significant/not significant” label treats some of them alike.
03Let width influence confidence in the decision
Interval width is an expression of precision, not study quality by itself. Wide intervals often arise from small samples, few events, high variability, or inefficient designs. They indicate that the data leave a broader range of effects unresolved. A “non-significant” result with a wide interval should not be described as proof of no effect if clinically important benefit or harm remains compatible with the data.
Conversely, a very large study can produce a narrow interval excluding the null around an effect too small to matter. Statistical detectability should not be mistaken for decision relevance.
The DELTA² guidance on sample-size calculations connects these ideas at the design stage. Even when a study is powered for a target difference, examining the expected confidence-interval width can reveal whether the eventual estimate will be precise enough for the intended decision.
04Report the correct interval
Match the interval to the estimand and effect measure. A risk difference, risk ratio, odds ratio, hazard ratio, mean difference, or absolute rate answers a different question. Whenever possible, include an absolute effect because relative effects can appear similar across settings with very different baseline risks.
Specify the confidence level, method, analysis population, and whether the interval is adjusted for multiplicity. For subgroup analyses, an interval around each subgroup estimate does not replace a formal assessment of interaction. For prediction, intervals around model parameters do not substitute for validation and uncertainty in predicted risk.
05Replace verdicts with calibrated conclusions
The American Statistical Association states that p-values do not measure effect size or importance and should not alone determine scientific or policy conclusions. CONSORT 2025 similarly places estimation and transparent reporting at the center of trial reports.
A useful result sentence therefore combines the effect measure, point estimate, interval, and interpretation: “The estimated reduction was 4 mm Hg (95% CI 1 to 7); the data are compatible with a small-to-moderate benefit, but the lower bound does not reach the prespecified 3 mm Hg threshold for a clearly important effect.” The exact wording depends on context, but the structure keeps uncertainty visible.
Before approving a table, remove the p-values temporarily. Can a reader still see the magnitude, direction, precision, and decision threshold? If not, the report is asking one number to carry more meaning than it contains.
References
- American Statistical Association Statement on P-Values ↗ — effect size, evidential context, full reporting, and limits of threshold decisions. Accessed 27 July 2026.
- CONSORT 2025 Statement ↗ — transparent reporting of trial estimates and analyses. Accessed 27 July 2026.
- DELTA² Guidance on Target Differences and Sample Size ↗ — target differences, power, expected interval width, and reproducible calculations. Accessed 27 July 2026.
- WHO Trial Registration Data Set ↗ — outcome results and measures of precision in trial summary results. Accessed 27 July 2026.
This article is educational and intended for research purposes. It does not provide individual medical advice, diagnosis, or treatment. No patient data were used.