Writing about numbers without overclaiming
The sentences that get past reviewers and should not.
Causal language where the design does not support it
The commonest overclaim in the field, and it usually happens in the discussion rather than the results. Correlational and uncontrolled designs support statements about association, prediction and covariation — not about effect, impact, influence, cause, leads to, produces or improves.
| Design | Warranted verbs | Not warranted |
|---|---|---|
| Cross-sectional correlation | associated with, related to, covaries with | predicts, causes, leads to, improves |
| Longitudinal, no manipulation | predicts, precedes, is prospectively associated with | causes, produces, results in |
| Uncontrolled pre-post | changed over the course of, was followed by | was effective, improved, worked |
| RCT | caused, produced, was more effective than | — (but specify: than what comparator?) |
| Meta-analysis of RCTs | the pooled effect of X versus Y was… | X is the best treatment for Y |
Six other things reviewers should catch and often do not
- 1Non-significant does not mean equivalent. A null result means the study did not detect a difference; only an equivalence or non-inferiority design, powered for it, supports "no difference".
- 2A significant difference between "significant" and "non-significant" is not a significant difference. To claim that an effect is larger in one group than another you must test the interaction, not compare two p values.
- 3Effect sizes need confidence intervals. A d of 0.6 with a CI from 0.05 to 1.15 is not the same finding as a d of 0.6 with a CI from 0.5 to 0.7, and reporting only the point estimate conceals the difference.
- 4Cohen’s benchmarks are conventions, not thresholds. Cohen said so. What counts as a meaningful effect depends on the outcome, the cost and the comparator.
- 5Statistical significance is not clinical significance. Report reliable change and the proportion crossing into the functional range, not only group means.
- 6Report the comparator in the sentence. "CBT produced a large effect" is uninterpretable; "CBT produced a large effect relative to a waitlist control" is a finding.
Sentences worth having in your fingers
Because participants were not randomly assigned, these associations cannot support a causal interpretation; alternative explanations including [specific confound] cannot be excluded.
The confidence interval includes effects ranging from clinically trivial to substantial, so the magnitude of the effect remains uncertain despite statistical significance.
Sources
Cumming, G. (2014). The new statistics: Why and how. Psychological Science, 25(1), 7–29
articleGelman, A., & Stern, H. (2006). The difference between "significant" and "not significant" is not itself statistically significant. The American Statistician, 60(4), 328–331
articleWasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133
guideline