Start with the difference, not d
A power calculation for a comparison of means should not begin with “let's use d = 0.5.” The useful question is more concrete: what size of difference should the study have a high probability of detecting? Define that target in the outcome's own unit and explain why it matters.1
Suppose 1.5 points on a 0–30 cognitive scale would change the clinical interpretation. If the expected pooled within-group SD in a comparable population is 3 points, the standardized effect for two independent groups is d = 1.5 / 3 = 0.50. You did not choose a “medium effect.” You translated a 1.5-point target into SD units.
Effect size is not one statistic
A mean difference, risk difference, risk ratio, and hazard ratio are all measures of effect. Standardized measures such as Cohen's d divide a difference by the relevant variability to produce a unitless number.2
Standardization helps comparisons and power calculations. It does not make an effect clinically meaningful. Whenever possible, report both the target difference in the outcome's own unit and its standardized equivalent.
d, dz, and f: use the right denominator
| Design | Input | Definition |
|---|---|---|
| Two independent groups | d | Target mean difference ÷ expected pooled within-group SD |
| Before–after in the same person | dz | Target mean change ÷ SD of the change scores |
| Three or more independent groups | f | Spread of group means ÷ within-group SD |
| Two proportions | p₁ and p₂ | The two expected proportions; effects may also be expressed as a risk difference, risk ratio, or odds ratio |
d and dz are not interchangeable.3 In a paired design, the change-score SD depends on the before and after variability and on their correlation.
If both raw SDs are 3 and the correlation is 0.70, the change-score SD is 2.32. A target change of 1.5 points gives dz ≈ 0.65. Dividing the same 1.5 by the raw SD to get 0.50 is not the same calculation. Nor is the change SD always smaller than the raw SD; with equal raw SDs, that is true only when the correlation exceeds 0.50.
What are Cohen's benchmarks for?
| Conventional label | d | f | For d, if the relevant SD is 3 |
|---|---|---|---|
| Small | 0.20 | 0.10 | 0.6 points |
| Medium | 0.50 | 0.25 | 1.5 points |
| Large | 0.80 | 0.40 | 2.4 points |
Cohen offered these as rough reference points for settings with no better basis, not as scientific boundaries that replace context.4 The same d can be negligible for one outcome and decision-changing for another.
The labels are also often applied to dz. But dz depends on the change-score SD and therefore on the within-person correlation. Two “medium” effects from different designs are not automatically comparable.
Why does n rise so quickly as the effect shrinks?
With the other assumptions fixed, the required sample size scales roughly with 1 / dz². Halving the target effect therefore requires about four times as many participants. The values below use a paired t test, two-sided α = 0.05, 80% power, and the noncentral t distribution.
| dz | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | 0.8 | 1.0 |
|---|---|---|---|---|---|---|---|
| Complete records | 199 | 90 | 52 | 34 | 24 | 15 | 10 |
A small but important target may require longer recruitment or additional sites; the curve makes that planning cost visible.
Where does a defensible target difference come from?
- Clinical importance: Define the smallest difference that would matter to patients, clinicians, or decision makers in the outcome's own unit. Use patient- or clinician-anchored thresholds where suitable.
- Comparable external data: Prefer several studies or a systematic review with a similar population, outcome, follow-up, and analysis. Do not copy the observed effect from one small paper as though it were the true value.
- Pilot data, used for the right job: A pilot can inform recruitment, retention, measurement feasibility, the SD, and the within-person correlation. Its observed treatment difference is usually too unstable to be the sole target for a definitive study.5
- An openly declared planning scenario: If no better basis exists, use Cohen's values as scenarios. Do not present them as observed or established effects.
A robust plan should not depend on one point estimate. Make the clinically important target the main scenario, then show the required n for smaller and larger plausible effects.
How should it appear in a protocol?
For a continuous primary outcome, a brief but reproducible methods statement can read like this:
“For the primary outcome, [measure], the target difference was [Δ and unit], justified by [clinical or patient-based rationale and source]. Based on [source], the relevant [pooled within-group/change-score] SD was [value], giving [d or dz] = [Δ / SD] = [value]. Under a [test and design], two-sided α = [value], power = [value], and [allocation ratio], [n] analyzable participants were required [per group/in total]. Across the prespecified [effect or SD range], the required total n ranged from [lower] to [upper].”
Fill every bracket. “A medium effect was assumed” is not a reproducible sample-size method by itself.
Three shortcuts that break the calculation
- Using the wrong denominator. Feeding an independent-groups d into a paired calculation, or confusing the change-score SD with the raw-score SD, gives the wrong n.
- Choosing the effect that makes the study feasible. A large difference seen in a small pilot—or a d selected to fit the budget—reduces power for smaller effects that may still matter. Those effects do not become impossible to observe; the probability of declaring them significant falls below 80%, and their confidence intervals remain wide.
- Computing observed power after the study. Post-hoc power based on the observed effect largely restates the p-value. Report the effect estimate and its confidence interval instead.7
Evidence behind the guide
- Cook JA, Julious SA, Sones W, et al. DELTA² guidance on choosing the target difference and undertaking and reporting the sample size calculation for a randomised controlled trial. BMJ. 2018;363:k3750.
- Lakens D. Calculating and reporting effect sizes to facilitate cumulative science. Front Psychol. 2013;4:863.
- Morris SB, DeShon RP. Combining effect size estimates in meta-analysis with repeated measures and independent-groups designs. Psychol Methods. 2002;7:105–125.
- Cohen J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Lawrence Erlbaum; 1988.
- Albers CJ, Lakens D. When power analyses based on pilot data are biased. J Exp Soc Psychol. 2018;74:187–195.
- Norman GR, Sloan JA, Wyrwich KW. Interpretation of changes in health-related quality of life. Med Care. 2003;41:582–592.
- Hoenig JM, Heisey DM. The abuse of power. Am Stat. 2001;55:19–24.