Set the unadjusted calculation first
We will use one example throughout: a two-arm, parallel, 1:1 randomized trial with a continuous primary outcome. The target standardized difference is d = 0.50, two-sided α = 0.05, and power is 80%. Using the noncentral t method, an unadjusted two-sample t test requires 64 analyzable participants per arm, 128 in total.
That 128 is not yet a recruitment target. We have not applied the success rule, the planned baseline adjustment, or attrition. The arithmetic may be exact; the result is only as credible as the assumptions entered.
Four hypotheses create two different problems
If four hypotheses are each tested at 0.05 and any one can make the trial successful, the family-wise Type I error increases. With independent tests, the probability of at least one false positive is 1 − (1 − 0.05)4 = 18.55%; correlation changes that exact rate. Bonferroni controls the family-wise error without requiring the dependence structure.2
| Confirmatory hypotheses | α per test | Total n |
|---|---|---|
| 1 · single primary | 0.0500 | 128 |
| 2 · either can qualify | 0.0250 | 156 |
| 4 · any can qualify | 0.0125 | 182 |
Calling an endpoint secondary does not remove multiplicity by itself. Confirmatory secondary claims need a prespecified hierarchy or gatekeeping procedure. FDR may suit a broad exploratory family; it is not a general substitute for a confirmatory error-control plan.
A baseline measure can buy real precision
In a two-arm randomized trial, ANCOVA uses a baseline measure that predicts the follow-up outcome. Here r is the anticipated within-treatment-arm baseline–follow-up correlation. In the simple Borm approach, the unadjusted per-arm size is multiplied by 1 − r², then one participant per arm is added as a small-sample correction.3
| Baseline–follow-up correlation | Total analyzable n | Change from unadjusted |
|---|---|---|
| Unadjusted | 128 | Reference |
| r = 0.40 | 110 | −18 people |
| r = 0.50 | 98 | −30 people |
| r = 0.60 | 84 | −44 people |
This is not a free discount. The approximation assumes a continuous outcome, a prespecified baseline covariate, an appropriate linear model, and a reasonable common-slope assumption. Take r from comparable external data or a cautious pilot estimate, and show a sensitivity range rather than one optimistic value.
Attrition inflation protects headcount, not bias
Our baseline-adjusted target is 84 analyzable participants. If the expected attrition proportion is q, the simplest recruitment calculation is ⌈84 / (1 − q)⌉.
| Expected attrition | Recruitment target | Calculation |
|---|---|---|
| 10% | 94 | ⌈84 / 0.90⌉ |
| 20% | 105 | 84 / 0.80 |
| 30% | 120 | 84 / 0.70 |
This formula preserves the expected number of completers; it does not guarantee 84. In a simple model where each participant completes independently with probability 0.80, recruiting 105 gives only about a 56% chance of reaching at least 84 completers. The smallest target reaching at least 90% assurance is 112, which gives 92.2% under this model. State whether the protocol uses an expected count or a chosen assurance level.
What effect is detectable with a fixed N?
Sometimes capacity is the starting point. With the same noncentral t method, an unadjusted two-group design limited to 100 participants in total has a minimum detectable effect at 80% power of about d = 0.57.5 This is a sensitivity summary, not a hard boundary where 0.56 is invisible and 0.58 is certain to be significant.
If the clinically important target is smaller, do not massage the effect assumption. Extend recruitment, add sites, use a more efficient defensible analysis, or design an explicit feasibility pilot. “Observed power” computed from the effect after the study does not answer this prospective question.
Multicenter is not the same as cluster-randomized
When a site or clinic is randomized—or the intervention is delivered at site level—outcomes inside a cluster are not independent. For a simple parallel cluster design with equal cluster sizes, the design effect is 1 + (m − 1) · ICC.67
Do not automatically apply this multiplier when people are individually randomized within centers; the answer depends on how the planned model handles center. In a cluster-randomized trial, the patient total is not enough: the number of clusters, unequal cluster sizes, cluster dropout, and ICC uncertainty all enter the final power analysis.
Run one scenario from start to finish
Our worked trial has one primary endpoint that defines success, a prespecified baseline-adjusted analysis, defensible external data for r = 0.60, and 20% expected attrition.
- 01Unadjusted two-arm calculationd = 0.50 · α = 0.05 · 80% power128
- 02Single-primary success ruleNo multiplicity adjustment128
- 03Baseline-adjusted analysisANCOVA · r = 0.60 · +2 correction84
- 0420% expected attrition84 / 0.80105
If either of two confirmatory hypotheses could define success, the unadjusted total would start at 156. With the same r = 0.60, the analyzable target would be 102 and the 20%-attrition recruitment target 128. One sentence in the success rule becomes 23 people in the field.
How to write it in the protocol
“For the [design], the [primary outcome and success rule] used a two-sided family α of [value] and power of [value]. Under [target difference/effect and source], the unadjusted analysis required [total/per-arm n]. With the prespecified [baseline covariate], a correlation of [r and source], and [ANCOVA method], the analyzable target was [n]. Expected attrition of [% and source] gave a recruitment target of ⌈n / (1 − q)⌉ = [N]. Across [assumption ranges], the target ranged from [lower] to [upper].”
Also report the unit of randomization, allocation ratio, software or numerical method, and rounding decisions. Readers need a reproducible path to N, not only the final number.1
Sources
- Chan A-W, Boutron I, Hopewell S, et al. SPIRIT 2025 statement: updated guideline for protocols of randomised trials. BMJ. 2025;389:e081477.
- U.S. Food and Drug Administration. Multiple Endpoints in Clinical Trials: Guidance for Industry. 2022.
- Borm GF, Fransen J, Lemmens WA. A simple sample size formula for analysis of covariance in randomized clinical trials. J Clin Epidemiol. 2007;60:1234–1238.
- Little RJ, Cohen ML, Dickersin K, et al. The design and conduct of clinical trials to limit missing data. Stat Med. 2012;31(28):3433–3443.
- Lakens D. Sample Size Justification. Collabra: Psychology. 2022;8(1):33267.
- Campbell MK, Piaggio G, Elbourne DR, Altman DG. CONSORT 2010 statement: extension to cluster randomised trials. BMJ. 2012;345:e5661.
- Hayes RJ, Bennett S. Simple sample size calculation for cluster-randomized trials. Int J Epidemiol. 1999;28(2):319–326.