On this page
A study finishes with 40 patients, the primary result isn't statistically significant, and a reviewer's comment reads: "the study appears underpowered to detect a clinically meaningful difference." That sentence is one of the most preventable rejections in clinical research — preventable months earlier, before a single patient was enrolled, with a calculation that typically takes under an hour.
Sample size and power analysis often get treated as a bureaucratic requirement for the ethics committee rather than a design decision that determines whether the study can succeed at all. This article covers the four inputs the calculation actually needs, and how to get each of them.
Why This Happens
It's seen as a form to complete, not a design choice
Because ethics committees require a sample size justification, it's easy to treat the number as something to produce for the application rather than a figure that should shape recruitment planning from the start.
The expected effect size isn't obvious in advance
Unlike the other three inputs, effect size requires either pilot data or a defensible estimate from prior literature — and without either on hand, the calculation feels impossible to start.
Statistical software looks more intimidating than it is
Tools like G*Power are free and designed for exactly this calculation, but without guidance on which of its many test options applies, they can feel more complex than the underlying decision actually is.
Common Mistakes We See
- Not calculating sample size at all. Recruiting "as many patients as feasible" within a study window is a resourcing decision, not a substitute for a power calculation.
- Using a fixed rule of thumb regardless of design. "30 per group" is sometimes cited as a default, but the correct number depends entirely on the expected effect size, variability, and statistical test — not a single memorized figure.
- Calculating power after the study is finished. Post-hoc power calculations, run to explain a non-significant result after the fact, are statistically discouraged and don't retroactively validate an underpowered study.
- Ignoring anticipated dropout. A sample size calculated for the final analysis needs to be inflated to account for expected loss to follow-up, or the completed study will end up underpowered regardless of the original number.
- Picking an effect size arbitrarily. Choosing a number that produces a convenient sample size, rather than one grounded in pilot data or prior literature, undermines the entire calculation.
Practical Recommendations
Every sample size calculation needs four inputs:
| Input | What it means | Typical value |
|---|---|---|
| Significance level (α) | The acceptable risk of a false-positive result | 0.05 |
| Power (1−β) | The probability of detecting a true effect if one exists | 0.80–0.90 |
| Effect size | The smallest difference considered clinically meaningful | From pilot data or prior published studies |
| Variability | Standard deviation of the outcome (for continuous outcomes) | From pilot data or prior literature |
G*Power (free) is sufficient for the large majority of standard designs — t-tests, ANOVA, chi-square, correlation, and regression-based calculations. Once the test is chosen at the protocol stage, the corresponding G*Power module usually asks for exactly these four inputs.
Quick Checklist
- The statistical test is chosen before the sample size calculation is run
- Effect size is based on pilot data or a specific prior study, not an arbitrary figure
- Significance level and power are explicitly stated (typically 0.05 and 0.80–0.90)
- The calculated number is inflated to account for anticipated dropout
- The calculation is documented in the protocol before ethics submission
- A statistician has reviewed the calculation, even briefly
- The final recruitment target matches the inflated, not the raw, calculated number
Need Expert Guidance?
Book a free call with PureMed and let our research specialists help you save valuable time, avoid costly research mistakes, and move confidently toward publication.
Frequently Asked Questions
What power is generally considered acceptable?
80% is the most commonly used minimum in clinical research, meaning an 80% chance of detecting a true effect if one exists. 90% is increasingly requested by some journals and funders for a stronger study.
Where do I get an effect size estimate before I've collected any data?
From a small pilot study if feasible, or from the closest comparable study in the published literature — using its reported effect size or variability as the basis for your calculation.
Does a small pilot study still need its own power calculation?
Not in the same way. Pilot studies are typically sized for feasibility (recruitment rate, protocol adherence) rather than statistical significance, and are explicitly described as such rather than powered to detect the main effect.
What if I can't realistically reach the calculated sample size?
Consider whether the study question can be answered with a different design that needs fewer participants, whether a multi-center collaboration is feasible, or whether the study should be framed explicitly as a pilot rather than a definitive trial.