Reality Bending Lab › News › "Your sample is too small": How to Answer to Reviewers
"Your sample is too small": How to Answer to Reviewers
Reviewers questioning the statistical power of a study, while often valid, is common and unoriginal. The concern is frequently combined with a demand for stringent multiple comparisons / tests "corrections". And it is not only reviewers: ethics committees and grant application forms carry the same blanket "justify your sample size" box, usually written as though every study were a two-group comparison with a single confirmatory test.
It is worth being precise about what low power actually does. A small sample does not make a test more likely to be significant on nothing: the false positive rate is whatever α was set to, regardless of N. What it does is make true effects easy to miss, and make the significant results that do come out less trustworthy: a smaller share of them are "real", and those that are come out exaggerated (in fact, effect size is often negatively correlated with sample size) and occasionally with the wrong sign (Type M and Type S errors; Button et al., 2013; Gelman & Carlin, 2014).
The standard answer to power concerns is preregistration (and registered-reports) with a proper power-analysis, and it can be a good answer. It is also, in my experience, a lot less universal than it is made to sound.
It works well for simple designs: two groups, one test, one effect size to guess. As soon as things get more complicated (nested random effects, several predictors, interactions), the analytical solutions run out and you have to go through simulation (Kumle et al., 2021), which is a small project of its own and requires you to make up a lot of parameters that you don't know. And in a neurophysiological experiment, where the effect has to survive a long preprocessing pipeline that adds uncertainty at every step, the "power" of the final test is often not really a number that you can honestly compute at all.
There is also a deeper problem. Power is a creature of the frequentist NHST framework: it is defined as the probability of rejecting the null hypothesis given that it is false. It only means something if the goal is a binary decision about a null hypothesis. If we are trying to move away from that way of doing science (and I think we should), then being asked for a power analysis is being asked to justify ourselves in a language that we are trying to stop speaking.
The whole power-analysis thing has honestly turned into a bit of a role-play theatre. The question is asked in a blanket fashion and gets answered in a blanket fashion: a G*Power screenshot, an effect size borrowed from a loosely related paper (if not conjured from thin air), and a ticked box. Nobody involved necessarily understands what the numbers mean, and quite often they don't mean anything.
Which is a shame, because there is a perfectly good answer available: practical constraints. Sample sizes are, in real life, set by the money, by the hours of testing that the staff can actually do, by the length of a funded project or of a student's placement, or by how many patients with that condition exist within driving distance. Saying that is legitimate: a resource-constrained justification is an accepted way of justifying a sample size (Lakens, 2022), and it is far more honest, and more informative, than a power calculation retrofitted to the N that we were always going to get.
Maximize the power you have
The best thing to do when the sample size is not massive is to make the most of what we have. The biggest gains are usually not in the analysis but in the design and the measurement. An unreliable task attenuates every effect measured with it, so a noisy paradigm and a small sample make each other worse (we wrote a whole post on how to assess task reliability). And use all the information that you have, in particular the information within participants: if you have multiple trials per participant, don't average over them! Averaging throws away the very thing that tells the model how certain each participant's score is. Use mixed-models instead.
Report uncertainty
This, I think, is the real answer to the criticism. A small sample does not invalidate an estimate, but widens it. Reporting and taking that into account is a big step forward. Fitting the model in a Bayesian framework gives you the whole posterior distribution of each effect, and you can report it: the median, the credible interval, the Probability of Direction (pd), how much of it falls inside a Region of Practical Equivalence (ROPE). With 20 participants the interval will be wide, and the reader can see that it is wide.
Moving under a Bayesian framework can also make non-significant resuls more informative. We can quantify the evidence against an effect, with Bayes Factors or with the posterior mass inside the ROPE. Priors can also be leveraged: narrower priors centred on 0 will pull the estimates towards it ("shrinkage"), which makes it harder for noise to look like an effect. And if a binary call really has to be made, then we can also say we will be stricter about it: a lower p-value threshold (Ioannidis, 2018), a Bayes Factor above 30 rather than 10, a wider ROPE, etc.
Read the results as a pattern
Many studies yield an essemble of results, that should be looked at together, including the non-significant and "trending" effects. In a small sample study, I would trust a conclusion built on 10 independent indices of "depression", all barely significant and all going in the same direction, much more than one super-significant (and possibly cherry-picked) result interpreted as proof of the effect of "depression" in general.
Be honest about it
Acknowledge the limited power instead of trying to hide it, and temper the conclusions accordingly. And if a number is wanted, report a sensitivity analysis (the smallest effect that the design could have detected) rather than observed power: post-hoc power is just a transformation of the p-value and tells you nothing you did not already know (Lakens, 2022).
What to actually write
Here's an example of how to address the issue:
We agree that the sample size limits the strength of the conclusions that can be drawn, and we have made this explicit in the Limitations. Recruitment was constrained by [reason], and increasing N was not feasible. Rather than claim that the study is adequately powered, we have re-analysed the data using trial-level mixed models, and we now report the full posterior of each effect (point estimate, credible interval, probability of direction), so that the remaining uncertainty is visible instead of being hidden behind a threshold. We have also added a sensitivity analysis, and removed the wording that presented any single test as decisive.
Note that this never claims that the sample is fine. It agrees with the reviewer, and then shows what was done about it, which is in my experience a much shorter road than arguing.
Science must be open, replicable & reproducible, slow, and based on methodological best-practices. But first and foremost, it must be honest and transparent.
References
- Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafò, M. R. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365-376.
- Gelman, A., & Carlin, J. (2014). Beyond power calculations: Assessing Type S (sign) and Type M (magnitude) errors. Perspectives on Psychological Science, 9(6), 641-651.
- Ioannidis, J. P. (2018). The proposal to lower P value thresholds to .005. JAMA, 319(14), 1429-1430.
- Kumle, L., Võ, M. L.-H., & Draschkow, D. (2021). Estimating power in (generalized) linear mixed models: An open introduction and tutorial in R. Behavior Research Methods, 53, 2528-2543.
- Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), 33267.
- Makowski, D., Ben-Shachar, M. S., Chen, S. H. A., & Lüdecke, D. (2019). Indices of effect existence and significance in the Bayesian framework. Frontiers in Psychology, 10, 2767.
Thanks for reading! Do not hesitate to share this post. 🤗