Reality Bending LabNews › "Your sample is too small": How to Answer to Reviewers

"Your sample is too small": How to Answer to Reviewers

2021-11-05 · Methods · Dominique Makowski

Reviewers questioning the statistical power of a study, while often valid, is common and unoriginal. The concern is frequently combined with a demand for stringent multiple comparisons / tests "corrections". And it is not only reviewers: ethics committees and grant application forms carry the same blanket "justify your sample size" box, usually written as though every study were a two-group comparison with a single confirmatory test.

It is worth being precise about what low power actually does. A small sample does not make a test more likely to be significant on nothing: the false positive rate is whatever α was set to, regardless of N. What it does is make true effects easy to miss, and make the significant results that do come out less trustworthy: a smaller share of them are "real", and those that are come out exaggerated (in fact, effect size is often negatively correlated with sample size) and occasionally with the wrong sign (Type M and Type S errors; Button et al., 2013; Gelman & Carlin, 2014).

The standard answer to power concerns is preregistration (and registered-reports) with a proper power-analysis, and it can be a good answer. It is also, in my experience, a lot less universal than it is made to sound.

It works well for simple designs: two groups, one test, one effect size to guess. As soon as things get more complicated (nested random effects, several predictors, interactions), the analytical solutions run out and you have to go through simulation (Kumle et al., 2021), which is a small project of its own and requires you to make up a lot of parameters that you don't know. And in a neurophysiological experiment, where the effect has to survive a long preprocessing pipeline that adds uncertainty at every step, the "power" of the final test is often not really a number that you can honestly compute at all.

There is also a deeper problem. Power is a creature of the frequentist NHST framework: it is defined as the probability of rejecting the null hypothesis given that it is false. It only means something if the goal is a binary decision about a null hypothesis. If we are trying to move away from that way of doing science (and I think we should), then being asked for a power analysis is being asked to justify ourselves in a language that we are trying to stop speaking.

The whole power-analysis thing has honestly turned into a bit of a role-play theatre. The question is asked in a blanket fashion and gets answered in a blanket fashion: a G*Power screenshot, an effect size borrowed from a loosely related paper (if not conjured from thin air), and a ticked box. Nobody involved necessarily understands what the numbers mean, and quite often they don't mean anything.

Which is a shame, because there is a perfectly good answer available: practical constraints. Sample sizes are, in real life, set by the money, by the hours of testing that the staff can actually do, by the length of a funded project or of a student's placement, or by how many patients with that condition exist within driving distance. Saying that is legitimate: a resource-constrained justification is an accepted way of justifying a sample size (Lakens, 2022), and it is far more honest, and more informative, than a power calculation retrofitted to the N that we were always going to get.

Maximize the power you have

The best thing to do when the sample size is not massive is to make the most of what we have. The biggest gains are usually not in the analysis but in the design and the measurement. An unreliable task attenuates every effect measured with it, so a noisy paradigm and a small sample make each other worse (we wrote a whole post on how to assess task reliability). And use all the information that you have, in particular the information within participants: if you have multiple trials per participant, don't average over them! Averaging throws away the very thing that tells the model how certain each participant's score is. Use mixed-models instead.

Report uncertainty

This, I think, is the real answer to the criticism. A small sample does not invalidate an estimate, but widens it. Reporting and taking that into account is a big step forward. Fitting the model in a Bayesian framework gives you the whole posterior distribution of each effect, and you can report it: the median, the credible interval, the Probability of Direction (pd), how much of it falls inside a Region of Practical Equivalence (ROPE). With 20 participants the interval will be wide, and the reader can see that it is wide.

Moving under a Bayesian framework can also make non-significant resuls more informative. We can quantify the evidence against an effect, with Bayes Factors or with the posterior mass inside the ROPE. Priors can also be leveraged: narrower priors centred on 0 will pull the estimates towards it ("shrinkage"), which makes it harder for noise to look like an effect. And if a binary call really has to be made, then we can also say we will be stricter about it: a lower p-value threshold (Ioannidis, 2018), a Bayes Factor above 30 rather than 10, a wider ROPE, etc.

Read the results as a pattern

Many studies yield an essemble of results, that should be looked at together, including the non-significant and "trending" effects. In a small sample study, I would trust a conclusion built on 10 independent indices of "depression", all barely significant and all going in the same direction, much more than one super-significant (and possibly cherry-picked) result interpreted as proof of the effect of "depression" in general.

Be honest about it

Acknowledge the limited power instead of trying to hide it, and temper the conclusions accordingly. And if a number is wanted, report a sensitivity analysis (the smallest effect that the design could have detected) rather than observed power: post-hoc power is just a transformation of the p-value and tells you nothing you did not already know (Lakens, 2022).

What to actually write

Here's an example of how to address the issue:

We agree that the sample size limits the strength of the conclusions that can be drawn, and we have made this explicit in the Limitations. Recruitment was constrained by [reason], and increasing N was not feasible. Rather than claim that the study is adequately powered, we have re-analysed the data using trial-level mixed models, and we now report the full posterior of each effect (point estimate, credible interval, probability of direction), so that the remaining uncertainty is visible instead of being hidden behind a threshold. We have also added a sensitivity analysis, and removed the wording that presented any single test as decisive.

Note that this never claims that the sample is fine. It agrees with the reviewer, and then shows what was done about it, which is in my experience a much shorter road than arguing.

Science must be open, replicable & reproducible, slow, and based on methodological best-practices. But first and foremost, it must be honest and transparent.

References


Thanks for reading! Do not hesitate to share this post. 🤗

Keep reading

Click to Open
Reality Bending Lab

Exploring the neuropsychology of reality and its distortions.

The Reality Bending Lab, led by Dominique Makowski at the University of Sussex in Brighton, UK, conducts world-leading research on reality perception, illusions, fake news, AI-beliefs, deception and its links with emotions, cognitive control and the Self. Using recording signals from the body (ECG, EDA…) and the brain (EEG), we analyse data using advanced modelling (Bayesian statistics, chaos theory, computational models) and develop open-source tools to improve neuropsychological science.

“Do not try to bend the spoon — that’s impossible. Instead, only try to realise the truth: there is no spoon. Then you'll see that it is not the spoon that bends, it is only yourself.”

People

Information

Where to find us

The Reality Bending Lab is based in the School of Psychology, at the University of Sussex in Brighton, UK.

Pevensey 1 - room 2B7, School of Psychology, University of Sussex, BN1 9QH

The University of Sussex campus in the South Downs University of Sussex crest

University of Sussex

Brighton is the sunniest city in the UK and, by some margin, its most cheerfully strange — on the sea, an hour from London. The University of Sussex sits just above it in the South Downs, and its School of Psychology is one of the largest and strongest in the country.

It is famous for its research on consciousness and interoception, with scientists like Anil Seth, Andy Clark, Zoltan Dienes and Hugo Critchley, its teaching of cutting-edge psychological methods (with stats rockstar Andy Field), its history of multidisciplinary research (arts and science), and its radical engagement in open and disruptive science.

Join this community

Hire us Join the Lab Like this website? See how to get yours