Study design
Sample size for a survey or questionnaire study: what to power for and what to report
Co-founder & CEO, SutrixSeptember 18, 2026 · 7 min read
Short answer
Power the study for one primary comparison, not for the survey as a whole. Decide the outcome, the comparison, and the smallest difference worth detecting, then compute the sample for that test at the usual 80 or 90 percent power and 5 percent two-sided alpha. Inflate the result for expected non-response, incomplete questionnaires, and any clustering in how participants were recruited. Report the inputs and the software in the methods so a reviewer can reproduce the number.
Start with one comparison
A questionnaire study can support dozens of comparisons, and a sample size cannot be calculated for all of them at once. Choose the primary one: the outcome score, the groups or exposure being compared, and the statistical test that will answer it. Everything else in the study is secondary or exploratory and does not drive the sample size, although it should be labelled that way in the analysis plan.
If the study is descriptive, for example estimating the prevalence of a symptom or the mean score in a population, the calculation targets the precision of that estimate instead. You choose how wide a confidence interval you can accept and solve for the sample that delivers it.
Choose an effect size you can defend
The effect size is the input that matters most and the one most often chosen carelessly. For a validated instrument, the best anchor is a published minimal clinically important difference for that score. Failing that, use the difference or correlation reported in prior studies of the same population. A generic small, medium, or large effect from a textbook is the weakest choice and should be a last resort, stated as such.
Along with the effect size you need its scale: the standard deviation of the score in a comparable sample for a continuous outcome, or the expected proportions in each group for a binary one. Instrument development papers and prior studies usually supply these.
Inflate for the realities of survey data
The number the calculator returns is the number of analyzable responses, not the number of people to invite. Divide by the expected response rate to get invitations. Then allow for participants who start but do not finish, and for questionnaires that cannot be scored because too many items are missing. Response and completion rates from your own pilot or from a comparable study in the same setting are the right inputs.
If participants were recruited in groups, such as clinics, teams, or classrooms, responses within a group tend to be more alike than responses across groups. That clustering reduces the effective sample size and the calculation should account for it with a design effect based on the expected within-cluster correlation and cluster size.
- Analyzable responses from the power calculation.
- Divided by the expected response rate for the number to invite.
- Plus an allowance for incomplete and unscorable questionnaires.
- Times a design effect if recruitment was clustered.
Validation studies follow different rules
If the purpose is to validate or adapt an instrument rather than to test a hypothesis, the sample size is governed by the psychometric methods. Factor analysis and item response models need larger samples than a group comparison, and the guidance is usually stated as a ratio of participants to items or as an absolute minimum. Reviews of published validation studies show wide variation in practice, so cite the specific guidance you followed.
Write it into the methods
One paragraph is enough: the primary outcome and comparison, the effect size and where it came from, the standard deviation or proportions used, alpha and power, the resulting analyzable sample, the inflation factors applied, and the software or formula. A reader should be able to reproduce the number from that paragraph alone. If the study was completed with fewer participants than planned, say so and report the achieved power or, better, focus on the confidence intervals.
Common questions
What if there is no prior study or MCID for my instrument?
Run a small pilot to estimate the standard deviation and response rate, and choose the smallest difference your clinical colleagues would act on. Report the choice openly as a judgment rather than borrowing a generic effect size.
Can I calculate the sample size after the data are collected?
A post hoc power calculation from the observed effect adds nothing the confidence interval does not already show, and reviewers know it. Report confidence intervals and, if useful, the sample size that was planned and why it was not reached.
Does a larger sample fix a poorly chosen instrument?
No. Sample size controls random error. A questionnaire that does not measure what you think it measures produces the same wrong answer more precisely.
Sources
- 1.Schulz KF, Grimes DA. Sample size calculations in randomised trials: mandatory and mystical. Lancet. 2005.
- 2.Charan J, Biswas T. How to calculate sample size for different study designs in medical research. Indian J Psychol Med. 2013.
- 3.Anthoine E, et al. Sample size used to validate a scale: a review of publications on newly-developed patient reported outcomes measures. Health Qual Life Outcomes. 2014.
- 4.Faul F, et al. G*Power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav Res Methods. 2007.
- 5.Killip S, Mahfoud Z, Pearce K. What is an intracluster correlation coefficient? Crucial concepts for primary care researchers. Ann Fam Med. 2004.