There's no fixed number — not 10, not 30, not 100 — that separates a sample "worth analyzing" from one that isn't. What actually determines it is whether your sample can detect an effect size you'd genuinely care about, and if it can't, whether you're asking a confirmatory question at all or something a small sample can still answer honestly.
The real question is power, not headcount
A statistical test doesn't fail because a sample is "small" in the abstract. It fails because the sample is too small, given the variability in your outcome, to reliably detect the size of effect you're actually looking for. That's what a power analysis computes: the smallest effect your design has a reasonable chance (conventionally 80%) of detecting, given your sample size and the outcome's expected variability. Two studies with the same headcount can be worlds apart — one adequately powered, one not — because the outcomes behave differently.
Say you're comparing a program group to a control group of 15 people each, on an outcome with a standard deviation of 10 points. That design is powered to reliably detect a difference of roughly 12 points or more between groups — a large effect. If you expect the program to move the outcome by 3 or 4 points, which is plausible for a lot of real interventions, that sample can't detect it. Running the test anyway doesn't make the question unanswerable; it just means a non-significant result tells you almost nothing about whether the program worked, only that this sample was never going to be able to tell you either way.
When a small sample is still worth analyzing
- It's a pilot, and you say so. A pilot study analyzed and reported as a feasibility check — can we recruit, does the measure behave sensibly, what's the rough variability we should plan around — is a legitimate use of a small sample. The mistake isn't running it. It's reporting the effect estimate as if it were a confirmed finding rather than a planning input for a properly powered follow-up.
- The effect you're looking for is genuinely large. A small sample can detect a big, obvious effect just fine. The problem is specifically with modest, realistic effect sizes, which is most of them.
- You're describing, not inferring. Reporting what happened in a small group — means, ranges, individual cases — without a significance test attached to a causal claim is honest and often useful. The trouble starts when a descriptive number gets a p-value bolted onto it and treated as proof.
- You're generating a hypothesis for a larger study, not testing one. A small sample can tell you which of several outcomes looks promising enough to build a properly powered study around. It just can't tell you, on its own, that the promising one is real.
A quick gut-check before you run the test
- What's the smallest effect that would actually matter? Not the effect you hope for — the smallest one that would change a decision. That number, not your headcount, is the real input to a power calculation.
- Do you know the outcome's variability, even roughly? From prior literature, a pilot, or a comparable dataset. Without it, a power calculation is a guess dressed up as math.
- Is the design confirmatory or exploratory? If it's exploratory, say so in how you report it, and don't let a single p-value stand in for a conclusion.
- Would a different model use the data more efficiently? Designs that pool information — multilevel models, Bayesian approaches with informative priors, matched or paired comparisons — often get more out of a small sample than a standard test that treats every observation as independent and uninformed by anything else.
If you're not sure whether the sample you have (or the one you're planning to collect) can actually answer the question you're asking, run the power analysis before you collect the data, not after — this is exactly where we help.