Sample Size and Statistical Power: A Bench Scientist's Guide
How to calculate sample size and statistical power before you run an experiment, so you stop wasting reagents on studies that were never built to detect the effect you care about.
Every bench scientist has run the experiment that ended in a shrug. The treatment looked a little better than control, the error bars overlapped, the p-value landed at 0.18, and you were left staring at the plate wondering whether the effect was real and your study just couldn't see it, or whether there was nothing there at all. That ambiguity is rarely bad luck. More often it is the predictable result of starting an experiment without asking the one question that determines whether the data can answer your question: how many samples does this experiment actually need?
Sample size calculation and statistical power are not statistician bureaucracy you bolt on for a grant or a regulatory filing. They are design decisions that determine, before a single pipette tip is used, whether your experiment is capable of detecting the effect you care about. This guide walks through the logic the way a scientist at the bench actually needs it: the levers you control, how to estimate the inputs from pilot data and the literature, the replicate trap that quietly wrecks otherwise careful work, and a worked example you can adapt to your own comparison.
Why underpowered studies waste time and money
An underpowered study is one that does not have enough samples to reliably detect a real effect of a given size. The cruelty of it is that you usually cannot tell from the inside. You run eight wells per group, the difference is in the right direction, the statistics say "not significant," and you conclude the treatment doesn't work. But a non-significant result from an underpowered study is uninformative, not negative. You have not shown the effect is absent. You have shown your experiment was too small to find it.
The costs compound. You burn reagents, instrument time, and weeks of a postdoc's life on a result that cannot be interpreted either way. Worse, underpowered experiments that do reach significance tend to overestimate the effect, because only the larger, noisier swings clear the bar. Chase that inflated estimate into a follow-up study and it evaporates. Multiply this across a lab and you have the everyday version of the reproducibility problem: not fraud, just a steady stream of experiments that were never built to detect what they were testing for.
Power analysis flips the order of operations. Instead of running the experiment and hoping, you decide up front what effect would be scientifically meaningful, estimate how noisy your system is, and calculate the sample size that gives you a fair chance of detecting that effect if it exists. It is the difference between designing an experiment and merely performing one.
The four levers: effect size, alpha, power, and variability
Sample size is not a number you pick. It falls out of four quantities that are mathematically locked together. Fix any three and the fourth is determined. Understanding these four levers is most of what statistical power actually is.
Once you internalize that these four are interdependent, the workflow becomes obvious. You set alpha and power by convention, you estimate effect size and variability from your knowledge of the system, and you solve for n. The hard part is never the arithmetic, which any power-analysis tool will do. The hard part is supplying honest estimates of the effect you expect and the noise you live with.
- Effect size - how big a difference you want to be able to detect. A larger effect is easier to see, so it needs fewer samples. This is a scientific judgment, not a statistical one: what difference would actually change a decision?
- Significance level (alpha) - your tolerance for a false positive, conventionally 0.05. Lowering it (say to 0.01) makes the test more conservative and demands more samples.
- Power (1 − beta) - your probability of detecting the effect if it is genuinely there. Conventionally set to 0.80. Higher power (0.90, 0.95) requires more samples.
- Variability - how much your measurements scatter, captured as standard deviation. Noisier assays need more samples to see through the noise. This is the lever you can actually shrink through better technique.
What "power = 0.80" actually means
Power is the probability that your experiment will return a statistically significant result when the effect you are looking for is real. Setting power to 0.80 means that if the true effect is at least as large as the one you designed for, you have an 80 percent chance of detecting it, and a 20 percent chance of missing it. That 20 percent is your false-negative rate, beta. You are explicitly accepting that one time in five, a real effect of that size will slip past you.
Eighty percent is a convention, not a law of nature, and it is worth deciding deliberately rather than by habit. If missing a real effect is expensive, a failed candidate quietly dropped, a safety signal overlooked, you may want 0.90 or 0.95, and you should budget the extra samples that requires. If the experiment is an early, cheap screen and a missed hit will resurface elsewhere, you might accept less. The point is that power is a knob you are choosing, and choosing 0.80 by default is fine as long as you know that is what you are doing.
A useful reframing: alpha protects you from being fooled by noise into seeing something that isn't there, and power protects you from being blind to something that is. A well-designed experiment takes both seriously. Controlling alpha at 0.05 while running at 50 percent power is like locking the front door and leaving the back wide open.
Estimating effect size and variance before you have the data
The chicken-and-egg objection comes up immediately: how can I estimate the effect size and variability of an experiment I haven't run yet? You are not guessing blindly. You almost always have three sources to draw on, and a defensible power analysis simply makes those sources explicit.
Pilot data is the strongest source. A small preliminary run, even a handful of replicates, gives you a direct read on the standard deviation of your assay under your hands, in your lab, with your reagents. You should be cautious using a pilot to estimate the effect size itself (small pilots produce wildly unstable effect estimates), but a pilot is genuinely valuable for pinning down variability. Prior literature is the second source: published studies in a comparable system give ballpark means and standard deviations, though assays and cell lines differ enough that you should treat these as approximate. The third source is the smallest effect that matters, which is a scientific decision you can often make from first principles. If a 10 percent improvement in yield wouldn't change what you do next, don't power the study to detect 10 percent; power it to detect the difference that would actually move a decision.
A practical habit: when you express effect size in standardized form (the difference in means divided by the standard deviation, often called Cohen's d), you can reason about it without committing to exact units. A standardized effect around 0.5 is moderate; around 0.8 is large. If your literature scan and pilot suggest the effect you care about sits near the boundary of "moderate," you immediately know you are in the territory that needs a respectable sample size, not three wells and a hope.
Biological vs technical replicates, and the trap that inflates n
This is where careful experiments most often go wrong, so it deserves to be stated plainly: technical replicates do not increase your sample size. Biological replicates do. Confusing the two is one of the most common ways an experiment that looks well-replicated is actually underpowered.
A biological replicate is an independent biological unit: a separate animal, a distinct cell-culture passage, an independently prepared protein batch, a different formulation lot. These capture the real, unavoidable biological variation your conclusion needs to generalize over. A technical replicate is the same biological sample measured more than once: three wells loaded from one lysate, duplicate injections of the same vial, repeat reads of the same plate. Technical replicates tell you how precise your measurement is. They say nothing about whether the next animal, the next passage, or the next batch will behave the same way.
Your n for a power analysis is the number of independent biological replicates, not the total number of measurements. If you run three animals and measure each in triplicate, your n is three, not nine. The triplicates make each animal's value more precise, which is useful, but they cannot substitute for biological units. Averaging technical replicates into one value per biological sample before analysis is usually the honest move; treating all nine readings as independent data points is pseudoreplication, and it will make your study look far more powered than it is.
- Biological replicate - an independent biological unit (animal, passage, batch, lot). This is what counts toward n.
- Technical replicate - the same sample measured multiple times. Improves measurement precision, does not increase n.
- Rule of thumb - average technical replicates to one value per biological sample, then count biological samples as your n.
A worked example: treatment vs control with a t-test
Here is a concrete, deliberately hypothetical example you can adapt. Suppose you are comparing a treated group against a control on a continuous readout, say expression of a target protein measured by a validated assay, and you will analyze the two groups with a two-sample t-test. The illustrative numbers below are made up to show the mechanics, not pulled from any real study.
From a small pilot, your control values average around 100 units with a standard deviation of about 20 units. You decide, on scientific grounds, that the smallest difference worth detecting is 20 units, a 20 percent shift, because anything smaller wouldn't change your next step. That gives a standardized effect size of 20 divided by 20, or d = 1.0, a large effect. You fix alpha at 0.05 (two-sided) and power at 0.80.
Plug those into a standard two-sample power calculation and you get roughly 17 biological replicates per group to detect a d of 1.0 at 80 percent power. Now watch how sensitive that is to your inputs. If the meaningful difference is only 10 units instead of 20 (d = 0.5), the required n per group jumps to around 64, because halving the effect size you want to detect roughly quadruples the samples. If your assay is noisier than the pilot suggested and the standard deviation is 30 rather than 20, n climbs again. This sensitivity is the whole lesson: the difference between a feasible experiment and an impossible one often lives in assumptions you made in thirty seconds, which is exactly why those assumptions deserve to be written down and defended.
The reverse calculation is just as useful at the bench. If your animal protocol or your budget caps you at 10 per group, run the math backward: with n = 10, alpha 0.05, and power 0.80, what effect size can you actually detect? If the answer is a difference far larger than anything biologically plausible, the honest conclusion is that the experiment as scoped cannot answer the question, and you should redesign it before you run it, not after.
More conditions, more comparisons: the adjustment problem
The clean two-group example rarely survives contact with a real project. You add a second dose, a vehicle control, a positive control, three time points, and suddenly you are not making one comparison but a dozen. Every comparison you run at alpha 0.05 carries its own 5 percent chance of a false positive, and those chances accumulate. Run 20 independent comparisons with no real effects and you expect about one to come back "significant" purely by chance. This is the multiple-comparisons problem, and it is not optional accounting.
Correcting for it, whether by a Bonferroni adjustment, a false-discovery-rate procedure, or an appropriate ANOVA structure with planned contrasts, tightens the effective threshold each individual test must clear. A tighter threshold means you need more samples per group to retain the same power. So adding conditions has a double cost: more groups to run, and a stiffer bar for each comparison. The practical implication is that you must plan your analysis, including which comparisons you will actually make and how you will adjust for them, before you size the study. Sample size and analysis plan are the same decision, not two.
The most expensive experiments in any lab aren't the ones that fail. They're the ones that can't tell you whether they failed.
Common mistakes to design out from the start
A few failure modes show up again and again, and all of them are cheaper to prevent than to discover three months in. Catching them is mostly a matter of asking the right question before the experiment rather than after.
The underlying pattern is that almost every one of these mistakes is a planning failure, not an execution failure. The pipetting was fine; the design was set up to be uninterpretable. Power analysis is the discipline that forces these questions to the front, where they are still cheap to fix.
- Counting technical replicates as n - the single most common cause of overstated power and pseudoreplication.
- Running a study with no power calculation at all, then reading a non-significant result as evidence of no effect.
- Powering for an effect size you wish were true rather than the smallest one that would actually matter.
- Estimating variability from an overly optimistic pilot, or from a published study run under cleaner conditions than yours.
- Adding conditions and comparisons without adjusting alpha or re-sizing the study to preserve power.
- Deciding the analysis after seeing the data, which quietly turns exploratory peeking into a multiple-comparisons minefield.
Letting the design tool carry the statistics
None of this is conceptually hard, but doing it correctly every time, for every comparison, under deadline, is where good intentions break down. That is precisely the gap Shadow AI is built to close. When you describe your research question in plain language, Shadow AI produces a bench-ready experiment design that treats sample size and statistical power as first-class parts of the plan, not an afterthought. It proposes a power level and significance threshold, reasons explicitly about biological versus technical replicates so your n reflects independent biological units, and recommends a sample size tied to the effect size you actually care about.
Crucially, it keeps the analysis plan and the sample size consistent with each other. If your design carries multiple conditions and comparisons, Shadow AI flags the multiple-comparisons issue and folds the correction into both the statistical plan and the recommended n, so you are not surprised at analysis time. You stay in control of the scientific judgments, what effect matters, how much risk you'll accept, while the tool handles the bookkeeping that is easy to get wrong by hand.
If you have ever finished an experiment unsure whether a flat result meant "no effect" or "not enough samples," that uncertainty was designed in long before the bench work began. Turning your next research question into a fully powered, analysis-ready design, with the sample size justified before you start, is exactly what Shadow AI is for. Bring the question; let the design carry the statistics.