Most teams take the sample size calculation seriously. They debate the assumptions, run the numbers several ways, and weigh the result against the budget.
Then the protocol is submitted, and the calculation feels settled. Enrollment opens. Everyone moves forward.
But that calculation was done while the protocol was being written, before this trial had produced a single data point. It often rests on earlier-phase results and published studies, which may come from somewhat different populations under somewhat different protocols.
Then the data start to come in, and some of those assumptions turn out to be wrong. More patients stop study treatment than expected. More withdraw from the study altogether. The primary endpoint is more variable than the published literature suggested. In a responder analysis, the response rate in the control arm is higher than the historical precedent you relied on. Any one of these can leave your trial underpowered — and you may not find out until the primary analysis.

⚠️ What goes into a sample size calculation — and where it goes wrong
A sample size calculation rests on several inputs: the expected treatment difference, the variability of the outcome measure, the Type I error rate, the desired power, and how many patients you expect to stop study treatment or withdraw from the study. In practice, there is one more: the budget.
The Type I error rate, the power, and the budget are choices. The rest are assumptions, not facts.
The expected treatment difference is usually the most consequential assumption and the most uncertain. It is typically drawn from earlier-phase data — a Phase 2 trial, a published study, sometimes both. But Phase 2 is noisy. Observed treatment differences frequently shrink between Phase 2 and Phase 3, a phenomenon sometimes called the winner's curse. If your Phase 3 trial is powered on an optimistic Phase 2 estimate, it may be underpowered before the first patient enrolls.
The budget brings its own pressure. The first sample size is often larger than the budget can support, so the assumptions get revisited. Sometimes that is the right call, because a first pass can be too conservative. But each revised assumption still has to stand on the data. A treatment difference that grows because the budget is tight is not a better estimate. It is a less powerful trial.
Variability is another common source of error. It is often estimated from prior studies in somewhat different populations, with somewhat different protocols. When variability is underestimated, the trial is underpowered even if the true treatment difference is exactly what you expected.
⚠️ Stopping treatment is not leaving the study
The assumptions about patients who don't complete the trial as planned are the ones most often wrong, and wrong in the most predictable direction. Sponsors consistently underestimate them, especially in long trials and in populations with multiple comorbidities.
They are also the assumptions most often lumped together into a single rate. But two different things can happen to a patient.
Discontinuing study treatment is not the same as withdrawing from the study. A patient who stops taking study drug can, and under a treatment-policy strategy should, continue to be followed. Their data still count.
Withdrawal from the study, whether through withdrawn consent or loss to follow-up, is different. Those patients' data are missing.
The protocol, the SAP, and the sample size assumptions should keep the two apart.
That matters for the sample size, because the two affect it differently. Withdrawal costs you data. Stopping treatment usually costs you treatment difference: patients who stop study drug are still followed, but they are unlikely to keep the full benefit, so the treatment difference you can expect to see shrinks. Most sample size sections handle both with a single inflation factor. I'll come back to why that falls short in a future issue.
🔍 What FDA reviewers look for
In the last issue, I wrote that an SAP should state every assumption behind the sample size, and where it came from. Here is why that matters so much to a reviewer.
When I reviewed an NDA or BLA, the sample size section told me a great deal about how carefully a team had thought through its trial design.
A rigorous sample size section documents every assumption explicitly — where it came from and why it was chosen. The strongest ones also include a table showing the sample size, or the power, under alternative assumptions. What happens to power if more patients stop treatment, or withdraw, than you assumed? What happens if the treatment difference is 20% smaller? Sponsors who included that table had clearly done the work.
What concerned me was a sample size section that stated assumptions without justification. A treatment difference of 3 points, with no reference to the data that supported it. A standard deviation of 8, with no explanation of where that number came from. An assumption that 10% of patients would not complete a 2-year trial in an elderly population, with no source.
These are not neutral choices. If they are wrong, the trial cannot answer its own question. When they appeared without justification, I looked harder at everything else.
⚠️ Interim analyses and formal re-estimation
Some trial designs allow a formal sample size re-estimation at a pre-specified interim look. Done properly, this is a legitimate tool. The key word is "properly."
A blinded sample size re-estimation — in which the pooled variance (for a continuous endpoint) or the overall response rate (for a binary endpoint) is examined without unblinding treatment assignments — is well accepted and generally has little or no effect on the Type I error rate.
An unblinded re-estimation, which adjusts the sample size based on the observed treatment difference, is more complex. It requires methods that control the Type I error rate, a firewall that keeps interim results away from the sponsor's trial team, and alignment with FDA in advance.
If you are considering either, it needs to be pre-specified in the protocol and SAP, with the rules clearly defined: at which interim look, what information will be examined, how the decision will be made, and how the final analysis will account for it. Deciding to re-estimate after the fact, or re-estimating without pre-specification, is a serious problem. It leads straight back to the reviewer's question from the last issue: what did the data look like when this decision was made?
💡 A working estimate, not a final answer
The sample size calculation is not a one-time event. It is your best estimate at a point in time, and it should be revisited as you learn about your population.
The sponsors who run into trouble finalize the calculation, file it, and never think about it again until the primary analysis falls short. At that point, the options are limited and expensive.
The sponsors who navigate this well build sensitivity into the design up front. They monitor blinded, pooled data — treatment discontinuation, study withdrawal, variability, the overall event rate — to see whether their assumptions are holding. And if their design allows a formal re-estimation, the rules are already in the protocol and SAP, and they have discussed them with FDA.
That conversation belongs before the trial starts, not after the data come in.
If you're second-guessing your power calculation before locking your protocol, I'm happy to talk through where you stand. Book a free 30-minute call.
Thank you for reading!
Lisa