In every NDA and BLA I reviewed, the statistical analyses were measured against one document, written before anyone saw the data: the Statistical Analysis Plan.

The SAP is a binding commitment to FDA about exactly what you are going to do with your data — which analyses, which populations, which methods, in which order — before you see the results. Pre-specification exists because it is the primary safeguard against the human tendency to make the data tell a more favorable story than it actually tells. When I read your SAP, I was reading your promise. And then I was going to find out whether you kept it.

Most sponsors know this. They are encouraged to finalize the SAP early, ideally when the protocol is finalized. In practice, many SAPs are not finalized until shortly before the blind is broken. ICH E9 permits that. But a late SAP invites the question every reviewer asks: had anyone seen the data when this was written? Even when the answer is no, the sponsor now has to show it.

Infographic: What reviewers see when they read your SAP. Signals credibility: primary endpoint precisely defined; assumptions stated and sourced; missing data approach fully specified; testing hierarchy pre-specified. Raises concerns: 'appropriate methods will be applied'; '90% power to detect a treatment difference' with no effect size, variability, or sources; 'multiple imputation will be used' with no further detail; multiple endpoints with no alpha adjustment.

What a rigorous SAP looks like

A strong SAP is specific to the point of being almost tedious. It names the primary endpoint and defines it precisely. It specifies the analysis population — and explains why that population, not another, is the appropriate one. It describes how missing data will be handled, not in a general sense ("multiple imputation will be used") but in enough detail that a statistician who had never spoken with your team could reproduce the analysis exactly.

It defines the estimand. Since ICH E9(R1), a rigorous SAP explicitly states the estimand for each objective — the precise scientific question the trial is designed to answer. The estimand specifies the population, the treatment, the endpoint, the summary measure, and critically, how intercurrent events are handled. Intercurrent events are the things that happen during a trial that complicate interpretation: patients discontinue treatment, use rescue medication, withdraw consent, or die before the endpoint is assessed. Why a patient withdrew consent — lack of perceived benefit, an adverse event, worsening disease — can matter as much as the fact that they withdrew. Every one of these needs a pre-specified strategy. An SAP that doesn't address intercurrent events explicitly has left a gap the reviewer will have to fill. That gap becomes a question.

It addresses multiplicity. If you are testing a primary endpoint and secondary endpoints, the SAP tells the reviewer the exact testing hierarchy — in what order, with what alpha allocation, using what method. It doesn't leave this to be determined later.

It states the assumptions behind the sample size calculation, analysis models, missing data methods, and simulations — and where those assumptions came from. This is where many SAPs fall short. A sample size built on an effect size and variability estimate that appear nowhere in the document. A mixed model with no stated covariance structure. A multiple imputation approach with no statement of what it assumes about why data are missing. Simulated operating characteristics with no description of the scenarios behind them. Every statistical method rests on assumptions. If the SAP doesn't state them, the reviewer can't judge whether they were reasonable — or whether your conclusions depend on them. The FDA reviewer must be able to reproduce your analyses, and that starts with knowing what you assumed.

It pre-specifies sensitivity analyses. These are not afterthoughts. They belong in the SAP, specified in the same detail as the primary analysis.

And critically — it matches the protocol. Every analysis in the SAP should be traceable to the trial design. If the SAP introduces an analysis that wasn't contemplated in the protocol, that needs to be explained. If the SAP was amended, the amendments are dated, documented, and submitted with clear rationale.

What a weak SAP signals to a reviewer

A vague SAP is not neutral. It raises a specific concern: that the sponsor is leaving room to make decisions after the data are in hand.

I have read SAPs that described the primary analysis as "appropriate statistical methods will be applied." That is not a pre-specified analysis. That is an announcement that you haven't decided yet — or that you are reserving the right to decide after you see the data.

I have read SAPs that omitted any mention of multiplicity adjustments in submissions that tested three secondary endpoints. I have read SAPs whose planned analyses did not address the stated hypothesis or research question. The objective asked one thing, and the analysis answered another. The estimand framework is helping close that gap, because it requires the question to be defined before the analysis is chosen.

In each case, the gap became a problem. Either the submission required extensive back-and-forth to clarify, or the resulting analyses were treated with less confidence than the sponsor had hoped.

A weak SAP creates work. For the sponsor, for the FDA reviewer, and for the program. It generates review questions that didn't need to exist. It can delay approval or contribute to a CRL.

The deviation problem

Even a strong SAP creates problems when the submission doesn't follow it.

Over the course of a long development program, the SAP gets finalized, the trial runs, the data come in — and somewhere between the SAP and the final statistical report, something shifts. An analysis is reported that wasn't in the SAP. A sensitivity analysis that was pre-specified doesn't appear. The handling of missing data looks different from what was written.

Each of these deviations has to be identified, explained, and justified. Some are acceptable — protocol amendments happen, unforeseen situations arise. But unexplained deviations from the SAP are one of the most reliable predictors of a difficult review. When I encountered them, my next question was always the same: when was this decision made, and what did the data look like at that point?

Keep a dated record of when the SAP was finalized and when the blind was broken. ICH E9 asks for exactly that, and it is the first thing that answers the reviewer's question.

Get FDA's eyes on it early

Most SAP problems — undefined terms, methods that don't fit the design or the question being asked — are far easier to fix before the first patient is enrolled than after the data are in.

The most effective step a sponsor can take is to submit the protocol and SAP to FDA for review before the study starts. At that stage, a review concern costs a revision. Raised during the NDA or BLA review, the same concern can cost months, or contribute to a CRL.

Once the SAP is finalized, treat it as fixed. Every deviation needs a written explanation that anticipates the hardest version of the reviewer's question, because that is the version the reviewer will ask.

If you're finalizing a protocol and SAP, or preparing to explain deviations in a submission, I'm happy to talk through where you stand. Book a free 30-minute call.

Thank you for reading!
Lisa