Blog

Missing data: make the assumptions reviewable

Missing data: make the assumptions reviewable

The percentage of missing observations is informative, but it does not determine the risk of bias by itself. Timing, reasons, treatment differences, and the relationship between missingness and outcomes all matter. A small amount of consequential missing information can be more important than a larger amount of unrelated administrative loss.

The EMA guideline on missing data provides a regulatory reference. This article proposes a review workflow for organizing the clinical and statistical questions that arise before selecting an approach.

1. Describe the missingness process

Separate an assessment that was never scheduled from one that was scheduled but missed. Distinguish a missed visit, an unusable measurement, a withdrawn participant, and a record that has not yet arrived. These categories imply different operational actions and different potential implications for analysis.

For a hypothetical patient-reported endpoint, plot assessment availability over time by treatment, then examine reasons and participant characteristics. This is a descriptive diagnostic, not a formal proof of a missing-data mechanism. It helps identify where the analysis assumptions require closer scrutiny.

2. Tie assumptions to the treatment-effect question

A missing-at-random assumption is conditional on the information included in the analysis; it is not established simply because the model includes several covariates. Discuss whether clinically relevant predictors of both outcome and observation are available and measured adequately.

ICH E9(R1) is the relevant reference for the estimand. The question determines which outcomes matter, including how intercurrent events are addressed. Missing-data handling should be specified in that context rather than added after the primary model has been chosen.

3. Challenge assumptions deliberately

A sensitivity analysis should explore departures that clinicians can interpret. For example, a hypothetical pattern-mixture exercise might vary an assumption about unobserved outcomes after discontinuation. The range and direction of changes need clinical justification; arbitrary numeric shifts do not automatically provide meaningful assurance.

Identify which assumptions are being challenged and which remain unchanged. Document implementation details, random seeds where applicable, and uncertainty calculations. If a result changes materially under a plausible departure, that is information to explain, not an inconvenience to hide.

4. Prevent the problem where possible

Improve scheduling, follow-up, and data collection before relying on analytical remedies. Keep permitted post-discontinuation assessments operationally feasible. Track reasons for noncollection in a way that supports interpretation without imposing unnecessary burden on participants or sites.

ICH E8(R1) is a useful design reference. A practical review package combines availability plots, reason summaries, the governing estimand, primary assumptions, and prespecified sensitivity analyses. This gives clinical and statistical reviewers a shared basis for discussing uncertainty instead of treating missingness as a single percentage.

Sources and further reading