A large observational dataset can contain detailed information while still being unsuitable for a particular question. Clinical care, billing, and research collection serve different purposes. The circumstances that produced a record influence what it can support analytically.
The FDA’s July 2024 EHR and medical claims guidance is a primary reference for assessing these data sources. The workflow below is a proposed way to make provenance and fitness-for-purpose questions explicit before modeling.
1. Explain the data-generating process
Identify the healthcare settings, contributing organizations, capture mechanisms, coverage period, and known gaps. Distinguish a recorded diagnosis from a confirmed clinical endpoint and a prescription from evidence that a medicine was taken.
For a hypothetical treatment cohort, a gap in claims might mean loss of coverage rather than absence of care. The analysis should not silently interpret that gap as a period without events. Review eligibility and capture rules before defining follow-up.
2. Validate key concepts
Assess how exposure, outcomes, covariates, and time references are identified. Document algorithms, code lists, version choices, and supporting validation evidence. A concept identified accurately in one setting may perform differently in another.
Review whether important confounders are captured with sufficient quality and timing. Statistical adjustment cannot recover a variable that was never measured adequately. Describe residual uncertainty rather than treating a long covariate list as proof of comparability.
3. Preserve transformation history
Keep the original representation, harmonization rules, linkage decisions, and exclusion logic recoverable. Data linkage can add information while introducing mismatch or selection risks. Examine those risks in the context of the intended analysis.
The FDA’s draft external-control guidance is additional reading when observational data are used for an external comparator. It remains draft guidance, not a universal acceptance standard for such designs.
4. Connect quality assessment to the question
A dataset may be fit for describing utilization and less fit for estimating a causal treatment effect. Specify the target population, treatment strategies, outcome, follow-up, and analytical assumptions before deciding which quality checks are sufficient.
ICH E9(R1) provides useful treatment-effect vocabulary, although applying causal methods to observational data requires additional design and identification work. Do not imply that adopting trial terminology removes confounding.
A reviewable real-world evidence package explains why the data are appropriate, how variables were constructed, what limitations remain, and how those limitations affect interpretation. Provenance is part of the scientific argument, not an administrative appendix added after the estimate is produced.