Two programs can agree and still answer the wrong question. Independent implementation is valuable, but agreement is only one form of evidence. A reliable quality-control strategy also examines specifications, clinical meaning, data provenance, and the consequences of an error.
ICH E8(R1) emphasizes attention to factors critical to study quality. We apply that principle here as a practical way to prioritize biometrics review, rather than treating every discrepancy as equally consequential.
1. Distinguish shared assumptions from independent checks
If both programmers use the same incorrect population rule, their outputs may match exactly. Independence therefore needs more than separate source files. Reviewers should challenge the governing rule and inspect selected participant records against the approved plan.
For a hypothetical efficacy endpoint, check the definition of baseline, the assessment window, and the handling of rescue treatment before comparing estimates. Agreement after an incorrect exclusion is not evidence that the clinical question was implemented correctly.
2. Use different evidence for different risks
A numerical reconciliation is appropriate for many calculations. Record-level review is useful for derivations and selection rules. Structural checks can identify duplicate keys or impossible relationships. Visual review is necessary for labels, footnotes, and information that disappears during formatting.
We recommend a control matrix that names the failure mode, its potential consequence, the evidence needed to detect it, and the responsible reviewer. It should explain why a check exists. Counting checks without understanding their coverage can create a misleading impression of assurance.
3. Design adversarial test cases
Include synthetic cases that stress boundaries: the last instant of a treatment-emergent window, a missing date component, two equally eligible assessments, or a participant who changes treatment. Expected outcomes should be written independently of the code under test.
CDISC ADaM helps frame the analysis data package, but study-specific statistical rules still need targeted tests. Do not assume that a standards validator can determine whether a baseline rule matches the SAP.
4. Close findings with evidence
A discrepancy is not closed merely because someone changed a program. Record the cause, affected outputs, correction, verification, and whether the issue indicates a broader pattern. A wrong denominator in one table may indicate a reusable population function that affects many displays.
ICH E6(R3) is relevant to the broader quality and oversight context. Our proposed operating practice is to keep review decisions attached to the artifact version they concern, with a clear owner and release consequence.
A balanced QC process asks two questions: do independent calculations agree, and is there sufficient evidence that the agreed result is the intended one? Keeping both questions visible makes review more informative and helps direct limited expert attention toward errors that could change a clinical interpretation.
A boundary-case review packet
For a hypothetical rule that includes events beginning from first dose through 30 days after last dose, prepare cases one day before, exactly on, and one day after each boundary. Add a partial onset date, a missing last-dose date, and an event that begins before dosing but worsens afterwards. The latter cases require approved study-specific rules; the test author should not invent them.
Write the expected classification and rationale before execution. Compare the output flag and contributing dates against that record. If the program passes ordinary cases but cannot explain an exception, the test has exposed either an implementation gap or an unresolved specification.
This packet tests consequences that an aggregate reconciliation can miss. It complements independent programming by challenging the meaning of the shared rule.