Already a DIA Member? Sign in. Not a member? Join.

Sign in

Forgot User ID? or Forgot Password?

Not a Member?

Create Account and Join

Menu Back to Poster-Presentations-Details

P237: An Agentic AI Fidelity Evaluation Framework for Real-World Evidence Workflows





Poster Presenter

      John Diaz-Decaro

      • Founder
      • Black Swan Causal Labs
        United States

Objectives

To propose a structured framework for evaluating end-to-end fidelity of agentic AI systems in RWE generation, decomposing performance across protocol reconstruction, code execution, and numerical concordance.

Method

A seven-subagent pipeline was architected across two phases (Protocol & Code Fidelity (S1–S4) and Execution & Reporting (S5–S7)) coordinated by a central orchestrator agent applying iterative quality gates with automated retry logic.

Results

Agentic AI systems capable of executing multi-step RWE workflows represent a fundamentally new evaluation challenge: standard benchmarks assess isolated tasks, not end-to-end reproducibility of pharmacoepidemiological analyses. This framework addresses that gap through three fidelity dimensions applied sequentially. Protocol Fidelity (S1-S2): A Protocol Reconstructor operationalizes published study designs using the HARPER template (extracting estimand definitions, inclusion/exclusion criteria, covariate specifications, and survey design parameters). A Scoring Gate evaluates completeness on a 0-100 scale; scores below 100% trigger structured retry cycles with itemized deficiency notes. Execution Fidelity (S3-S4): A Code Generator produces executable Python or R implementing the approved protocol. An Inspector Gate evaluates generated code against a methodological checklist. Failed checks trigger revision cycles. Numerical Concordance (S5-S6): A Concordance Evaluator compares AI-generated estimates against published results using three thresholds: point estimate relative difference <=10% (PASS) or <=20% (WARN); CI Jaccard overlap >=0.80 (PASS) or >=0.60 (WARN); and directional concordance. Deviations are classified into four categories: DATA_HANDLING, MODEL_SPEC, ESTIMAND_DRIFT, and NUMERICAL. The orchestrator logs all gate decisions and transitions, producing a structured Fidelity Report with a composite score. The orchestrator presents gate results to a human reviewer at each fidelity check, with stage transitions requiring explicit approval prior to proceeding. In an initial proof-of-concept using a published NHANES-based observational study (Jacobs et al., 2023), protocol reconstruction identified two critical gaps, an omitted covariate and an unspecified subgroup regression model consistent with the MODEL_SPEC and ESTIMAND_DRIFT failure categories. Full execution and concordance evaluation will be presented at the meeting.

Conclusion

As agentic AI systems (and by extension, generative AI tools) are incorporated and used in RWE workflows and processes, there is a need to carefully define quality checks to ensure traceability, reproducibility and transparency. Thus, evaluating whether agentic AI systems can faithfully reproduce RWE studies requires a framework, one that mirrors the structure of the analytic workflow itself. By incorporating fidelity checks into protocol reconstruction, code execution, and numerical concordance, and embedding iterative quality gates with structured retry logic, this framework provides an auditable, reproducible approach to AI performance evaluation in pharmacoepidemiology. The failure taxonomy (DATA_HANDLING, MODEL_SPEC, ESTIMAND_DRIFT, NUMERICAL) offers a principled vocabulary for characterizing where and why AI systems deviate from published results, which enables targeted improvement rather than binary pass/fail results. The orchestrator agent's decision log provides a complete audit trail consistent with emerging regulatory expectations for AI transparency in evidence generation. This framework contributes to ongoing discussions within professional societies (eg ISPE, ISPOR) and regulatory agencies on standards for AI-assisted RWE. As the field moves toward agentic workflows for pharmacoepidemiological research, validated evaluation infrastructure is a prerequisite for responsible deployment. Pilot results will be presented at the meeting.

Be informed and stay engaged.

Don't miss an opportunity - join our mailing list to stay up to date on DIA insights and events.