P224: From PDF to Computable Protocol: Operationalizing ICH M11 for Automated EHR Extraction and Data Provenance
Poster Presenter
Ron Fitzmartin
Principal Consultant
Decision Analytics United States
Objectives
1. Explain how a machine-readable ICH M11 protocol can drive USDM-aligned data capture and protocol-to-data workflows.
2. Assess methods for extracting structured clinical trial data from EHR images with automated PII redaction and source provenance.
3. Evaluate how protocol-driven workflows can reduce manual transcription and SDV while producing SDTM-ready outputs.
Method
Objective
To evaluate a cloud-based workflow in which a machine-readable ICH M11 protocol generates USDM-aligned capture forms and supports extraction from EHR images with automated PII redaction, enabling near-zero, exception-based SDV and SDTM-ready outputs.
Method
In a secure cloud workspace, an ICH M11 protocol is rendered to XML/JSON and used to generate USDM-aligned extraction forms. Data are extracted from de-identified synthetic EHR screenshots generated in a hospital sandbox. Automated PII redaction is applied and each extracted value is linked to its redacted source image and audit record. Monitoring is performed using an exception-based approach and outputs are transformed to SDTM.
Results
Traditional EDC workflows rely on manual EHR-to-EDC transcription and extensive SDV. This study evaluates a protocol-driven workflow where a machine-readable ICH M11 protocol generates USDM-aligned capture screens and enables source-linked extraction from synthetic EHR screenshots. Automated PII redaction is applied, and each extracted value is anchored to a redacted source image and audit trail, producing SDTM-ready outputs and an auditable eSource archive.
Performance will be assessed using predefined sources and outputs. Extraction accuracy will be evaluated using a gold-standard sample (~250 fields) comparing extracted values to manual entries. PII redaction accuracy will be measured through manual review of ~175 screenshots. Provenance integrity will verify linkage between extracted values, redacted images, and audit records. Data-entry and SDV reduction will be estimated across ~2400 extracted fields and monitor records.
Benchmark targets include =95% extraction accuracy, =99% identifier masking, 100% provenance integrity, =80% reduction in manual data entry, and =90% reduction in SDV workload.