P320: Data Curation Best Practices and Innovations for Real World Evidence Generated Using Electronic Health Record Sourced Data
Poster Presenter
Patrick Rodriguez
Policy Analyst
Duke-Margolis Institute For Health Policy (DMI) United States
Objectives
To explore operational and practical considerations for EHR-sourced data curation; identify best practices and tools to address curation-specific regulatory considerations; examine the emerging role of artificial intelligence (AI) in supporting curation workflows; and develop policy recommendations.
Method
From February to October 2025, the Duke-Margolis RWE Collaborative engaged in monthly meetings, a workshop, peer-reviewed literature and regulatory guidance review, and expert presentations to characterize EHR-sourced data curation challenges and identify practical and policy solutions.
Results
The RWE Collaborative approached and considered data curation as an iterative, governance-intensive process that 1) occurs throughout the EHR-sourced data life cycle, and 2) must preserve source data’s clinical meaning while enabling interoperability and analytic use. Best practices were identified as implementing the following measures: predefined curation procedures, standardized terminologies, audit trails, addressing unaccounted bias, and provenance documentation to enhance regulatory utility. Ten policy recommendations were developed, with three (3) centered on defining EHR-sourced data curation, four (4) centered on real-world best practices and tools to address regulatory concerns, and four (4) centered on the role of AI in EHR-sourced data curation. For example, when defining EHR-sourced data curation, curators should establish and adhere to pre-defined SOPs for iterative quality checks across the data pipeline (i.e., from source to final analytic file) with transparent documentation of curation processes, quality testing, and justifications for data transformations. When considering real-world best practices and tools to address regulatory concerns, sponsors and data vendors should adopt a proactive transparency and regulatory readiness mindset—engaging early and often with the FDA to ensure data management practices align with regulatory expectations. Lastly, when considering the role of AI in EHR-sourced data curation, a shared framework for AI transparency and provenance should be established for EHR data curation, leveraging FHIR provenance to maintain an auditable record of both AI-driven and human reviewer actions. Any processing of free text notes by AI/natural language processing should be verified against originals to ensure accuracy and prevent hallucinations and AI should not autonomously make definitive EHR data changes.
Conclusion
Advancing regulatory considerations for EHR-sourced data curation requires correct conceptualization of data curation, effective real-world best practices and tools, and AI integration. Our findings and policy recommendations draw on engagement with the RWE Collaborative, academic literature, and regulatory guidance. Stakeholders involved in EHR-sourced data curation can consider our actionable recommendations to foster reliable data curation activities that align with key regulatory expectations. Aligning curation practices with regulatory expectations and thoughtfully integrating AI can strengthen the quality, relevance, and reliability of EHR-sourced data. These recommendations form a framework for advancing trustworthy, regulatory fit-for-use EHR-sourced data curation. By strengthening methods and transparency around EHR-sourced data curation, stakeholders can more clearly demonstrate and convey data relevance, reliability, and quality for a variety of purposes, including regulatory.