P342: Machine Learning–Derived RWE for Scientifically Appropriate and Statistical Sound Promotional Materials: A Case Study
Poster Presenter
Rose Chang
Vice President
Analysis Group, Inc. United States
Objectives
To demonstrate ML tools, such as latent class analysis (LCA), applied to RWD can generate evidence meeting FDA’s scientifically appropriate and statistically sound standard for promotional materials by identifying meaningful patient profiles and translating them into online disease education tools.
Method
Using Optum® de-identified Electronic Health Record data set (2017–2022), LCA identified profiles among 376,004 female outpatients treated for uncomplicated urinary tract infections (uUTI). Model fit statistics selected the optimal cluster number and treatment failure was compared across clusters.
Results
LCA identified unobserved (latent) subgroups based on patterns across demographic, clinical, and healthcare utilization variables, enabling multivariable segmentation beyond traditional stratification.
Four distinct clinically interpretable patient profiles were identified. The largest cluster (46.5%) comprised relatively healthy patients with non-recurrent UTI and minimal prior antibiotic exposure. A second cluster (22.7%) represented patients with recurrent UTI and high prior antibiotic use. A third cluster (14.0%) included access-disadvantaged patients characterized by higher Medicaid enrollment and substantial emergency department utilization. The fourth cluster (16.8%) consisted of older patients with elevated comorbidity burden, including diabetes and obesity.
Compared with the healthy reference cluster, all other clusters had significantly higher odds of treatment failure following empiric oral antibiotic treatment (odds ratios 1.40–1.54; all p<0.001). Elevated treatment failure risk was observed not only among patient clusters with elevated traditional clinical risk factors (e.g., recurrent infection, prior antibiotic exposure, high comorbidity burden), but also among the patient cluster defined primarily by social and healthcare access characteristics. These findings highlight meaningful heterogeneity in risk profiles, healthcare access, and comorbidity burden within the uUTI population that is not apparent using traditional stratification approaches.
These statistically derived clusters were translated into the “Faces of uUTI” framework. Each cluster was represented as a hypothetical patient archetype, assigned a name and narrative profile reflecting its defining characteristics (e.g., age, recurrent infection history, comorbidity burden). This human-centered translation enables clinicians to intuitively recognize and contextualize heterogeneous risk patterns observed in practice that could contribute to antibiotic treatment failures.
Conclusion
This uUTI case study illustrates how AI and ML methods applied to large-scale real-world data can generate actionable patient segmentation and link distinct profiles to treatment failure outcomes.
Importantly, the value extended beyond analytical differentiation and outcome modeling. Translating multivariable clustering outputs into intuitive, named patient archetypes enabled complex analytic findings to be communicated in a manner accessible to clinicians, patients, and other stakeholders. By bringing data-derived profiles to life through narrative-based representations, the “Faces of uUTI” framework facilitated clearer understanding of heterogeneous risk patterns, including those driven by social and healthcare access factors that may otherwise be under-recognized.
Broadly, this approach demonstrates how real-world data powered by AI and machine learning can be operationalized for drug company website educational materials that meet the FDA’s scientifically appropriate and statistically sound standard – not only to quantify heterogeneity, but to meaningfully communicate it, thereby improving stakeholder engagement and effectively communicating disease complexity across therapeutic areas.