P234: Comparative Validation of the Article Screener Agent of an Agentic AI Literature Review Platform Against Manual Review
Poster Presenter
Julien Heidt
Associate Director, Applied AI Science
IQVIA United States
Objectives
• To compare the performance of the full-text screener agent of an agentic AI literature review platform against manual review by an epidemiologist as the gold standard.
• To quantify time saved by using the agentic AI literature review platform versus manual screening.
Method
A targeted literature review (TLR) was conducted to characterize the definition and assessment methods for infant developmental delays in pregnancy studies. The title/abstract and full text of the identified studies were independently reviewed against pre-specified eligibility criteria by AI and the
Results
A total of 730 studies were identified from PubMed for title/abstract review. The following validation metrics were calculated upon comparison of the AI decisions against the manual review: true positive (TP): 29, true negative (TN): 664, false positive (FP): 29, false negative: 8, positive predictive value (PPV, A.K.A precision): 50%, negative predictive value (NPV): 99%, sensitivity: 78%, specificity: 96%, F1: 61%, and accuracy: 95%. Reasons for misclassification by AI included: not recognizing some outcomes as developmental delay (e.g. weight or height measurement), misclassifying non-pharmaceutical interventions (e.g. Covid or environmental exposure) as qualifying exposures even though only vaccines, pharmaceuticals and alcohol were permitted, and at times misclassifying prospective design as retrospective due to patient enrollment occurring in the past. The manual review of 730 title/abstracts took 10 hours while AI processing time was less than 5 minutes.
A total of 37 articles qualified for full-text review by the epidemiologist. Full-text could not be obtained for 6 articles; therefore 31 articles proceeded to independent full-text review by the AI platform and the epidemiologist. The following metrics were calculated based on full-text screening: TP: 26, TN: 3, FP: 1, FN:1, PPV: 96%, NPV: 75%, sensitivity: 96%, specificity: 75% (1 misclassification out of the 3 TN), F1 score 96%, and accuracy 94%. Misclassifications by AI included: inclusion of a study with minimal details on developmental delay and misclassifying retrospective design as prospective. The manual review of the 31 full-text articles took 7 hours while the AI processing time was less than 5 minutes.
Conclusion
The accelerating volume of medical publications necessitates scalable and systematic approaches to literature review, for which AI-based tools may provide support. Validation of AI-assisted literature reviews is crucial in ensuring methodological rigor. The agentic AI platform demonstrated high PPV, NPV, sensitivity, specificity, F1 score and accuracy, especially for full-text screening, while offering 99% time savings in the screening stage of the literature review. These performance metrics were obtained despite the challenging nature of the research question whereby the outcome component of the PICO framework was itself the research question and hence was not defined in detail in the eligibility criteria to avoid imposing contemporary or investigator-driven assumptions on the review. The lower PPV observed during title/abstract screening reflects the platform’s intentionally conservative design, which prioritizes recall by retaining studies when eligibility is uncertain to avoid missing relevant evidence. Given the substantial time savings achieved, the capture of additional articles at this stage would not meaningfully increase reviewer workload. AI screening can support TLR workflows by rapidly screening articles at the title/abstract and full-text review stages of a literature review to help researchers keep pace with emerging research. The results of the study are particularly applicable to drug research where timely and reproducible evidence synthesis is essential.