Already a DIA Member? Sign in. Not a member? Join.

Sign in

Forgot User ID? or Forgot Password?

Not a Member?

Create Account and Join

Menu Back to Poster-Presentations-Details

P113: XAI-RegSci Index: A Five-Axis Framework for Quantitative Evaluation of Explainable AI in Regulatory Science





Poster Presenter

      Yuma Tsuboi

      • PhD student
      • Waseda University
        Japan

Objectives

To develop and validate a five-axis quantitative framework (XAI-RegSci Index) that evaluates explainable AI across Faithfulness, Robustness, Justification, Traceability, and Reproducibility, addressing the lack of standardized XAI quality metrics for regulatory science decision-making.

Method

SHAP explanations for three ML models (XGBoost, Logistic Regression, Random Forest) were evaluated on four biomedical datasets (n=569–1,357) using five axis-specific metrics aligned with FDA/EU AI guidance. Validated via 5-fold CV with bootstrap CIs and Dirichlet Monte Carlo (n=10,000)

Results

On the Breast Cancer Wisconsin benchmark, the XAI-RegSci Index scores were: Logistic Regression 0.909, XGBoost 0.886, and Random Forest 0.884. Logistic Regression achieved the highest Faithfulness (0.840) and Reproducibility (0.989), while Random Forest scored highest in Traceability (0.878). All models demonstrated strong Justification (0.955–0.974) and Robustness (0.877–0.884). Five-fold cross-validation on BCW confirmed stability: Faithfulness showed the greatest variability (LR: 0.875±0.092, RF: 0.483±0.164), while Robustness (0.977–1.000), Justification (0.957–0.970), and Reproducibility (0.896–0.983) remained consistently high across folds. Generalizability was tested on three additional datasets. Heart Failure (n=299): overall scores ranged from 0.823 (XGBoost) to 0.859 (Random Forest). METABRIC genomic data (n=1,357, 500 genes): scores ranged from 0.663 (LR) to 0.691 (XGBoost). GDSC pharmacogenomics (n=962, 500 genes): scores ranged from 0.661 (LR) to 0.677 (RF). High-dimensional genomic datasets showed systematically lower scores, primarily driven by reduced Faithfulness in high-feature spaces. Sensitivity analysis using Dirichlet Monte Carlo simulation (n=10,000 random weight vectors) demonstrated that Logistic Regression ranked first in 79.4% of samples on BCW, with no ranking reversals observed across three predefined weight presets (equal, regulatory-focused, performance-focused). Construct validity testing confirmed the framework's discriminative power: real SHAP explanations scored significantly higher than shuffled or random baselines across all five axes, confirming that the index captures meaningful explanation quality rather than artifacts.

Conclusion

The XAI-RegSci Index provides the first integrated, quantitative framework for evaluating XAI quality specifically designed for regulatory science contexts. By operationalizing five complementary axes—Faithfulness, Robustness, Justification, Traceability, and Reproducibility—the framework addresses key requirements from emerging FDA and EU AI/ML regulatory guidance documents. Three key findings emerge. First, simpler models (Logistic Regression) can achieve higher overall explainability scores than complex models (XGBoost, Random Forest), primarily through superior Faithfulness and Reproducibility. This provides quantitative evidence for the interpretability-complexity tradeoff relevant to regulatory submissions. Second, Faithfulness is the most discriminating and variable axis, suggesting it should receive particular attention in regulatory evaluation of AI explanations. Third, the framework generalizes across clinical and genomic datasets, though high-dimensional settings present systematic challenges that the index transparently captures. The framework has practical implications for regulatory submissions involving AI/ML. It offers sponsors a standardized vocabulary and scoring system for demonstrating explanation quality, and provides reviewers at agencies such as FDA, EMA, and PMDA with quantitative benchmarks for assessing XAI trustworthiness. The five-axis decomposition allows targeted identification of specific explanation weaknesses rather than relying on a single aggregate metric. Limitations include the current focus on tabular data and SHAP-based explanations. Future work will extend the framework to image and text modalities, additional explanation methods, and prospective validation in regulatory review settings.

Be informed and stay engaged.

Don't miss an opportunity - join our mailing list to stay up to date on DIA insights and events.