A framework for rigorous evaluation of model-agnostic explainability methods: multi-metric statistical benchmarking, operational protocol, and reproducibility
DOI:
https://doi.org/10.69850/rimi.vi3.307Palabras clave:
Explainable AI, Benchmark Methodology, Reproducibility, Statistical Evaluation, Model-Agnostic ExplanationsResumen
Evaluating explainability methods requires more than a single faithfulness proxy. We present a modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end. On the UCI Adult benchmark, we use a staged evidence protocol: a calibration/reproducibility stage (EXP1) followed by a primary com-parative/robustness benchmark (EXP2); the current merged recovery snapshot contains 299 committed result artifacts (99.7% artifact coverage) plus a 30-row SHAP recovery batch, yielding 275 analyzable unique runs out of 300 planned cells (91.7%). Across complete model-size blocks (5 models, N ∈ {50, 100, 200}), Friedman tests indicate significant method differences for fidelity (χ2 = 42.12, p = 3.78 × 10−9), stability (χ2 = 40.68, p = 7.65 × 10−9), sparsity (χ2 = 35.64, p = 8.92 × 10−8), faithfulness gap (χ2 = 45.00, p = 9.25 × 10−10), and runtime (χ2 = 30.44, p = 1.12 × 10−6). SHAP leads on fidelity/stability, DiCE leads on sparsity, and LIME remains fastest overall. We release the framework, operation protocol, and artifacts with explicit data-quality caveats for reproducible benchmark use under a quantitative-only claim scope.
Descargas
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Jonathan Herrera Vasquez; Miguel Herrero Uceda

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial-CompartirIgual 4.0.
