A framework for rigorous evaluation of model-agnostic explainability methods: multi-metric statistical benchmarking, operational protocol, and reproducibility

Autores/as

DOI:

https://doi.org/10.69850/rimi.vi3.307

Palabras clave:

Explainable AI, Benchmark Methodology, Reproducibility, Statistical Evaluation, Model-Agnostic Explanations

Resumen

Evaluating explainability methods requires more than a single faithfulness proxy. We present a modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end. On the UCI Adult benchmark, we use a staged evidence protocol: a calibration/reproducibility stage (EXP1) followed by a primary com-parative/robustness benchmark (EXP2); the current merged recovery snapshot contains 299 committed result artifacts (99.7% artifact coverage) plus a 30-row SHAP recovery batch, yielding 275 analyzable unique runs out of 300 planned cells (91.7%). Across complete model-size blocks (5 models, N ∈ {50, 100, 200}), Friedman tests indicate significant method differences for fidelity (χ2 = 42.12, p = 3.78 × 109), stability (χ2 = 40.68, p = 7.65 × 109), sparsity (χ2 = 35.64, p = 8.92 × 108), faithfulness gap (χ2 = 45.00, p = 9.25 × 1010), and runtime (χ2 = 30.44, p = 1.12 × 106). SHAP leads on fidelity/stability, DiCE leads on sparsity, and LIME remains fastest overall. We release the framework, operation protocol, and artifacts with explicit data-quality caveats for reproducible benchmark use under a quantitative-only claim scope.

Descargas

Publicado

2026-09-01

Número

Sección

Téchne