Evaluates the quality of feature importance explanations for deep learning models in healthcare mortality prediction. It measures how well an interpretability method identifies truly predictive features by observing model performance degradation when those features are removed. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUC of performance curve.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill mimic-if-interpretability-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Mimic If Interpretability Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-mimic-if-interpretability-eval)More formats (shields.io, HTML) on the badges page.
---
name: mimic-if-interpretability-eval
description: Evaluates the quality of feature importance explanations for deep learning models in healthcare mortality prediction. It measures how well an interpretability method identifies truly predictive features by observing model performance degradation when those features are removed. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUC of performance curve.
metadata:
skill_kind: dataset_eval
source_arxiv: 2102.06761
bibtex_key: meng2021mimicif
confidence: high
---
# mimic-if-interpretability-eval
> MIMIC-IF: Interpretability and Fairness Evaluation of Deep Learning Models on MIMIC-IV Dataset — Meng et al. (2021) (arXiv:2102.06761, 2021)
## What this evaluates
Evaluates the quality of feature importance explanations for deep learning models in healthcare mortality prediction. It measures how well an interpretability method identifies truly predictive features by observing model performance degradation when those features are removed.
## Datasets
- **MIMIC-IV** — total ?; splits: train (-1), val (-1), test (-1)
## Metrics
- `AUC of performance curve` **(primary)** — range: [0, 1]
- Area under the curve (AUC) of the performance-vs-feature-drop-ratio plot. Computed via trapezoidal rule over ratios 0.1 to 1.0. Lower values indicate faster performance degradation and better feature importance ranking.
## Input / output format
**Input**: Patient feature sequence X in R^{TxF} or summary vector x in R^F, flattened to R^{d_in} for evaluation.
**Output**: Binary probability of in-hospital mortality, plus a non-negative feature importance score vector s(x) in R^{d_in} for each interpretability method.
## Scoring recipe
```python
def compute_roar_auc(interpreter, X_train, y_train, X_test, y_test):
scores = interpreter.get_importance(X_train, X_test)
curve = []
for ratio in [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0]:
X_train_m = replace_top_features(X_train, scores, ratio, fill=np.mean(X_train, axis=0))
X_test_m = replace_top_features(X_test, scores, ratio, fill=np.mean(X_train, axis=0))
model = retrain(X_train_m, y_train)
preds = model.predict(X_test_m)
perf = auc_pr(preds, y_test) # or auc_roc
curve.append(perf)
return np.trapz(curve, x=[0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0])
```
## Common pitfalls
- Ground-truth feature importance is unavailable for healthcare tasks, so ROAR performance degradation is used as a proxy rather than direct correlation with true importance.
- Gradient-based methods are inapplicable to models with categorical embeddings (e.g., AutoInt), requiring alternative methods like ShapleySampling or ArchDetect.
- Lower AUC on the performance curve indicates better interpretability, which is counter-intuitive compared to standard accuracy metrics.
## Evidence (verbatim from paper)
> Since acquiring the ground-truth feature importance is challenging for mortality prediction tasks, we evaluate one feature importance estimation by gradually dropping most important features it gives at certain ratios from the dataset and observe the degradation of the model’s performance. The larger the degradation is, the better the estimation is, since it identifies the features most helpful for the model on the task. More specifically, we use ROAR (remove and retrain) proposed in[[7]] for evaluation. For each interpretability method, we replace the most important features of certain fractions of each data sample with a fixed uninformative value. We conduct this in both training and test sets. Then we retrain the model with the modified training set and evaluate its classification performance on the modified test set. By retraining the model on datasets with features removed, ROAR ensures that train and test data comes from a similar distribution and reduces the impact on the model’s performance of data distribution discrepancy, so that the degradation of performance is caused by the removal of information instead of the shift of data distribution. We evaluate each interpretabil
## Citation
```bibtex
@misc{meng2021mimicif,
title={MIMIC-IF: Interpretability and Fairness Evaluation of Deep Learning Models on MIMIC-IV Dataset},
author={Meng et al. (2021)},
year={2021},
note={arXiv:2102.06761}
}
```
- arXiv: 2102.06761
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!