Probes a clinical language model's ability to predict binary adverse outcomes from free-text EHR notes while simultaneously calibrating its prediction uncertainty. It evaluates both discriminative accuracy and probabilistic calibration across multiple clinical risk stratification tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill cura-mimic-iv-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Cura Mimic Iv Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-cura-mimic-iv-eval)More formats (shields.io, HTML) on the badges page.
---
name: cura-mimic-iv-eval
description: Probes a clinical language model's ability to predict binary adverse outcomes from free-text EHR notes while simultaneously calibrating its prediction uncertainty. It evaluates both discriminative accuracy and probabilistic calibration across multiple clinical risk stratification tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.
metadata:
skill_kind: dataset_eval
source_arxiv: 2604.14651
bibtex_key: wang2026cura
confidence: high
---
# cura-mimic-iv-eval
> CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Prediction — Wang et al. (arXiv:2604.14651, 2026)
## What this evaluates
Probes a clinical language model's ability to predict binary adverse outcomes from free-text EHR notes while simultaneously calibrating its prediction uncertainty. It evaluates both discriminative accuracy and probabilistic calibration across multiple clinical risk stratification tasks.
## Datasets
- **MIMIC-IV** — total ?; splits: 5-fold cross-validation (-1)
## Metrics
- `AUROC` **(primary)** — range: [0, 1]
- Area under the Receiver Operating Characteristic curve, measuring the trade-off between true positive rate and false positive rate across all classification thresholds.
- `AUPRC` — range: [0, 1]
- Area under the Precision-Recall Curve, summarizing the trade-off between precision and recall across thresholds, particularly sensitive to class imbalance.
- `Brier score` — range: [0, 1]
- Mean squared difference between predicted probabilities and actual binary outcomes: $\frac{1}{N}\sum_{i=1}^N (p_i - y_i)^2$. Lower is better.
- `NLL` — range: other
- Negative Log-Likelihood: $-\frac{1}{N}\sum_{i=1}^N [y_i \log(p_i) + (1-y_i) \log(1-p_i)]$. Measures probabilistic calibration and sharpness.
- `AURC` — range: [0, 1]
- Area Under the Risk-Coverage curve, evaluating how well uncertainty estimates correlate with prediction errors across varying coverage levels.
## Input / output format
**Input**: Single free-text clinical note from electronic health records.
**Output**: Binary label $y \in \{0, 1\}$ indicating a pre-defined adverse outcome within a fixed horizon, along with a predicted probability/uncertainty score.
## Scoring recipe
```python
import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
def compute_metrics(y_true, y_prob):
auroc = roc_auc_score(y_true, y_prob)
auprc = average_precision_score(y_true, y_prob)
brier = np.mean((y_prob - y_true) ** 2)
nll = -np.mean(y_true * np.log(y_prob + 1e-8) + (1 - y_true) * np.log(1 - y_prob + 1e-8))
sorted_idx = np.argsort(y_prob)
coverage = np.arange(1, len(y_true) + 1) / len(y_true)
errors = 1 - y_true[sorted_idx]
aurc = np.trapz(errors, coverage)
return {'AUROC': auroc, 'AUPRC': auprc, 'Brier': brier, 'NLL': nll, 'AURC': aurc}
```
## Common pitfalls
- Results are averaged over 5 cross-validation folds, not a single held-out test split; failing to report fold-wise variance obscures result stability.
- Calibration metrics (Brier, NLL, AURC) require well-calibrated predicted probabilities, not hard class predictions or logits.
- AURC evaluates uncertainty alignment across varying coverage levels, not at a fixed decision threshold like 0.5.
## Evidence (verbatim from paper)
> We evaluate model performance using two categories of metrics: (1) Discrimination: Area Under the ROC Curve (AUROC) and Area Under the Precision-Recall Curve (AUPRC); and (2) Calibration \& Uncertainty: Brier score, Negative Log-Likelihood (NLL), and Area Under the Risk-Coverage curve (AURC). All reported results represent the mean and standard deviation across five cross-validation folds.
## Citation
```bibtex
@misc{wang2026cura,
title={CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Prediction},
author={Wang et al.},
year={2026},
note={arXiv:2604.14651}
}
```
- arXiv: 2604.14651
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!