Evaluates an LLM-based agent's ability to automate cohort selection, feature extraction, and clinical code standardization across heterogeneous Electronic Medical Record (EMR) databases. It probes schema-aware SQL reasoning, iterative query planning, and robustness to unseen database structures without manual rule engineering. Use when the user wants to benchmark on MIMIC-III, eICU, SICdb, or asks about evaluating this task. Reports F1.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill emr-agent-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Emr Agent Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-emr-agent-eval)More formats (shields.io, HTML) on the badges page.
---
name: emr-agent-eval
description: Evaluates an LLM-based agent's ability to automate cohort selection, feature extraction, and clinical code standardization across heterogeneous Electronic Medical Record (EMR) databases. It probes schema-aware SQL reasoning, iterative query planning, and robustness to unseen database structures without manual rule engineering. Use when the user wants to benchmark on MIMIC-III, eICU, SICdb, or asks about evaluating this task. Reports F1.
metadata:
skill_kind: dataset_eval
source_arxiv: 2510.00549
bibtex_key: lee2025emragent
confidence: high
---
# emr-agent-eval
> EMR-AGENT: Automating Cohort and Feature Extraction from EMR Databases — Lee et al. (2025) (arXiv:2510.00549, 2025)
## What this evaluates
Evaluates an LLM-based agent's ability to automate cohort selection, feature extraction, and clinical code standardization across heterogeneous Electronic Medical Record (EMR) databases. It probes schema-aware SQL reasoning, iterative query planning, and robustness to unseen database structures without manual rule engineering.
## Datasets
- **MIMIC-III** — total ?; splits: test (-1)
- **eICU** — total ?; splits: test (-1)
- **SICdb** — total ?; splits: test (-1)
## Metrics
- `F1` **(primary)** — range: [0, 1]
- Harmonic mean of precision and recall: 2 * (precision * recall) / (precision + recall). Used as the headline metric for both cohort/feature selection and code mapping tasks.
- `Accuracy` — range: [0, 1]
- Proportion of correctly predicted cohort/feature selections out of total instances.
- `Balanced Accuracy` — range: [0, 1]
- Average of recall obtained on each class, specifically used for the code mapping task to handle class imbalance.
## Input / output format
**Input**: Natural language requests for cohort/feature extraction or code mapping, combined with database schema metadata (10 sample values per column), external clinical documents/knowledge, and evaluation memos.
**Output**: SQL queries (for cohort/feature selection) and standardized clinical codes (for code mapping).
## Scoring recipe
```python
def compute_metrics(predictions, gold):
tp = sum(1 for p, g in zip(predictions, gold) if p == g and p == 1)
fp = sum(1 for p, g in zip(predictions, gold) if p == 1 and g == 0)
fn = sum(1 for p, g in zip(predictions, gold) if p == 0 and g == 1)
precision = tp / (tp + fp) if (tp + fp) > 0 else 0
recall = tp / (tp + fn) if (tp + fn) > 0 else 0
f1 = 2 * precision * recall / (precision + recall) if (precision + recall) > 0 else 0
accuracy = sum(1 for p, g in zip(predictions, gold) if p == g) / len(gold)
return f1, accuracy
# For code mapping, balanced accuracy is computed per class and averaged.
```
## Common pitfalls
- Baselines are evaluated with naive prompt adaptations to PostgreSQL, which may not reflect optimal prompt engineering for each method.
- The evaluation relies on strict adherence to PhysioNet Data Use Agreements, limiting public reproducibility of the exact database instances.
- Temperature scheduling and retry limits are hardcoded per component, meaning results are not directly comparable across different LLM backbones without re-tuning.
## Evidence (verbatim from paper)
> On MIMIC-III, CFSA achieves an F1 of 0.94, surpassing single-prompt baselines (e.g., ICL-PLUQ, 0.749 F1) as well as more complex pipelines. Even under more complex and unseen schemas such as eICU and SICdb (Section[4.1]), where baseline F1 scores fall below 0.53 and 0.51, respectively, CFSA maintains high performance (0.93 and 0.81), demonstrating strong generalizability.
## Citation
```bibtex
@misc{lee2025emragent,
title={EMR-AGENT: Automating Cohort and Feature Extraction from EMR Databases},
author={Lee et al. (2025)},
year={2025},
note={arXiv:2510.00549}
}
```
- arXiv: 2510.00549
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!