This benchmark evaluates a model's ability to verify factual accuracy in long-form medical texts by recursively decomposing claims into a verification tree. It probes fine-grained fact-checking across six medical domains, requiring the model to distinguish between factual and deliberately falsified claims while accounting for contextual dependencies. Use when the user wants to benchmark on Med-Critics, or asks about evaluating this task. Reports accuracy.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill med-critics-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Med Critics Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-med-critics-eval)More formats (shields.io, HTML) on the badges page.
---
name: med-critics-eval
description: This benchmark evaluates a model's ability to verify factual accuracy in long-form medical texts by recursively decomposing claims into a verification tree. It probes fine-grained fact-checking across six medical domains, requiring the model to distinguish between factual and deliberately falsified claims while accounting for contextual dependencies. Use when the user wants to benchmark on Med-Critics, or asks about evaluating this task. Reports accuracy.
metadata:
skill_kind: dataset_eval
source_arxiv: 2501.10642
bibtex_key: huang2025iterative
confidence: high
---
# med-critics-eval
> Iterative Tree Analysis for Medical Critics — Huang et al. (2025) (arXiv:2501.10642, 2025)
## What this evaluates
This benchmark evaluates a model's ability to verify factual accuracy in long-form medical texts by recursively decomposing claims into a verification tree. It probes fine-grained fact-checking across six medical domains, requiring the model to distinguish between factual and deliberately falsified claims while accounting for contextual dependencies.
## Datasets
- **Med-Critics** — total 980; splits: test (980)
## Metrics
- `accuracy` **(primary)** — range: [0, 1]
- Measures the proportion of correctly verified claims or samples against the ground truth labels. Calculated as the number of correct predictions divided by the total number of evaluated claims.
- `F1@K` — range: [0, 1]
- Evaluates both precision and recall at the top-K retrieved or verified facts. Computed as the harmonic mean of precision and recall considering only the top-K predictions.
## Input / output format
**Input**: Long-form medical text passages (typically 20-30 sentences) containing atomic claims, with a subset deliberately falsified to simulate misinformation.
**Output**: Binary factuality labels (factual/non-factual or accept/reject) for each extracted atomic claim, organized hierarchically as a verification tree.
## Scoring recipe
```python
def compute_accuracy(predictions, gold_labels):
correct = sum(1 for p, g in zip(predictions, gold_labels) if p == g)
return correct / len(gold_labels)
def compute_f1_at_k(top_k_preds, top_k_golds, k):
tp = sum(1 for p, g in zip(top_k_preds[:k], top_k_golds[:k]) if p == g == 1)
fp = sum(1 for p, g in zip(top_k_preds[:k], top_k_golds[:k]) if p == 1 and g == 0)
fn = sum(1 for p, g in zip(top_k_preds[:k], top_k_golds[:k]) if p == 0 and g == 1)
prec = tp / (tp + fp) if (tp + fp) > 0 else 0
rec = tp / (tp + fn) if (tp + fn) > 0 else 0
return 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
```
## Common pitfalls
- Long-form factuality lacks a definitive set of facts, making ground-truth construction and claim extraction highly sensitive to LLM prompts and tokenization.
- Medical facts are highly context-dependent; evaluating claims in isolation often yields false positives, requiring subtree-level contextual grounding.
- The number and nature of extracted claims are unpredictable in generative models, complicating fair comparison and consistent scoring across methods.
## Evidence (verbatim from paper)
> To evaluate our method’s performance in assessing the factuality of long-form medical texts, we employ several key metrics. The accuracy measures the discrepancy between the ground truth and the factual verifications, providing a basic accuracy assessment. We also use the $F_{1}@K$ metric Wei et al. ([2024]), which evaluates both precision and recall.
## Citation
```bibtex
@misc{huang2025iterative,
title={Iterative Tree Analysis for Medical Critics},
author={Huang et al. (2025)},
year={2025},
note={arXiv:2501.10642}
}
```
- arXiv: 2501.10642
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!