Evaluates pathology foundation models on fine-grained tissue classification of liver cancer patches and whole-slide tumor/non-tumor segmentation. It measures how well models can quantify tissue composition (e.g., fibrosis, necrosis, inflammation) within clinically defined regions to support reproducible digital pathology analysis. Use when the user wants to benchmark on HepatoBench, or asks about evaluating this task. Reports F1-score, Dice coefficient.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill hepatobench-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Hepatobench Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-hepatobench-eval)More formats (shields.io, HTML) on the badges page.
---
name: hepatobench-eval
description: Evaluates pathology foundation models on fine-grained tissue classification of liver cancer patches and whole-slide tumor/non-tumor segmentation. It measures how well models can quantify tissue composition (e.g., fibrosis, necrosis, inflammation) within clinically defined regions to support reproducible digital pathology analysis. Use when the user wants to benchmark on HepatoBench, or asks about evaluating this task. Reports F1-score, Dice coefficient.
metadata:
skill_kind: dataset_eval
source_arxiv: 2604.22858
bibtex_key: xiao2026hepatobench
confidence: high
---
# hepatobench-eval
> A Digital Pathology Resource for Liver Cancer Quantification with Datasets, Benchmarks, and Tools — Xiao et al. (2026) (arXiv:2604.22858, 2026)
## What this evaluates
Evaluates pathology foundation models on fine-grained tissue classification of liver cancer patches and whole-slide tumor/non-tumor segmentation. It measures how well models can quantify tissue composition (e.g., fibrosis, necrosis, inflammation) within clinically defined regions to support reproducible digital pathology analysis.
## Datasets
- **HepatoBench** — total ?; splits: train (-1), val (-1), test (-1)
## Metrics
- `F1-score` **(primary)** — range: [0, 1]
- Harmonic mean of precision and recall. Computed per class for the seven tissue categories and macro-averaged across classes for the final score.
- `Dice coefficient` **(primary)** — range: [0, 1]
- 2 * |A ∩ B| / (|A| + |B|), where A and B are the predicted and ground truth segmentation masks. Measures spatial overlap between predicted and annotated tumor/non-tumor regions.
- `Accuracy` — range: [0, 1]
- Proportion of correctly classified patches out of the total number of patches.
- `IoU` — range: [0, 1]
- Intersection over Union: |A ∩ B| / |A ∪ B|, measuring the overlap ratio between predicted and ground truth segmentation masks.
## Input / output format
**Input**: 150×150 pixel image patches at 20× magnification for classification; whole-slide images (WSIs) for segmentation.
**Output**: Per patch: one of seven tissue class labels (TUM, FIB, INF, NEC, NOR, REA, STE). Per WSI: a binary mask delineating tumor vs. non-tumor regions.
## Scoring recipe
```python
def compute_metrics(preds, gold):
acc = sum(p == g for p, g in zip(preds, gold)) / len(gold)
f1s = []
for cls in ['TUM','FIB','INF','NEC','NOR','REA','STE']:
tp = sum(1 for p,g in zip(preds,gold) if p==cls and g==cls)
fp = sum(1 for p,g in zip(preds,gold) if p==cls and g!=cls)
fn = sum(1 for p,g in zip(preds,gold) if p!=cls and g==cls)
prec = tp/(tp+fp) if (tp+fp)>0 else 0
rec = tp/(tp+fn) if (tp+fn)>0 else 0
f1s.append(2*prec*rec/(prec+rec) if (prec+rec)>0 else 0)
macro_f1 = sum(f1s)/len(f1s)
inter = np.sum(pred_mask & gold_mask)
union = np.sum(pred_mask | gold_mask)
iou = inter/union if union>0 else 0
dice = 2*inter/(np.sum(pred_mask)+np.sum(gold_mask)) if (np.sum(pred_mask)+np.sum(gold_mask))>0 else 0
return {'accuracy': acc, 'f1_macro': macro_f1, 'iou': iou, 'dice': dice}
```
## Common pitfalls
- Data splits are strictly performed at the slide/patient level to prevent information leakage from overlapping patches derived from the same WSI.
- WSIs used to derive patches for the classification dataset are explicitly excluded from the test set to avoid confounding the end-to-end pipeline evaluation.
- Inference uses overlapping sliding windows, requiring careful stitching to avoid boundary artifacts that could skew mask continuity and metric scores.
## Evidence (verbatim from paper)
> For the patch-level seven-class classification task, we evaluate model performance using Accuracy (Acc), Precision, Recall and F1-score. For whole-slide image (WSI) region segmentation, we report Intersection-over-Union (IoU) ,Dice coefficient (Dice), Precision, Recall and F1-score as the primary evaluation metrics.
## Citation
```bibtex
@misc{xiao2026hepatobench,
title={A Digital Pathology Resource for Liver Cancer Quantification with Datasets, Benchmarks, and Tools},
author={Xiao et al. (2026)},
year={2026},
note={arXiv:2604.22858}
}
```
- arXiv: 2604.22858
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!