This evaluation benchmarks 2D medical image segmentation models across three diverse clinical datasets (polyp, skin lesion, cardiac ultrasound). It probes the cross-domain transferability and segmentation accuracy of general-purpose vision models versus specialized medical architectures. Use when the user wants to benchmark on NeoPolyp, CAMUS, ISIC'18, or asks about evaluating this task. Reports mDSC.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill medical-segmentation-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Medical Segmentation Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-medical-segmentation-eval)More formats (shields.io, HTML) on the badges page.
---
name: medical-segmentation-eval
description: This evaluation benchmarks 2D medical image segmentation models across three diverse clinical datasets (polyp, skin lesion, cardiac ultrasound). It probes the cross-domain transferability and segmentation accuracy of general-purpose vision models versus specialized medical architectures. Use when the user wants to benchmark on NeoPolyp, CAMUS, ISIC'18, or asks about evaluating this task. Reports mDSC.
metadata:
skill_kind: dataset_eval
source_arxiv: 2603.13044
bibtex_key: borst2026generalpurpose
confidence: high
---
# medical-segmentation-eval
> Are General-Purpose Vision Models All We Need for 2D Medical Image Segmentation? A Cross-Dataset Empirical Study — Borst et al. (2026) (arXiv:2603.13044, 2026)
## What this evaluates
This evaluation benchmarks 2D medical image segmentation models across three diverse clinical datasets (polyp, skin lesion, cardiac ultrasound). It probes the cross-domain transferability and segmentation accuracy of general-purpose vision models versus specialized medical architectures.
## Datasets
- **NeoPolyp** — total ?; splits: 5-fold CV (-1)
- **CAMUS** — total ?; splits: 5-fold CV (-1)
- **ISIC'18** — total ?; splits: 5-fold CV (-1)
## Metrics
- `mDSC` **(primary)** — range: percent
- Mean Dice Similarity Coefficient. Calculated as the average Dice score across all segmentation classes, where Dice = 2 * |prediction ∩ ground_truth| / (|prediction| + |ground_truth|).
- `mIoU` — range: percent
- Mean Intersection over Union. Average IoU across all classes, where IoU = |prediction ∩ ground_truth| / |prediction ∪ ground_truth|.
- `mRec` — range: percent
- Mean Recall. Average recall across all segmentation classes.
- `mPrec` — range: percent
- Mean Precision. Average precision across all segmentation classes.
## Input / output format
**Input**: 2D medical images (RGB endoscopic/skin or grayscale ultrasound) with corresponding pixel-level segmentation masks.
**Output**: Pixel-wise segmentation masks matching input dimensions, with class labels per dataset annotation scheme.
## Scoring recipe
```python
def compute_mDSC(pred_mask, gold_mask, num_classes):
dice_scores = []
for c in range(num_classes):
pred_c = (pred_mask == c)
gold_c = (gold_mask == c)
intersection = np.sum(pred_c & gold_c)
union = np.sum(pred_c | gold_c)
dice = (2.0 * intersection) / (union + 1e-6)
dice_scores.append(dice)
return np.mean(dice_scores) * 100 # Paper reports as percentage
```
## Common pitfalls
- Performance varies significantly across classes; e.g., non-neoplastic polyps (C1) are consistently harder to segment than neoplastic ones.
- Evaluations use 5-fold cross-validation rather than a fixed held-out test set, so results should be interpreted as averaged fold metrics.
- Cross-dataset performance gaps are not uniform; GP-VMs show the largest advantage on NeoPolyp but only marginal gains on ISIC'18 and CAMUS.
## Evidence (verbatim from paper)
> Table 3 reports the 5-fold CV results, using mDSC as main performance metric. Measured by the average mDSC across all three datasets, the top-performing models are exclusively GP-VMs: VW-MiT (91.0%), VW-Conv and TransNeXt (both 90.9%), followed by InternImage (90.8%) as well as SegNeXt and SegFormer (both 90.7%).
## Citation
```bibtex
@misc{borst2026generalpurpose,
title={Are General-Purpose Vision Models All We Need for 2D Medical Image Segmentation? A Cross-Dataset Empirical Study},
author={Borst et al. (2026)},
year={2026},
note={arXiv:2603.13044}
}
```
- arXiv: 2603.13044
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!