Evaluates a model's ability to perform syndrome differentiation in Traditional Chinese Medicine by classifying clinical records into one of 148 predefined syndromes. It probes the model's capacity to handle domain-specific medical terminology and imbalanced multi-class classification. Use when the user wants to benchmark on TCM-SD, or asks about evaluating this task. Reports Macro-F1.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill tcmsd-sd-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Tcmsd Sd Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-tcmsd-sd-eval)More formats (shields.io, HTML) on the badges page.
---
name: tcmsd-sd-eval
description: Evaluates a model's ability to perform syndrome differentiation in Traditional Chinese Medicine by classifying clinical records into one of 148 predefined syndromes. It probes the model's capacity to handle domain-specific medical terminology and imbalanced multi-class classification. Use when the user wants to benchmark on TCM-SD, or asks about evaluating this task. Reports Macro-F1.
metadata:
skill_kind: dataset_eval
source_arxiv: 2203.10839
bibtex_key: ren2022tcmsd
confidence: high
---
# tcmsd-sd-eval
> TCM-SD: A Benchmark for Probing Syndrome Differentiation via Natural Language Processing — Mucheng Ren et al. (arXiv:2203.10839, 2022)
## What this evaluates
Evaluates a model's ability to perform syndrome differentiation in Traditional Chinese Medicine by classifying clinical records into one of 148 predefined syndromes. It probes the model's capacity to handle domain-specific medical terminology and imbalanced multi-class classification.
## Datasets
- **TCM-SD** — total 54152; splits: dev (-1), test (-1); repo https://github.com/Borororo/ZY-BERT
## Metrics
- `Macro-F1` **(primary)** — range: [0, 1]
- The unweighted mean of the F1-score calculated independently for each of the 148 syndrome classes. It treats all classes equally regardless of support.
- `Accuracy` — range: [0, 1]
- The proportion of correctly predicted syndromes out of the total number of instances.
## Input / output format
**Input**: Concatenated text: '[CLS] Chief Complaint [SEP] Medical History [SEP]'
**Output**: A single predicted syndrome label from a fixed set of 148 candidate labels, derived from the [CLS] token representation.
## Scoring recipe
```python
def compute_metrics(preds, golds, num_classes=148):
acc = sum(p == g for p, g in zip(preds, golds)) / len(golds)
f1_scores = []
for c in range(num_classes):
tp = sum(1 for p, g in zip(preds, golds) if p == c and g == c)
fp = sum(1 for p, g in zip(preds, golds) if p == c and g != c)
fn = sum(1 for p, g in zip(preds, golds) if p != c and g == c)
prec = tp / (tp + fp) if (tp + fp) > 0 else 0
rec = tp / (tp + fn) if (tp + fn) > 0 else 0
f1 = 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
f1_scores.append(f1)
macro_f1 = sum(f1_scores) / num_classes
return {'accuracy': acc, 'macro_f1': macro_f1}
```
## Common pitfalls
- The dataset is highly imbalanced, so Accuracy can be misleadingly high while Macro-F1 remains low.
- Syndromes are context-dependent and abstract, requiring domain-specific pre-training to map symptoms to structured labels effectively.
- Models must strictly follow the [CLS]...[SEP] token format for input concatenation to function correctly.
## Evidence (verbatim from paper)
> We selected the multi-class classification task as the primary form of SD to directly compare the performances of the existing models against the TCM-SD dataset, and used the accuracy and Macro-F1 as evaluation metrics. Specifically, the chief complaint and medical history were concatenated as the inputs, i.e. [CLS] Chief Complaint [SEP] Medical History [SEP], where [CLS] and [SEP] are special tokens used for classification and separation. Then the model predicts the target syndromes from 148 candidate labels based on the representation of [CLS] token. However, each syndrome in the TCM-SD dataset should have the same importance. Thus, the Macro-F1 is a more accurate metric to evaluate the performances of the models.
## Citation
```bibtex
@misc{ren2022tcmsd,
title={TCM-SD: A Benchmark for Probing Syndrome Differentiation via Natural Language Processing},
author={Mucheng Ren et al.},
year={2022},
note={arXiv:2203.10839}
}
```
- arXiv: 2203.10839
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!