Evaluates LLMs on Chinese public security domain tasks including text classification, information extraction, question answering, and text generation. Probes domain-specific accuracy, reliability, and contextual understanding in high-stakes law enforcement scenarios. Use when the user wants to benchmark on Weibo Sentiment Analysis, Rumor Detection, Telecommunication Fraud Detection, Drug-Related Case Reports, Public Security Case Reading Comprehension, Public Security Case Summary, or asks ab...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill cpsdbench-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Cpsdbench Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-cpsdbench-eval)More formats (shields.io, HTML) on the badges page.
---
name: cpsdbench-eval
description: Evaluates LLMs on Chinese public security domain tasks including text classification, information extraction, question answering, and text generation. Probes domain-specific accuracy, reliability, and contextual understanding in high-stakes law enforcement scenarios. Use when the user wants to benchmark on Weibo Sentiment Analysis, Rumor Detection, Telecommunication Fraud Detection, Drug-Related Case Reports, Public Security Case Reading Comprehension, Public Security Case Summary, or asks about evaluating this task. Reports F1-Score.
metadata:
skill_kind: dataset_eval
source_arxiv: 2402.07234
bibtex_key: tong2024cpsdbench
confidence: high
---
# cpsdbench-eval
> CPSDBench: A Large Language Model Evaluation Benchmark and Baseline for Chinese Public Security Domain — Tong et al. (2024) (arXiv:2402.07234, 2024)
## What this evaluates
Evaluates LLMs on Chinese public security domain tasks including text classification, information extraction, question answering, and text generation. Probes domain-specific accuracy, reliability, and contextual understanding in high-stakes law enforcement scenarios.
## Datasets
- **Weibo Sentiment Analysis** — total 600; splits: test (-1)
- **Rumor Detection** — total 400; splits: test (-1)
- **Telecommunication Fraud Detection** — total 400; splits: test (-1)
- **Drug-Related Case Reports** — total ?; splits: test (-1)
- **Public Security Case Reading Comprehension** — total ?; splits: test (-1)
- **Public Security Case Summary** — total ?; splits: test (-1)
## Metrics
- `Accuracy` — range: [0, 1]
- Calculated as (TP+TN)/(TP+TN+FP+FN). Represents the proportion of correctly classified instances out of the total.
- `Precision` — range: [0, 1]
- Calculated as TP/(TP+FP). Measures the proportion of positive predictions that are actually correct.
- `Recall` — range: [0, 1]
- Calculated as TP/(TP+FN). Measures the proportion of actual positives that are correctly identified.
- `F1-Score` **(primary)** — range: [0, 1]
- Harmonic mean of Precision and Recall: 2 * (Precision * Recall) / (Precision + Recall). Standard for classification and IE tasks.
- `Hybrid IE Metric` — range: [0, 1]
- A two-stage score combining exact match and fuzzy match. Exact match is 1 if pred==gold else 0. Fuzzy match uses Levenshtein distance; if distance ≤ threshold, score is 1. Final score is a weighted sum of exact and fuzzy scores. High-precision entities (names, amounts) use exact match only.
## Input / output format
**Input**: Chinese text instances from public security scenarios (e.g., social media posts, case reports, fraud messages, legal questions), wrapped in task-specific prompts containing role definition, task description, input specifications, and operational constraints.
**Output**: Task-dependent outputs: discrete class labels for classification; structured entity/relation tuples for information extraction; natural language yes/no or open-ended answers for question answering; and coherent summary paragraphs for text generation.
## Scoring recipe
```python
def compute_metrics(preds, golds, task_type):
if task_type == 'classification':
acc = sum(p == g for p, g in zip(preds, golds)) / len(golds)
return acc
elif task_type == 'ie':
exact_scores = []
fuzzy_scores = []
for p, g in zip(preds, golds):
exact = 1.0 if p == g else 0.0
exact_scores.append(exact)
if exact == 0:
lev = levenshtein_distance(p, g)
fuzzy_scores.append(1.0 if lev <= THRESHOLD else 0.0)
else:
fuzzy_scores.append(0.0)
return 0.5 * sum(exact_scores) + 0.5 * sum(fuzzy_scores)
return None
```
## Common pitfalls
- Exact match vs. semantic equivalence: LLM outputs may differ literally but be semantically correct (e.g., '9:00 am' vs 'around 9:00 am'). The hybrid metric accounts for this via Levenshtein distance.
- High-precision entities (names, amounts) in fraud/IE tasks strictly use exact match only, ignoring fuzzy matches, which can penalize minor phrasing variations.
- Dataset sizes are small (dozens to hundreds) due to commercial API costs, limiting statistical power and generalizability of results.
## Evidence (verbatim from paper)
> For text classification tasks, we have chosen Accuracy, Precision, Recall, and F1-Score as evaluation metrics. ... For information extraction tasks, the typical metrics are Precision, Recall, and F1-Score. ... we have designed a hybrid evaluation metric. It comprises two steps: firstly, calculating the exact match score between the LLM’s output and the label. Secondly, for predictions that are not exact matches, we calculate the Levenshtein distance... Finally, we obtain a comprehensive score by weighting these two types of scores.
## Citation
```bibtex
@misc{tong2024cpsdbench,
title={CPSDBench: A Large Language Model Evaluation Benchmark and Baseline for Chinese Public Security Domain},
author={Tong et al. (2024)},
year={2024},
note={arXiv:2402.07234}
}
```
- arXiv: 2402.07234
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!