Installs into .claude/skills of the current project.
Are you the author of Arxiv 2609 15855v1 K Bench A Clinically Calibrated Benchmark For Eval?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-arxiv-2609-15855v1-k-bench-a-clinically-calibrated)
--
name: arxiv-2609-15855v1-k-bench-a-clinically-calibrated-benchmark-for-eval
description: 'K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations (arXiv: 2609.15855v1)'
metadata:
{
"arxiv_id": "2609.15855v1",
"utility": 1.0,
"title": "K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations",
"authors": "Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova",
"url": "http://arxiv.org/abs/2609.15855v1"
}
--
# K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
**arXiv ID:** 2609.15855v1
**Authors:** Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova
**URL:** http://arxiv.org/abs/2609.15855v1
**Utility Score:** 1.00
## Summary
This skill was automatically generated from the arXiv paper titled "K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations" (ID: 2609.15855v1).
## Usage
This skill can be used to reference the paper's concepts, methodologies, or findings in agent workflows.
## References
- arXiv: http://arxiv.org/abs/2609.15855v1