
Claude Skills by qhjqhj00
github.com/qhjqhj00Probes whether black-box behavioral distillation preserves safety alignment in medical LLMs. It measures functional fidelity on benign medical prompts and quantifies safety violations and refusal failures on adversarial inputs using an automated moderation classifier. Use when the user wants to benchmark on Medical QA Datasets (MedQA, PubMedQA, MedMCQA, EMRQA), Handcrafted Red-Teaming Suite, GQ-Generated Harmful Prompts, or asks about evaluating this task. Reports Violation Rate.
Evaluates deep learning models for medical image segmentation by quantifying region overlap and boundary alignment between predicted masks and expert-annotated ground truth. Use when the user wants to benchmark on LA (Left Atrium), Pancreas CT, BraTS 2019, ACDC, or asks about evaluating this task. Reports DSC.
This evaluation protocol assesses the segmentation accuracy and robustness of 3D medical imaging models under varying data quality and training strategies. It specifically probes how aleatoric uncertainty quantification can guide data filtering and dynamic loss weighting to improve performance across diverse anatomical structures and imaging modalities. Use when the user wants to benchmark on LiTS, TotalSegmentator, WORD, FeTA 2022, KiTS23, or asks about evaluating this task. Reports Dice score.
This evaluation benchmarks 2D medical image segmentation models across three diverse clinical datasets (polyp, skin lesion, cardiac ultrasound). It probes the cross-domain transferability and segmentation accuracy of general-purpose vision models versus specialized medical architectures. Use when the user wants to benchmark on NeoPolyp, CAMUS, ISIC'18, or asks about evaluating this task. Reports mDSC.
This protocol evaluates medical image segmentation models by measuring how accurately they predict anatomical structures across multiple modalities and organs. It specifically probes the ability of adaptive skip connections and differentiable search strategies to maintain or improve segmentation quality under different training regimes and backbone architectures. Use when the user wants to benchmark on ACDC, AMOS, BraTS, KiTS, or asks about evaluating this task. Reports macro-average soft Dice.
This evaluation probes the ability of large language models to generate accurate and faithful summaries of medical texts under high out-of-vocabulary (OOV) conditions. It specifically measures how tokenization fragmentation and domain-specific terminology affect summarization quality and concept preservation across multiple medical benchmarks. Use when the user wants to benchmark on PubMedQA, EBM, BioASQ-M, BioASQ-S, or asks about evaluating this task. Reports Rouge-L.
Evaluates a vision-language model's ability to localize tumors in medical images and generate structured clinical reports. Probes spatial accuracy of coordinate prediction and the clinical utility/quality of automated radiology reports. Use when the user wants to benchmark on Unspecified medical imaging dataset, or asks about evaluating this task. Reports positional_deviation.
Evaluates the medical visual question answering capabilities of multimodal large language models across diverse imaging modalities and general medical knowledge domains. Use when the user wants to benchmark on VQA-RAD, SLAKE (English CLOSED), PathVQA, PMC-VQA, MMMU (Health & Medicine track), OmniMedVQA (open access), or asks about evaluating this task. Reports accuracy.
Evaluates whether multimodal medical vision-language models actually rely on image content to answer questions, or if they exploit text-only shortcuts. It measures visual grounding by comparing model performance and prediction stability across real, blank, and shuffled image conditions. Use when the user wants to benchmark on PathVQA, PMC-VQA, SLAKE, VQA-RAD, or asks about evaluating this task. Reports VRS (Visual Reliance Score), IS (Image Sensitivity).
Evaluates transformer-based models on identifying medication mentions, classifying medication-related events, and determining contextual attributes (e.g., negation, temporality, certainty) from clinical narratives. Use when the user wants to benchmark on Challenge test dataset, or asks about evaluating this task. Reports micro-averaged F1-score.
Probes the visual reasoning reliability and robustness of multimodal medical foundation models by presenting pairs of visually distinct but semantically confused medical images. It measures whether models can correctly answer questions about each image individually and consistently across the pair, revealing shortcut learning and hallucination tendencies. Use when the user wants to benchmark on MediConfusion, or asks about evaluating this task. Reports Set accuracy.
Evaluates multi-turn, multi-modal dialogue capabilities for radiology patient education, testing how well models personalize explanations based on hidden patient profiles and ground visual annotations in medical images. It probes the alignment between textual explanations and drawn/image-marked evidence, as well as safety and scope adherence in medical contexts. Use when the user wants to benchmark on MedImageEdu, or asks about evaluating this task. Reports MedImageEdu Overall.
Predicts driver visual attention (fixation locations) in critical driving scenarios by modeling goal-directed attention sequences. It evaluates how well a model's predicted saliency map matches human gaze patterns across spatial and spatiotemporal cues. Use when the user wants to benchmark on DR(eye)VE, BDD-A, DADA-2000, EyeCar, or asks about evaluating this task. Reports CC.
Evaluates the ability of medical vision-language models to align expert clinical terminology with patient-accessible layman language while preserving diagnostic accuracy. It measures lexical overlap, readability, clinical factuality, and zero-shot image-text retrieval performance. Use when the user wants to benchmark on MedLayBench-V, or asks about evaluating this task. Reports Recall@K (R@1, R@5, R@10).
Evaluates a model's ability to answer medical visual questions across diverse imaging modalities (CT, MRI, X-ray, etc.) and generalizes to out-of-domain benchmarks. It probes the model's capacity for latent visual reasoning and robust cross-modality transfer without relying on external tools or retrieval augmentation. Use when the user wants to benchmark on OmniMedVQA, SLAKE, VQA-RAD, PMC-VQA, MMMU (Health & Medicine), MedXpertQA, or asks about evaluating this task. Reports accuracy.
Evaluates biomedical multimodal foundation models across visual question answering, image captioning, image generation, visual chat, and interleaved text-image generation. It probes the model's ability to understand medical images, reason over clinical reports, and generate clinically grounded multimodal responses. Use when the user wants to benchmark on VQA-RAD, SLAKE, PathVQA, QuiltVQA, PMC-VQA, PathMMU, ProbMed, OmniMedVQA, PMC-OA, MIMIC-CXR, Quilt-1M, LLaVA-Med, MedMax-Instruct, or asks a...
Evaluates a model's ability to answer medical multiple-choice questions, testing both domain-specific knowledge retrieval and deep medical reasoning capabilities. Use when the user wants to benchmark on MedMCQA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the robustness of deep learning image classifiers against realistic, domain-specific corruptions in medical imaging. It measures how well models maintain performance when tested on corrupted versions of standard medical datasets compared to clean data. Use when the user wants to benchmark on MedMNIST-C, or asks about evaluating this task. Reports BE, rBE.
Evaluates the generalization capability and scaling efficiency of self-supervised vision foundation models on a diverse suite of 12 biomedical image classification tasks. It probes how model capacity, data diversity, and pretraining objectives affect downstream diagnostic performance when using a frozen feature extractor. Use when the user wants to benchmark on MedMNIST (12 benchmarks), or asks about evaluating this task. Reports MCC.
Evaluates the ability of machine learning models to classify biomedical images across diverse modalities, tasks, and scales. It probes generalization capabilities by testing on standardized 2D and 3D images resized to 28×28 or 28×28×28, covering binary, multi-class, multi-label, and ordinal regression tasks. Use when the user wants to benchmark on PathMNIST, ChestMNIST, DermaMNIST, OCTMNIST, PneumoniaMNIST, RetinaMNIST, BreastMNIST, BloodMNIST, TissueMNIST, OrganAMNIST, OrganCMNIST, OrganSMNI...
Evaluates 3D medical image segmentation backbones on diverse anatomical structures across CT and MR modalities. Probes the model's ability to learn robust spatial representations and generalize to fine-tuning on small and large-scale clinical datasets. Use when the user wants to benchmark on Pediatric CT-Seg, Stanford Knee MR, Toothfairy, Stanford Brain Mets, PANTHER Pancreatic Tumor, CTSpine1k, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
This benchmark evaluates the ability of process reward models and general critic models to detect factual, logical, and clinical errors at individual reasoning steps in medical question-answering. It probes step-level correctness verification and case-level chain validation, emphasizing the identification of clinically critical mistakes such as missing contraindications, flawed diagnostic logic, or premature conclusions. Use when the user wants to benchmark on MedPRMBench, or asks about evalu...
Probes multimodal large language models' ability to assess medical image quality through low-level visual attribute detection and no-reference or comparative reasoning. It evaluates how well models identify image degradations, describe clinical attributes, and compare quality across different imaging modalities. Use when the user wants to benchmark on MedQ-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models' robustness and metacognitive reliability when processing medical images with various quality degradations (e.g., blur, noise, motion, artifacts) across different clinical capability dimensions. Use when the user wants to benchmark on MedQ-Deg, or asks about evaluating this task. Reports accuracy.
Evaluates medical multiple-choice question answering capability of 4B-parameter LLMs, specifically comparing the impact of domain fine-tuning versus retrieval-augmented generation (RAG) on accuracy. Use when the user wants to benchmark on MedQA-USMLE, or asks about evaluating this task. Reports Majority-vote accuracy.
Evaluates a medical AI system's clinical reasoning behavior, focusing on uncertainty handling, deferral, and safety rather than raw answer accuracy. It probes the model's ability to maintain clinician-aligned reasoning, avoid speculative completions, and preserve context across diverse medical queries. Use when the user wants to benchmark on MedQuAD benchmark, or asks about evaluating this task. Reports Benchmark Completion Rate.
Evaluates large language models' clinical reasoning capabilities across three stages: examination recommendation, diagnostic decision-making, and treatment planning. It assesses both the accuracy of final medical outputs and the quality of the underlying reasoning steps using factuality, completeness, and efficiency metrics. Use when the user wants to benchmark on MedR-Bench, or asks about evaluating this task. Reports Accuracy.
Probes multimodal large language models' fine-grained capabilities in medical imaging across anatomical regions, imaging modalities, and cognitive hierarchies. It assesses reasoning reliability, shortcut behavior, and foundational perceptual skills to reveal how models handle clinical VQA beyond aggregate performance. Use when the user wants to benchmark on MedRCube, or asks about evaluating this task. Reports MedRCube Score.
Evaluates large language models' ability to detect, localize, and correct clinical errors in medical texts across Japanese and English. It probes cross-lingual medical reasoning, precise error identification, and accurate text correction capabilities. Use when the user wants to benchmark on MedRECT-ja, MedRECT-en, or asks about evaluating this task. Reports Error Detection F1.
Evaluates LLMs' ability to perform medical question answering under realistic retrieval-augmented generation (RAG) conditions. It probes four key capabilities: handling insufficient or noisy context, integrating multi-source information via sub-questions, detecting factual errors in retrieved documents, and standard RAG performance. Use when the user wants to benchmark on MedRGB, or asks about evaluating this task. Reports accuracy.
Evaluates the segmentation accuracy and inference efficiency of lightweight, promptable medical image models across diverse imaging modalities. It probes the trade-off between mask quality (overlap and boundary alignment) and computational speed on both 2D and 3D medical data. Use when the user wants to benchmark on MedSAMSlicer Competition Dataset, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates promptable 3D medical image and video segmentation across diverse organs, lesions, and imaging modalities. It probes spatial consistency across 3D slices and temporal continuity across video frames using bounding box prompts. Use when the user wants to benchmark on Holdout 3D Test Set, CAMUS, SUN, or asks about evaluating this task. Reports Dice similarity coefficient (DSC).
Evaluates multimodal models on multi-grained video description and fine-grained temporal/perceptual visual reasoning using long-form medical videos. Use when the user wants to benchmark on SVU-31K, or asks about evaluating this task. Reports CI, DO, CU, TU.
This benchmark evaluates sequential visual grounding in medical imaging, specifically testing a model's ability to perform cross-image semantic alignment, detect differences between sequential scans, and identify consistent regions across time. Use when the user wants to benchmark on MedSG-Bench, or asks about evaluating this task. Reports average IoU.
Evaluates multimodal large language models on their ability to generate differential diagnoses (DDx) and select final diagnoses (FDx) for complex clinical cases. It probes cross-modal evidence calibration, testing how models weigh textual versus visual clinical evidence, and measures their sensitivity to specific evidence types. Use when the user wants to benchmark on MEDSYN, or asks about evaluating this task. Reports FDx SelectionAcc. (%).
Evaluates vision-language models' ability to interpret multiple medical images, integrate cross-view evidence, and perform stepwise clinical reasoning for differential diagnosis. It probes visual grounding, evidence alignment, and reasoning depth beyond simple answer matching. Use when the user wants to benchmark on MedThinkVQA, or asks about evaluating this task. Reports Stepwise Reasoning Evaluation.
Evaluates heterogeneous medical video understanding across classification, grounding, and captioning tasks. Probes a model's ability to perform clinical safety checks, predict surgical actions, assess skills, localize temporal/spatial events, and generate precise medical video descriptions. Use when the user wants to benchmark on MedVidBench (Standard), or asks about evaluating this task. Reports accuracy.
Evaluates a multimodal LLM's ability to perform disease classification, visual grounding, and zero-shot reasoning across diverse medical imaging modalities (X-ray, CT, MRI, ultrasound, endoscopy, video) and tasks. Use when the user wants to benchmark on Chest X-ray (10 test sets from 5 public + 5 private), ImageCAS, ASOCA, CCA200, LPPolypVideo/SUN-SEG/CVC-12k, EndoVis18, LDPolyVideo, ISIC16, HAM10000, TN3K, BUID, TBX11K, RSNA Pneumonia, Luna16, DeepLesion, ADNI, LGG, Object-CXR, or asks about...
Evaluates vision-language models on quantitative medical image analysis, specifically anatomical structure detection, tumor/lesion size estimation, and angle/distance measurement. It probes the models' ability to perform precise spatial localization and numeric regression in a clinical context. Use when the user wants to benchmark on MedVision, or asks about evaluating this task. Reports IoU>0.5.
Evaluates expert-level medical reasoning and clinical understanding using real-world board exam questions, patient records, and multimodal clinical data. Use when the user wants to benchmark on MedXpertQA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates vision-language models on multimodal educational exam questions in Persian and English. It specifically probes visual grounding, reasoning capabilities, and robustness to missing or mismatched visual cues across different prompting strategies. Use when the user wants to benchmark on MEENA (PersianMMMU), or asks about evaluating this task. Reports accuracy.
This evaluation protocol assesses the robustness of automatic speech recognition (ASR) systems in multi-talker meeting scenarios under varying microphone configurations (close-talk, distant, and source-separated). It measures transcription accuracy across different levels of speaker overlap and acoustic conditions, while also analyzing how diarization errors propagate to downstream ASR performance. Use when the user wants to benchmark on LibriCSS, AMI, AliMeeting, or asks about evaluating thi...
Evaluates the quality of abstractive and extractive summarization systems on city council meeting transcripts and videos. It probes a model's ability to capture key decisions, maintain factual accuracy, and produce fluent, coherent, and non-redundant summaries. Use when the user wants to benchmark on MeetingBank, or asks about evaluating this task. Reports Average Score.
Evaluates large language models' scientific reasoning capabilities across general science, specialized domains (chemistry, CS, medicine, physics), and mathematical problem-solving. It tests the model's ability to follow chain-of-thought prompting, extract precise answers (including units), and correctly identify multiple-choice options. Use when the user wants to benchmark on MegaScience Evaluation Suite (MMLU, GPQA-Diamond, MMLU-Pro, SuperGPQA, SciBench, OlympicArena, ChemBench, CS-Bench, Me...
Evaluates the quality and evidential support of automatically generated Wikipedia question-answer pairs and their cited source documents. It probes whether sources actually contain the information claimed in passages, and measures the strength of support for QA pairs. Use when the user wants to benchmark on MegaWika, or asks about evaluating this task. Reports answerability.
Evaluates a 4B multi-modal medical agentic model's ability to perform clinical reasoning, tool use, and multi-step interaction across radiology, pathology, and clinical domains. Probes strategy selection (when to use tools vs direct reasoning) and execution policy under various agent frameworks. Use when the user wants to benchmark on MIMIC-CXR-VQA, ChestAgentBench, PathVQA, SLAKE, VQA-RAD, OmniMed, MedXpertQA, MedQA, PubMedQA, NEJM, NEJM Ext., MIMIC-IV, MedQA Ext., or asks about evaluating t...
This evaluation protocol assesses a model's ability to classify Android and Windows PE binaries as benign or malicious using only static features. It probes robustness, generalization across malware families, and discriminative power under concept drift and evolving threat scenarios. Use when the user wants to benchmark on CIC-AndMal2020, BODMAS, EMBOD, or asks about evaluating this task. Reports Accuracy.
Evaluates the on-device runtime performance, energy efficiency, thermal behavior, and quality of experience of LLM inference across mobile and edge platforms under varying quantization and framework configurations. Use when the user wants to benchmark on OpenAssistant/oasst1 (filtered subset), or asks about evaluating this task. Reports throughput.
This benchmark evaluates multimodal long-term conversational memory in MLLM agents across multi-session dialogues. It probes the agent's ability to extract, adapt, reason over, and manage evolving visual and textual information, including handling temporal dependencies, conflicting updates, and knowledge gaps. Use when the user wants to benchmark on Mem-Gallery, or asks about evaluating this task. Reports answer correctness.
Evaluates the click-through rate prediction capability of deep learning recommendation models under varying memory budgets and compression techniques. It probes how well alternative representation schemes (like Bloom filter encoding) maintain recommendation accuracy while drastically reducing embedding table sizes. Use when the user wants to benchmark on Avazu, Criteo-Kaggle, Criteo-Terabyte, or asks about evaluating this task. Reports ROC-AUC.