
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the physiological realism and clinical fidelity of synthetic 12-lead ECGs by measuring how often expert clinicians can correctly distinguish them from real clinical recordings, and how accurately they can diagnose specific pathologies in both synthetic and real signals. Use when the user wants to benchmark on MedalCare-XL, or asks about evaluating this task. Reports accuracy.
Evaluates clinical prediction capabilities on unstructured notes and structured EHR data. It benchmarks zero-shot LLMs, finetuned BERTs, and conventional ML/DL models on mortality, readmission, and length-of-stay prediction tasks. The setup tests out-of-the-box prompting versus task-specific finetuning across diverse model families. Use when the user wants to benchmark on MIMIC-IV, TJH, or asks about evaluating this task. Reports AUROC.
Evaluates clinical text-to-SQL capabilities by requiring models to generate executable BigQuery queries that perform multi-table joins, temporal reasoning, and patient-similarity cohort analysis on electronic health record data. Use when the user wants to benchmark on CLINSQL, or asks about evaluating this task. Reports Execution Score.
This evaluation protocol probes a model's ability to perform continual learning (CIL) using vision-language models (CLIP) without catastrophic forgetting. It measures how well the model retains knowledge from previous tasks while adapting to new ones, specifically testing stability-plasticity trade-offs under varying task splits and replay settings. Use when the user wants to benchmark on ImageNetR, ImageNetA, CIFAR-100, or asks about evaluating this task. Reports Last.
Evaluates zero-shot classification performance of CLIP-based vision-language models on chest X-rays, assessing fairness across demographic subgroups (age, sex, race) and robustness to spurious correlations (presence of chest drains in pneumothorax cases). Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Reports AUPRCadj.
Evaluates cross-modal (text-image) retrieval and text-only embedding performance. Probes zero-shot retrieval accuracy, semantic similarity, and overall text embedding capability across diverse benchmarks. Use when the user wants to benchmark on CLIP Benchmark, MTEB, or asks about evaluating this task. Reports r@5.
Evaluates the zero-shot robustness of CLIP models to natural distribution shifts by measuring classification accuracy on four ImageNet-derived datasets. It probes how pre-training data composition and quality affect generalization to out-of-distribution images like sketches, renditions, and novel viewpoints. Use when the user wants to benchmark on ImageNet-V2, ImageNet-R, ImageNet-Sketch, ObjectNet, or asks about evaluating this task. Reports accuracy.
Evaluates the naturalness and quality of synthesized speech across single-speaker, multi-speaker, and multi-emotion TTS models. It measures how closely generated audio matches human ground truth in terms of overall speech quality and emotional similarity. Use when the user wants to benchmark on Baker, AISHELL3, LJSpeech, LibriTTS, Emotional Speech Dataset (ESD), or asks about evaluating this task. Reports MOS.
Compute the CLIPImageQualityAssessment metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CLIPImageQualityAssessment, or asks how to score with CLIPImageQualityAssessment.
Compute the CLIPScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CLIPScore, or asks how to score with CLIPScore.
Evaluates a model's ability to determine whether two code snippets share the same semantics or to retrieve relevant code snippets from a repository. It probes semantic code similarity and code retrieval capabilities. Use when the user wants to benchmark on BigCloneBench, POJ-104, or asks about evaluating this task. Reports Overall.
Evaluates the ability of audio classifiers to distinguish between real human speech and AI-generated cloned voices across single and multi-speaker scenarios. It also probes robustness against adversarial audio laundering, including additive noise and AAC transcoding, to assess how well different feature representations (learned, spectral, perceptual) generalize and resist degradation. Use when the user wants to benchmark on ElevenLabs (EL), Uberduck (UD), WaveFake (WF), TIMIT-ElevenLabs, or a...
Evaluates a universal adversarial perturbation framework designed to defend against zero-shot voice cloning TTS models. It measures how well the perturbation degrades cloned audio quality and speaker similarity while preserving the perceptual fidelity of the original protected speech. Use when the user wants to benchmark on VCTK, LibriSpeech ASR, LibriTTS-R, LJSpeech, Common Voice, or asks about evaluating this task. Reports DSR.
Evaluates the ability of voice cloning models to preserve speaker identity and acoustic characteristics across different speech conditions, including neutral and emotional speech. It measures how closely generated audio matches the reference speaker's embedding and signal properties without human intervention. Use when the user wants to benchmark on LS test-clean, TESS, or asks about evaluating this task. Reports cosine similarity (WavLM).
Evaluates cross-modal retrieval and zero-shot classification capabilities for remote sensing imagery (SAR and multispectral optical) paired with text descriptions. It probes how well unified semantic embeddings align heterogeneous geospatial data with natural language for crisis event and land cover analysis. Use when the user wants to benchmark on CrisisLandMark, or asks about evaluating this task. Reports nDCG@1000.
Evaluates a robot's ability to select effective grasp poses on hanging, folded cloth to maximize unfolding coverage. It probes material-aware perception, geometric reasoning, and policy-based grasp generation under complex folding configurations. Use when the user wants to benchmark on ICRA 2024 Cloth Competition dataset, or asks about evaluating this task. Reports relative coverage.
Evaluates the robustness of image classification models when trained on datasets with inherent label noise and class imbalance. It probes how effectively learning algorithms can filter mislabeled samples and adapt to skewed class distributions without manual curation. Use when the user wants to benchmark on Clothing1mPP, or asks about evaluating this task. Reports Noise Rate.
Evaluates zero-shot language-queried audio source separation on diverse environmental sounds using natural language captions. The benchmark tests isolation of a target sound from a concatenated background mixture. Use when the user wants to benchmark on Clotho v2, or asks about evaluating this task. Reports SDRi.
This benchmark evaluates how effectively different Virtual Machines (VMs) can execute specific application workloads by ranking them according to weighted hardware attributes. It probes the capability to map domain-specific application requirements to underlying infrastructure performance characteristics. Use when the user wants to benchmark on Cloud VM Benchmarking Suite, or asks about evaluating this task. Reports S_i.
Evaluates the ability of systems to detect context-aware anomalies in cloud environments by jointly analyzing system logs and performance metrics, and to classify the specific anomaly scenario. It also tests generalization to point-level anomalies using log-only or metric-only data. Use when the user wants to benchmark on CloudAnoBench, BGL, Thunderbird, HDFS_v1, or asks about evaluating this task. Reports F1-score.
Evaluates time series forecasting models, particularly pre-trained Transformers, on cloud operations data. It probes zero-shot generalization, architectural efficiency, and scaling behavior against classical and deep learning baselines. Use when the user wants to benchmark on azure2017, borg2011, ali2018, or asks about evaluating this task. Reports sMAPE.
Evaluates continual learning capabilities on Visual Question Answering (CLVQA) by measuring how well a model retains knowledge from previous tasks while learning new ones across scene-incremental and function-incremental settings. Use when the user wants to benchmark on CLOVE, or asks about evaluating this task. Reports average accuracy (%).
Evaluates a model's ability to perform generative error correction (GER) for automatic speech recognition by reformulating the task as a cloze test. The model must select the correct hypothesis from a 5-best N-best list to minimize word error rate while maintaining source speech fidelity. Use when the user wants to benchmark on HyPoradise (GER benchmark), or asks about evaluating this task. Reports WER (%).
Evaluates Chinese language understanding across nine diverse tasks, including text classification, natural language inference, semantic similarity, and machine reading comprehension. It probes a model's ability to handle Chinese-specific linguistic phenomena, whole-word masking, and token-level vs. global understanding through a standardized fine-tuning pipeline. Use when the user wants to benchmark on CLUE, or asks about evaluating this task. Reports Accuracy.
Evaluates few-shot learning capabilities of pre-trained language models across sentence classification, question answering, and named entity recognition tasks. It measures how well models adapt with limited labeled examples (10, 20, 30 shots) compared to fully supervised settings and human performance. Use when the user wants to benchmark on SST-2, MNLI, NER, MRC, or asks about evaluating this task. Reports macro-averaged results.
Compute the ClusterAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ClusterAccuracy, or asks how to score with ClusterAccuracy.
Evaluates a contrastive multi-level graph neural network for session-based recommendation by measuring its ability to predict the next item in a user session using pairwise and high-order transition patterns. Use when the user wants to benchmark on Tmall, Diginetica, Nowplaying, or asks about evaluating this task. Reports Recall@K.
Evaluates the transferability and quality of self-supervised multiview representations by measuring downstream performance on image classification, video action recognition, and semantic segmentation tasks. Use when the user wants to benchmark on ImageNet, UCF-101, HMDB-51, NYU-Depth-V2, STL-10, or asks about evaluating this task. Reports Top-1 classification accuracy (%).
Evaluates Chinese medical text embedding models across retrieval, reranking, and semantic textual similarity tasks. It probes the model's ability to capture domain-specific semantic alignment while measuring the trade-off between retrieval accuracy and inference efficiency. Use when the user wants to benchmark on CMedTEB, or asks about evaluating this task. Reports Avg.
Evaluates audio-text LLMs on music instruction following by framing traditional music information retrieval (MIR) tasks as prompts. It measures how accurately models follow instructions to perform classification, regression, captioning, and sequential audio analysis tasks. Use when the user wants to benchmark on CMI-Bench, or asks about evaluating this task. Reports Accuracy.
Evaluates music reward models on their ability to align with human aesthetic judgments and follow compositional multimodal instructions (text, lyrics, audio). It probes both absolute musicality scoring and relative pairwise preference ranking across diverse generation models. Use when the user wants to benchmark on PAM, MusicEval, Music Arena, CMI-Pref, or asks about evaluating this task. Reports Linear Correlation Coefficient (LCC), Spearman Rank Correlation (SRCC), Kendall-Tau (K-Tau), Pair...
Evaluates the quality of text-to-speech models trained on the CML-TTS dataset across seven low-resource languages. It probes speaker similarity preservation and text fidelity in synthesized audio under both seen and unseen speaker (zero-shot) conditions. Use when the user wants to benchmark on CML-TTS, or asks about evaluating this task. Reports SECS.
Evaluates large language models' Chinese language understanding and multitask knowledge across 67 subjects spanning STEM, humanities, social sciences, and China-specific domains. It probes memorization, reasoning, and instruction-following capabilities in a multiple-choice question-answering format. Use when the user wants to benchmark on CMMLU, or asks about evaluating this task. Reports macro average accuracy.
Evaluates the zero-shot cross-modality transfer capability of open-vocabulary object detectors from RGB to X-ray imaging. It measures how well pre-trained RGB detectors can localize and classify objects in X-ray images without any fine-tuning or labeled X-ray data. Use when the user wants to benchmark on DET-COMPASS, PIXray, PIDray, CLCXray, DvXray, HiXray, or asks about evaluating this task. Reports AP.
Evaluates the quality of generated responses in document-grounded conversations, specifically measuring how well models leverage external document context to produce engaging and fluent multi-turn dialogue. It assesses both automatic language modeling metrics and human-perceived response quality. Use when the user wants to benchmark on CMU.DoG, or asks about evaluating this task. Reports Perplexity.
Evaluates large language models' ability to generate valid, optimized molecular structures (SMILES) that satisfy multiple conflicting pharmacological and physicochemical property constraints. The benchmark probes the model's capacity for multi-objective reinforcement alignment, scaffold preservation, and strict adherence to property-wise improvement margins under both in-domain and out-of-distribution settings. Use when the user wants to benchmark on C-MuMOInstruct, or asks about evaluating t...
Evaluates the trade-offs between predictive accuracy, model compression, and dynamic inference efficiency of CNN optimization techniques (pruning, quantization, early-exit) for edge deployment. It probes how different architectures handle static compression versus input-adaptive latency reduction under hardware-constrained conditions. Use when the user wants to benchmark on Unspecified classification dataset, or asks about evaluating this task. Reports accuracy (%).
This evaluation benchmarks CNN-based models on automatic music tagging, measuring their ability to predict multiple genre, instrument, and mood labels from audio spectrograms. It assesses both standard classification performance and robustness to audio transformations like time-stretching and pitch shifting. Use when the user wants to benchmark on MagnaTagATune, Million Song Dataset, MTG-Jamendo, or asks about evaluating this task. Reports ROC-AUC.
Evaluates abstractive summarization quality by scoring generated summaries against human-written references and expert rubric-based scores. It measures how well automatic metrics correlate with human judgments across different summarization systems. Use when the user wants to benchmark on CNNDM, or asks about evaluating this task. Reports COMET.
Tests the model's ability to classify CNS-active versus CNS-inactive drugs and enrich active compounds from large virtual screening databases. It probes generalization on small-sample molecular datasets using external validation. Use when the user wants to benchmark on CNS Drug Dataset, or asks about evaluating this task. Reports AUC.
Evaluates the capability of models to detect section boundaries in clinical notes at the token level. It probes how well different architectures handle structured sentence-level segmentation versus unstructured freetext narrative variability in medical records. Use when the user wants to benchmark on MIMIC-IV Clinical Notes, or asks about evaluating this task. Reports Token-level F1.
Evaluates a joint deep learning architecture for simultaneous monocular depth estimation and semantic segmentation on aerial drone imagery, measuring prediction accuracy and inference speed against single-task and joint baselines. Use when the user wants to benchmark on MidAir, Aeroscapes, or asks about evaluating this task. Reports mIoU.
Evaluates long-horizon agentic reasoning, tool-augmented decision making, and factual grounding under conflict-aware verification. It probes how well systems can audit divergent reasoning steps, maintain structured knowledge, and produce accurate answers across multi-hop and interdisciplinary tasks. Use when the user wants to benchmark on GAIA, HLE, Chinese-SimpleQA, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of AI-generated image detectors to distinguish between real and synthetic images across diverse generative models, lossy compression formats, and real-world in-the-wild sources. It specifically probes out-of-distribution generalization and robustness to common post-processing transformations like JPEG compression, blurring, and noise. Use when the user wants to benchmark on Co-SpyBench, Co-SpyBench/in-the-wild, AIGCDetectBenchmark, GenImage, or asks about evaluating this...
Evaluates a model's ability to mitigate gender bias in downstream NLP tasks by measuring performance disparities across demographic groups. It probes whether models assign equal similarity scores to gender-swapped sentence pairs, maintain neutrality in natural language inference, and classify occupations without gender-based true positive rate gaps. Use when the user wants to benchmark on Bias-STS-B, Bias-NLI, Bias-in-Bios, or asks about evaluating this task. Reports average absolute differen...
Evaluates the ability of debiasing frameworks to transform biased (MNAR) recommendation data into unbiased (MAR) representations, measuring how well debiased rankings align with ground-truth user preferences. It probes whether reweighting or perturbation mechanisms successfully mitigate selection and staleness biases without degrading predictive performance. Use when the user wants to benchmark on Coat, or asks about evaluating this task. Reports AUC.
Evaluates large language models' ability to generate correct, compilable COBOL code from natural language specifications, and to translate bidirectionally between COBOL and Java. It probes functional correctness, compilation reliability, and practical utility for legacy system modernization. Use when the user wants to benchmark on COBOLEval, COBOLCodeBench, COBOL-JavaTrans, or asks about evaluating this task. Reports Pass@1.
Evaluates neural abstractive summarization models on their ability to generate factual, fluent, and relevant narrative summaries of randomized controlled trials (RCTs) from Cochrane systematic reviews. Probes the models' susceptibility to hallucination and their capacity to correctly infer the directionality of clinical findings. Use when the user wants to benchmark on Cochrane RCT Summaries, or asks about evaluating this task. Reports Manual Factuality.
Evaluates the ability of vision-language models to re-rank candidate image captions based on their semantic alignment with extracted visual context. It probes how well models can leverage object-level visual information to improve caption relevance and accuracy. Use when the user wants to benchmark on COCO Captions (Karpathy test split), or asks about evaluating this task. Reports BERTScore (B-S).
This benchmark evaluates whether language models appropriately refuse or comply with contextually ambiguous, incomplete, unsupported, or safety-related requests. It probes a model's ability to distinguish between benign queries that should be answered and problematic queries that should be declined, while avoiding exaggerated over-refusal on safe prompts. Use when the user wants to benchmark on CoCoNot, or asks about evaluating this task. Reports compliance rate.