
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates multimodal models' ability to answer questions about charts by extracting visual and textual information. It probes robustness to missing labels and geometric perturbations, distinguishing between simple pattern matching and true structural reasoning. Use when the user wants to benchmark on ChartQA, Charixv, or asks about evaluating this task. Reports Relaxed Accuracy (RA).
Evaluates automatic chart-to-text summarization models on their ability to generate accurate, fluent, and informative summaries from chart metadata and data tables. It probes factual correctness, trend capture, and hallucination resistance across short and long summary formats. Use when the user wants to benchmark on ChartSumm, or asks about evaluating this task. Reports BLEU.
Evaluates visual language models on complex chart understanding and reasoning tasks, including data extraction, comparison, and trend analysis across diverse chart types and domains. Use when the user wants to benchmark on ChartQA-Pro, CharXiv, ChartMuseum, ChartX, ChartBench, EvoChart, or asks about evaluating this task. Reports average score.
Evaluates multimodal large language models' ability to understand real-world charts through descriptive and reasoning tasks. It probes capabilities like information extraction, pattern recognition, counting, compositional reasoning, and robustness to chart complexity (e.g., multiple subplots) and unanswerable queries. Use when the user wants to benchmark on CharXiv, or asks about evaluating this task. Reports accuracy.
Evaluates large language models by collecting human preference votes on pairwise responses to real-world prompts, then ranks them using Bradley-Terry models to measure alignment and real-world utility. Use when the user wants to benchmark on Chatbot Arena, or asks about evaluating this task. Reports BT coefficients.
Evaluates ChatGPT's zero-shot multitask capabilities across summarization, machine translation, sentiment analysis, question answering, dialogue, and misinformation detection. It probes the model's generalization, reasoning, multilingual understanding, and task-specific performance without fine-tuning. Use when the user wants to benchmark on CNN/DM, SAMSum, FLoRes-200, NusaX, bAbI, EntailmentBank, CLUTRR, StepGame, Pep-3k, COVID-Social, COVID-Scientific, MultiWOZ2.2, OpenDialKG, or asks about...
Evaluates ChatGPT's few-shot generation and reasoning capabilities across multiple NLP tasks, including question answering, commonsense reasoning, natural language inference, and sentiment analysis. It probes the model's ability to follow task-specific formalizations, leverage retrieved demonstrations, and mitigate hallucination through self-verification. Use when the user wants to benchmark on SQuADv2, TQA, MRQA-OOD, CSQA, StrategyQA, RTE, CommitmentBank, SST-2, IMDB, Yelp, or asks about eva...
Evaluates multi-turn conversational question answering over both long and short documents, including text-only and tabular contexts. It probes a model's ability to retrieve relevant context, handle topic switching, perform arithmetic reasoning, and correctly identify unanswerable queries. Use when the user wants to benchmark on Doc2Dial, QuAC, QReCC, TopiOCQA, INSCIT, CoQA, DoQA, ConvFinQA, SQA, HybridDial, or asks about evaluating this task. Reports Average F1/EM across 10 datasets.
Evaluates a model's ability to perform multi-round multimodal referring and grounding, requiring logical consistency across dialogue turns while generating accurate text responses and bounding box coordinates for visual instances. Use when the user wants to benchmark on CB-LC, RefCOCOg, COCO 2017, or asks about evaluating this task. Reports BERT(·).
Evaluates a multimodal time series foundation model on zero-shot forecasting, context-guided forecasting, and time series question answering. It probes the model's ability to handle discretized numerical time series alongside textual prompts for prediction and feature recognition without task-specific fine-tuning. Use when the user wants to benchmark on Electric, Exchange, Traffic, Weather, and ETT datasets, Context-guided multimodal datasets, Synthesized TSQA dataset, or asks about evaluatin...
Evaluates the effectiveness of bottom-up versus top-down transformations for solving linear constrained Horn clauses (CHCs) using software verification workflows. It measures how many verification tasks a solver can successfully resolve within a strict time limit. Use when the user wants to benchmark on CHC-COMP21 LIA-Lin track, or asks about evaluating this task. Reports solved_tasks.
Evaluates the perceptual quality and human preference alignment of generative image models by comparing their discrete visual token distributions and sequence statistics against human ratings. Use when the user has predictions and gold and needs to compute CHD, CMMS.
Evaluates large language models' ability to process, generate, and retrieve molecular information across multiple modalities (SMILES, InChI, SELFIES, graphs, captions, IUPAC names, images). It probes cross-modal compatibility, chemical knowledge acquisition, and molecular property prediction capabilities. Use when the user wants to benchmark on ChEBI-20-MM, or asks about evaluating this task. Reports ROC_AUC.
Evaluates large language models' ability to perform stepwise chemical reasoning and molecular structure manipulation. It probes capabilities in molecular understanding, functional group/ring recognition, scaffold extraction, and SMILES-based molecule editing and reaction prediction. Use when the user wants to benchmark on ChemCoTBench, or asks about evaluating this task. Reports accuracy.
Evaluates LLM-based chemistry agents on four core computational chemistry tasks: text-to-molecule generation, molecule-to-text captioning, molecular property prediction, and reaction product prediction. It probes the model's ability to translate between chemical language and structures, predict physicochemical properties, and correct tool errors via hierarchical agent stacking. Use when the user wants to benchmark on ChemLLMBench, or asks about evaluating this task. Reports Accuracy.
Evaluates multimodal large language models' ability to solve Olympiad-level theoretical chemistry problems requiring visual perception, chemical reasoning, and structured problem-solving. It specifically probes the visual perception bottleneck in chemistry tasks and tests the effectiveness of multi-agent orchestration and structured visual enhancement. Use when the user wants to benchmark on ChemO, or asks about evaluating this task. Reports normalized rubric-based score.
Evaluates a model's ability to play chess and solve tactical puzzles without explicit search, testing its capacity for long-horizon planning and generalization to novel board states. Use when the user wants to benchmark on Lichess puzzles, or asks about evaluating this task. Reports Lichess Elo.
Evaluates a CNN's ability to classify chest X-ray images into disease categories (COVID-19, pneumonia, tuberculosis, normal) using various preprocessing techniques. It probes robustness across different dataset sizes and class distributions. Use when the user wants to benchmark on Multiclass Chest X-ray Dataset, Hamad Medical Corporation Tuberculosis Dataset, Pneumonia Dataset, NIH Chest X-ray Dataset, or asks about evaluating this task. Reports AUC.
This evaluation probes a model's ability to generate clinically coherent and accurate natural language reports from chest X-ray images. It measures both surface-level linguistic similarity to ground-truth radiology reports and the clinical correctness of extracted pathological findings. Use when the user wants to benchmark on Indiana U. Chest X-Ray, MIMIC-CXR, or asks about evaluating this task. Reports BLEU-1.
Evaluates a model's ability to perform multi-label classification of 14 thoracic diseases on chest X-ray images and localize pathological regions using attention maps. It probes whether anatomically grounded feature weighting improves disease detection over global feature fusion or saliency-based methods. Use when the user wants to benchmark on Chest X-ray14, or asks about evaluating this task. Reports AUROC.
Evaluates an AI agent's ability to perform multi-step medical reasoning and tool orchestration for chest X-ray interpretation. It probes capabilities across seven clinically relevant categories: detection, classification, localization, comparison, relationship, diagnosis, and characterization. Use when the user wants to benchmark on ChestAgentBench, or asks about evaluating this task. Reports Accuracy (%).
Evaluates instance-level detection and segmentation of thoracic diseases on chest X-rays. It probes a model's ability to localize and classify 13 disease categories while handling domain-specific challenges like ambiguous boundaries and disease co-occurrence. Use when the user wants to benchmark on ChestX-Det, DR-private, or asks about evaluating this task. Reports AP_50^bb.
Evaluates deep learning models for binary classification of chest X-rays into normal versus pneumonia categories, while also assessing the spatial interpretability of model predictions using Grad-CAM heatmaps. Use when the user wants to benchmark on Chest X-Rays dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates the effectiveness of attribute-neutralization models in removing demographic bias (sex and age) from chest X-ray images while preserving diagnostic utility for 15 disease findings. It probes the trade-off between demographic leakage suppression and clinical performance across varying edit intensities. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AI-Judge AUC.
Evaluates deep learning models for multi-label chest X-ray pathology classification. It probes the capability of CNN architectures to detect 14 specific thoracic diseases, testing the impact of transfer learning, network depth, and fusion of non-image patient demographics (age, gender, view position) on diagnostic accuracy. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to detect and localize multiple pathologies in high-resolution chest X-ray images. It specifically probes the model's robustness to severe class imbalance and its capacity to leverage explicit spatial location information for pathology classification. Use when the user wants to benchmark on ChestX-Ray14, PLCO, or asks about evaluating this task. Reports AUC.
Evaluates weakly-supervised multi-label classification and spatial localization of eight common thoracic diseases on chest X-rays. It probes a model's ability to detect disease presence from image-level labels and localize pathological regions using only bounding box annotations during testing. Use when the user wants to benchmark on ChestX-ray8, or asks about evaluating this task. Reports AUC.
Evaluates a vision-language model's ability to perform interactive localization, region classification, and text generation on chest X-rays. It probes zero-shot multitask capabilities, including sentence grounding, pathology detection, and customizable report generation. Use when the user wants to benchmark on MS-CXR, VinDrCXR, NIH8, CIG, MIMIC-CXR, or asks about evaluating this task. Reports mAP.
Evaluates vision-language foundation models on chest X-ray interpretation across two axes: image perception (view classification, disease identification/classification, VQA, reasoning) and textual understanding (findings generation, summarization). It probes clinical reasoning, visual grounding, and medical text generation capabilities. Use when the user wants to benchmark on MIMIC-CXR, CheXpert, SIIM, RSNA, OpenI, SLAKE, Rad-Restruct, or asks about evaluating this task. Reports accuracy.
Evaluates a chest X-ray vision-language foundation model across zero-shot classification, cross-modal retrieval, and adapted downstream tasks (classification, segmentation, report generation). It specifically probes data and compute efficiency, as well as the model's ability to represent long-tailed thoracic diseases without aggressive scaling. Use when the user wants to benchmark on SIIM-PTX, Pneumonia2017, TBX11K, CheXpert, MIMIC-CXR, ChestX-ray14, VinDr-CXR, VinDr-PCXR, or asks about evalu...
Evaluates text-to-image generative models for synthetic chest radiograph generation across three dimensions: fidelity (image quality and mode coverage), privacy (memorization and re-identification risks), and utility (downstream clinical performance for classification and report generation). It probes whether synthetic medical images can match real data in clinical tasks while preserving patient privacy. Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Re...
Evaluates the diagnostic accuracy of predictive algorithms and human radiologists on chest X-ray images for detecting atelectasis. It specifically probes whether human expertise provides actionable, non-redundant information on input subsets where algorithms are algorithmically indistinguishable. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
Evaluates an automated rule-based pipeline's ability to extract clinical observations from free-text radiology reports. It specifically tests the system's capacity to classify mentions as positive, negative, or uncertain, and to aggregate them into structured labels for 14 predefined observations. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports F1 score.
Evaluates a multi-task deep learning framework for thorax disease classification and weakly-supervised localization on chest X-rays. It probes the model's ability to detect multiple pathologies and accurately localize abnormal regions without requiring pre-annotated bounding boxes during training. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, MIMIC-CXR, or asks about evaluating this task. Reports AUC.
Evaluates the generalization of chest X-ray deep learning models to two clinically relevant distribution shifts: smartphone photographs of digital X-rays (introducing visual artifacts like blur and glare) and external institutional data. It probes robustness to imaging degradation and cross-institutional heterogeneity in multi-label pathology detection. Use when the user wants to benchmark on CheXpert, NIH, or asks about evaluating this task. Reports MCC.
This evaluation probes the transferability of ImageNet-pretrained CNN architectures to chest X-ray interpretation, measuring how model size, architecture family, and pretraining affect classification performance and parameter efficiency on the CheXpert dataset. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports CheXpert AUC.
Evaluates the ability of recurrent graph neural networks to forecast spatiotemporal epidemiological time series. It probes how well models capture spatial dependencies between adjacent regions and temporal dynamics like seasonality and zero-inflation over multiple forecasting horizons. Use when the user wants to benchmark on Chickenpox Cases in Hungary, or asks about evaluating this task. Reports mean squared error.
Evaluates fundamental lexical comprehension and generation capabilities of multilingual LLMs across thousands of languages. It probes word-level translation, context-aware translation, translation-conditioned language modeling, and bag-of-words machine translation tasks to measure basic linguistic competence beyond high-resource languages. Use when the user wants to benchmark on ChiKhaPo, or asks about evaluating this task. Reports language score.
Evaluates LLM safety alignment across four child developmental stages (ages 6–17) using simulated agents grounded in developmental psychology. It probes how models handle sensitive contexts, boundary-testing, and age-specific cognitive limitations in multi-turn interactions. Use when the user wants to benchmark on ChildSafe Dataset, or asks about evaluating this task. Reports semantic_safety_score.
Evaluates the effectiveness of a time-domain speech enhancement frontend in improving automatic speech recognition performance on noisy and reverberant speech. It probes the model's ability to enhance speech without introducing distortion that degrades downstream ASR accuracy. Use when the user wants to benchmark on CHiME-2, or asks about evaluating this task. Reports WER.
This benchmark evaluates speech enhancement models by measuring how well they clean noisy speech features before they are processed by a downstream automatic speech recognition (ASR) system. It probes the model's ability to preserve speech structure and reduce noise in challenging real-world far-field conditions. Use when the user wants to benchmark on CHiME-3, or asks about evaluating this task. Reports WER.
Evaluates the robustness of Automatic Speech Recognition (ASR) models under various real-world noise conditions (bus, cafe, pedestrian, street junction) using adapter-based fine-tuning and speech enhancement front-ends. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.
Evaluates multi-channel end-to-end speech recognition systems in noisy, real-world environments using neural beamforming front-ends combined with CTC-CRF acoustic models. It probes the model's ability to leverage single-channel data via pre-training, data scheduling, or simulation to improve robustness and accuracy. Use when the user wants to benchmark on CHiME4, AISHELL-4, or asks about evaluating this task. Reports WER.
Evaluates automatic speech recognition systems on real-world meeting and close-talking microphone speech. It measures how well acoustic models trained on raw waveforms generalize to challenging multi-microphone environments compared to traditional feature-based baselines. Use when the user wants to benchmark on CHiME4, AMI, or asks about evaluating this task. Reports WER.
Evaluates the robustness of automatic speech recognition systems in noisy environments by measuring word error rates on enhanced speech from speaker extraction models across matched and mismatched noisy conditions. Use when the user wants to benchmark on CHiME-4, VoiceBank-DEMAND, WHAM!, or asks about evaluating this task. Reports WER.
Evaluates end-to-end speech recognition and speech enhancement performance in noisy, reverberant multi-channel conditions. It probes the model's ability to jointly dereverberate, denoise, and transcribe speech using self-supervised learning representations. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.
Evaluates a model's ability to identify scientific idea recombinations in research abstracts, extract the involved scientific concepts (entities), and classify the type of recombination relation (inspiration or blend) between them. Use when the user wants to benchmark on CHIMERA, or asks about evaluating this task. Reports Precision, Recall, F1.
This benchmark evaluates the safety of large language models in Chinese by testing their ability to correctly classify text as safe or unsafe across multiple sensitive categories. It probes whether models can reliably detect harmful, policy-violating, or sensitive content in a Chinese-language context using both generation-based and perplexity-based strategies. Use when the user wants to benchmark on ChineseSafe, or asks about evaluating this task. Reports accuracy.
Evaluates Chinese large language models' world knowledge, academic understanding, and multi-dimensional alignment after pretraining or instruction fine-tuning. It probes the model's ability to follow instructions, reason across domains, and maintain safety and helpfulness standards in Chinese. Use when the user wants to benchmark on C-Eval, CMMLU, Alignbench, or asks about evaluating this task. Reports accuracy.
Evaluates the knowledge, reasoning, instruction-following, and safety alignment capabilities of Chinese instruction-tuned LLMs across academic, professional, open-ended, and safety-critical domains. Use when the user wants to benchmark on C-Eval, CMMLU, BELLE-EVAL, SafetyBench, or asks about evaluating this task. Reports log-likelihood.