
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates large language models on mathematical reasoning across diverse difficulty levels and languages. It probes the model's ability to convert natural language word problems into structured symbolic representations and execute step-by-step logical derivations without external solvers. Use when the user wants to benchmark on AQUA, MultiArith, GSM8K, MMLU-Redux, Olympiad Bench (English), GaoKao, Olympiad Bench (Chinese), or asks about evaluating this task. Reports exact match.
This benchmark evaluates large language models on formal combinatorial mathematics reasoning within the Lean 4 proof assistant. It probes the model's ability to generate correct, compilable proof scripts and accurately solve fill-in-the-blank combinatorial problems under rigorous automated verification. Use when the user wants to benchmark on CombiBench, or asks about evaluating this task. Reports pass@N.
Evaluates multimodal discrete mathematical reasoning, specifically the ability to parse and solve combinatorial problems involving graphs, grids, and geometric diagrams. It also probes susceptibility to deliberately crafted distractors in multiple-choice formats versus genuine solution construction. Use when the user wants to benchmark on CombiGraph-Vis, or asks about evaluating this task. Reports avg@8.
Probes the semantic fidelity and grammatical correctness of automated translation pipelines when converting English benchmarks into low-resource languages. It measures how well translation methods preserve task structure and downstream model performance consistency. Use when the user wants to benchmark on FLORES, WMT24++, MMLU, or asks about evaluating this task. Reports COMET.
Evaluates the accuracy and overhead of an integrated thermal simulation toolchain (CoMeT) for modeling processor-memory thermal dynamics across 2D, 2.5D, and 3D architectures. It probes the tool's ability to capture thermal coupling, leakage power effects, and DVFS/DTM interactions under diverse compute and memory-intensive workloads. Use when the user wants to benchmark on PARSEC 2.1, SPLASH-2, SPEC CPU2017, or asks about evaluating this task. Reports Temperature.
Evaluates how well models generate or complete commit messages given code diffs and optional historical context. It probes the model's ability to follow coding conventions, match ground truth exactly, and maintain semantic similarity under varying context lengths. Use when the user wants to benchmark on CMG_test, or asks about evaluating this task. Reports ExactMatch@1.
Evaluates the extent to which authors fulfill promises made during peer review rebuttals in their final camera-ready papers, and classifies unfulfilled commitments by severity and difficulty. Use when the user wants to benchmark on ICLR 2025, EMNLP 2024, or asks about evaluating this task. Reports fulfillment rate.
Evaluates multilingual automatic speech recognition (ASR) capabilities, specifically testing speaker generalization and low-resource language adaptation via transfer learning from an English model. It measures how well a model can transcribe audio from diverse, crowdsourced speakers across multiple languages with varying data sizes. Use when the user wants to benchmark on Common Voice, or asks about evaluating this task. Reports character error rate.
Evaluates the image quality and text-image alignment of a text-to-image diffusion model trained on Creative-Commons licensed data, benchmarking it against Stable Diffusion 2 using both automated distribution metrics and human pairwise preference. Use when the user wants to benchmark on MS COCO, PartiPrompts, or asks about evaluating this task. Reports User preference rate.
Evaluates an object detection model's ability to locate and classify form field widgets (text inputs, checkboxes/radio buttons, and signatures) on scanned or digital form pages. It probes sensitivity to input resolution and robustness across different languages and document domains. Use when the user wants to benchmark on CommonForms, or asks about evaluating this task. Reports mAP50-95.
This benchmark probes the commonsense reasoning capabilities of vision-language models by evaluating their ability to match images to text riddles (or vice versa) where the subject entity is replaced with a demonstrative pronoun. It specifically tests relational knowledge retrieval and generalization to unseen knowledge triples. Use when the user wants to benchmark on DANCE Diagnostic Set, or asks about evaluating this task. Reports Acc@50.
This benchmark evaluates a model's ability to answer multiple-choice questions that require real-world commonsense knowledge. It specifically probes whether models can distinguish a correct answer from semantically plausible but factually incorrect distractors based on spatial, causal, or physical reasoning. Use when the user wants to benchmark on CommonsenseQA, or asks about evaluating this task. Reports accuracy.
Evaluates machine translation quality across multiple language pairs and specialized tasks (general translation, terminology-constrained, and automatic post-editing). It measures how well encoder-decoder and decoder-only models generate accurate and fluent target sentences. Use when the user wants to benchmark on ComMT, or asks about evaluating this task. Reports SacreBLEU.
Evaluates the stability, uncertainty quantification, and accuracy of consensus-based community detection algorithms against ground-truth partitions on synthetic and real-world benchmark networks. Use when the user wants to benchmark on Zachary's Karate Network, LFR Benchmark, Ring of Cliques (RC) Benchmark, or asks about evaluating this task. Reports NMI.
This evaluation protocol probes the ability of fake image detectors to generalize across a wide variety of generative models and architectures. It measures how well classifiers trained on diverse synthetic data can distinguish real from generated images in both in-distribution and out-of-distribution settings. Use when the user wants to benchmark on Wang et al. [129], Ojha et al. [90], Synthbuster [7], GenImage [137], Community Forensics (Ours), or asks about evaluating this task. Reports mAP.
Evaluates a model's ability to align latent representations across different conditions (e.g., batch effects, treatment, demographic attributes) while preserving task-relevant information. It measures local mixing quality using nearest-neighbour and silhouette metrics, and assesses predictive utility via classification accuracy on held-out labels. Use when the user wants to benchmark on Tumour / Cell Line, Stimulated / untreated single-cell PBMCs, Single-cell RNA-seq data integration (PBMCs),...
Evaluates a model's ability to perform comprehensive hierarchical document structure analysis, including detecting page objects, predicting reading order across multiple groups, extracting tables of contents, and reconstructing the overall document hierarchy. Use when the user wants to benchmark on Comp-HRDoc, PubLayNet, DocLayNet, HRDoc, or asks about evaluating this task. Reports segmentation-based mAP.
Evaluates Open Information Extraction systems on their ability to extract compact, clause-level facts from text. It measures precision, recall, and F1 using token-level matching against gold triples, with a focus on avoiding over-specific extractions and handling overlapping constituents. Use when the user wants to benchmark on CaRB, Wire57, BenchIE, or asks about evaluating this task. Reports F1.
Evaluates a fairness-aware ensemble learning framework on recidivism risk prediction. It probes the model's ability to balance predictive accuracy against multiple group fairness constraints across racial demographics in a counterfactual causal setting. Use when the user wants to benchmark on COMPAS, or asks about evaluating this task. Reports MSE.
Evaluates Multimodal Large Language Models' ability to comprehend composite images (charts, collages, tables, code) and natural images, covering text recognition, visual reasoning, and conversational capabilities. Use when the user wants to benchmark on SEEDBench*, TextVQA, MMBench, MME, LLaVABench, ChartQA, DocVQA, InfoVQA, WebSRC, MathVista, OCRBench, or asks about evaluating this task. Reports Average score.
Evaluates the ability of large language models to generate correct, executable Python solutions for competitive programming problems. It probes algorithmic reasoning, code synthesis, and adherence to problem constraints under strict time and complexity limits. Use when the user wants to benchmark on LiveCodeBench, CodeContests, or asks about evaluating this task. Reports pass@1.
Evaluates an agent's ability to play competitive Pokémon Singles under partial observability and long-horizon uncertainty. It measures strategic decision-making, team building, and adaptation against heuristic opponents, search-based engines, LLM agents, and human players on a ranked ladder. Use when the user wants to benchmark on Competitive Pokémon Singles (CPS) on Pokémon Showdown, or asks about evaluating this task. Reports win rate.
Compute the completeness_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute completeness_score, or asks how to score with completeness_score.
Compute the CompletenessScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CompletenessScore, or asks how to score with CompletenessScore.
Compute the ComplexScaleInvariantSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ComplexScaleInvariantSignalNoiseRatio, or asks how to score with ComplexScaleInvariantSignalNoiseRatio.
Probes a model's ability to perform heterogeneous question answering by integrating information from multiple sources (knowledge bases, text, tables, infoboxes) across diverse domains and complex question intents. It specifically tests whether systems can fuse complementary structured and unstructured data to answer self-contained, human-generated questions. Use when the user wants to benchmark on CompMix, or asks about evaluating this task. Reports answer exact match.
Evaluates cross-lingual safety degradation in LLMs by measuring how well models refuse harmful prompts and avoid generating unsafe content across English and five Indic languages. Use when the user wants to benchmark on CompositeHarm, or asks about evaluating this task. Reports Refusal Rate (RR), Attack Success Rate (ASR).
This benchmark evaluates systematic generalization in abstract spatial reasoning by testing whether models can infer and compose geometric transformations (e.g., translation, rotation, reflection) from limited few-shot examples. It specifically probes out-of-distribution compositionality by training on known transformation primitives and level-1 compositions, then testing on novel level-2 compositions. Use when the user wants to benchmark on Compositional-ARC, or asks about evaluating this ta...
Evaluates the ability of on-device LLMs to perform two distinct tasks simultaneously in a single forward pass (compositional multi-tasking), such as summarization combined with translation or tone adjustment, while maintaining strict efficiency constraints. Use when the user wants to benchmark on Compositional Multi-tasking Benchmark, or asks about evaluating this task. Reports LLM judge (LLM-J).
This benchmark evaluates hardware-software co-design trade-offs for compound AI applications by measuring end-to-end latency, energy consumption, and accuracy across multi-modal workflows like video QA, evolutionary code generation, and RAG. It probes how different hardware configurations and software optimizations impact system performance under varying latency targets and workload patterns. Use when the user wants to benchmark on Google FRAMES benchmark, or asks about evaluating this task. ...
Evaluates the comprehensiveness and fine-grained accuracy of detailed image captions generated by vision-language models. It probes object detection, attribute binding, directional relationship modeling, and perception of tiny objects through hierarchical scene graph alignment and dedicated VQA tasks. Use when the user wants to benchmark on CompreCap, or asks about evaluating this task. Reports S_unified.
Evaluates the trade-offs between model accuracy and resource efficiency when applying various DNN compression techniques on mobile hardware. It measures how different compression methods affect inference speed, energy consumption, and storage footprint across standard vision and audio datasets. Use when the user wants to benchmark on CIFAR-10, MNIST, CIFAR-100, ImageNet, UbiSound, Har, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of transformer-based architectures to model long-range dependencies efficiently by compressing past hidden states into a fixed-size memory. It probes sequence modeling capabilities across text, audio, and visual domains, measuring how well compressed representations preserve salient information for next-token prediction and task completion. Use when the user wants to benchmark on Enwiki8, WikiText-103, PG-19, DMLab-30 (rooms_select_nonmatching_object), or asks about eval...
Evaluates the compute-optimal fine-tuning recipe for repurposing decoder-only LLMs into text embedding models. It measures how different computational budgets and fine-tuning methods affect both training contrastive loss and downstream retrieval/similarity performance. Use when the user wants to benchmark on BAAI BGE, MTEB, or asks about evaluating this task. Reports contrastive loss.
Evaluates computer-use agents on desktop task completion in online and offline settings, and measures the precision of a video-to-action module in detecting GUI events and extracting interaction parameters from screen recordings. Use when the user wants to benchmark on OSWorld-Verified, AgentNetBench, Video2Action Held-out Test Set, or asks about evaluating this task. Reports task success rate, step success rate.
Evaluates the robustness of open-set anomaly segmentation models under complex, real-world driving conditions. It probes a model's ability to detect out-of-distribution objects across diverse landforms and adverse weather while correctly ignoring non-driving-area elements and void regions. Use when the user wants to benchmark on ComsAmy, or asks about evaluating this task. Reports AuPRC.
Evaluates large vision-language models on chain-of-thought reasoning that requires generating both textual explanations and intermediate or final images. It probes the model's ability to perform four specific visual operations (creation, deletion, update, and selection) and align its multi-modal reasoning steps with ideal visual states. Use when the user wants to benchmark on CoMT, or asks about evaluating this task. Reports F1 score.
Evaluates the ability of neural models to predict human-assigned translation quality scores for Indian language pairs, measuring alignment with crowd-sourced DA+SQM ratings. Use when the user has predictions and gold and needs to compute Pearson correlation, Spearman correlation.
Evaluates the ability of machine learning and flow-based intrusion detection systems to accurately classify network traffic flows as benign or malicious. It probes the model's capacity to generalize across real-world benchmarks and synthetically generated, automatically labeled traffic for multi-step attack scenarios. Use when the user wants to benchmark on CICIDS17, ConCap ssh-patator, or asks about evaluating this task. Reports tpr.
Evaluates the compositional generalization capability of text-to-image models by testing their ability to generate images that satisfy multiple, simultaneously specified visual concepts (objects, colors, shapes, spatial relationships, etc.) within a single prompt. The benchmark probes model robustness to increasing compositional complexity (k) and reveals limitations in handling less frequent concept combinations. Use when the user wants to benchmark on ConceptMix, or asks about evaluating th...
Probes a model's ability to generate syntactically valid Java member functions from natural language documentation, conditioned on a full class environment including variable types, method signatures, and their interdependencies. It evaluates context-aware code generation, identifier disambiguation, and code reusability. Use when the user wants to benchmark on CONCODE, or asks about evaluating this task. Reports Exact match accuracy.
Compute the ConcordanceCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ConcordanceCorrCoef, or asks how to score with ConcordanceCorrCoef.
Measures how consistently a modeling approach's performance ranking holds across different question answering benchmarks. It probes whether improvements in QA models generalize across datasets with varying data collection procedures, passage/question distributions, and targeted linguistic phenomena. Use when the user wants to benchmark on SQuAD, NewsQA, NaturalQuestions, DROP, HotpotQA, QAMR, or asks about evaluating this task. Reports concurrence (Spearman's τ).
Evaluates a conditional latent diffusion framework's ability to synthesize task-specific LoRA parameters for NLP and image style-transfer tasks. It probes whether generated parameters can match or exceed standard fine-tuning and model-averaging baselines across diverse domains. Use when the user wants to benchmark on GLUE benchmark, SemArt, WikiArt, or asks about evaluating this task. Reports Average accuracy.
Evaluates in-game toxicity detection using a dual-level NLU framework that jointly predicts utterance-level toxicity intent and token-level semantic slots. It probes a model's ability to understand contextual, game-specific language and distinguish between explicit, implicit, and action-based toxicity. Use when the user wants to benchmark on CONDA, or asks about evaluating this task. Reports UCA.
Evaluates a conditional unigram tokenizer's cross-lingual alignment quality and its impact on downstream machine translation and language modeling tasks. It measures intrinsic tokenization properties, alignment accuracy, and task-specific performance metrics. Use when the user wants to benchmark on NLLB, MultiParaCrawl, WMT2020, Flores, WMT2020 test set, or asks about evaluating this task. Reports chrF++.
Evaluates a model's ability to perform conditional multi-hop reasoning in biomedical question answering, specifically how well it modulates clinical answers based on patient-specific constraints like comorbidities, contraindications, and special population factors. Use when the user wants to benchmark on CondMedQA, or asks about evaluating this task. Reports performance.
Evaluates continual learning methods for facial expression recognition under incremental, non-i.i.d. data settings. It probes a model's ability to learn new expressions sequentially while preserving prior knowledge, measuring both forward adaptation and backward forgetting. Use when the user wants to benchmark on CK+ (Extended Cohn-Kanade), or asks about evaluating this task. Reports Average Accuracy Score.
Evaluates the statistical validity (False Discovery Rate control) and detection sensitivity (statistical power) of cross-conformal anomaly detection methods against split-conformal baselines across datasets of varying sizes and dimensionalities. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports False Discovery Rate (FDR).
This evaluation protocol assesses the ability of 3D medical image segmentation models to control false negative rates under user-specified risk constraints while maintaining spatial precision. It benchmarks a model-agnostic conformal prediction calibration method against fixed heuristic thresholds across multiple anatomical datasets. Use when the user wants to benchmark on KiTS21, LiTS, NIH-LN ABD, LIDC-IDRI, MDSC-Colon, MDSC-Pancreas, or asks about evaluating this task. Reports ECR.