
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the ability of speech quality reward models to correctly rank pairs of audio samples based on their Mean Opinion Score (MOS). It probes fine-grained perceptual discrimination and cross-dataset generalization in preference-based audio modeling. Use when the user wants to benchmark on BVCC, NISQA, SingMOS, SOMOS, TMHINT-QI, VMC’23, or asks about evaluating this task. Reports accuracy.
Evaluates the naturalness, speaker similarity, and real-time synthesis speed of a Mandarin speech cloning system across diverse practical application scenarios. Use when the user has predictions and gold and needs to compute MOS (naturalness & similarity).
Evaluates a peer-to-peer wireless power transfer (P2P-WPT) framework's ability to balance energy across a mobile crowd while minimizing energy loss and maximizing network energy retention. It probes how well mobility and social-aware peer selection algorithms perform in dynamic, simulated environments. Use when the user wants to benchmark on MoSaBa Simulation Scenario, or asks about evaluating this task. Reports Total network energy.
Evaluates a model's ability to perform next-item recommendation in cross-domain sequential settings by decomposing user intent into orthogonal preference components. It probes how well the model leverages shared and domain-specific signals across multiple item categories to predict future interactions. Use when the user wants to benchmark on Amazon Reviews (Movie–Book), Amazon Reviews (Movie–Music), Douban (Movie–Book), or asks about evaluating this task. Reports NDCG@10.
Probes Visually Rich Document Understanding (VRDU) capabilities across three tasks: document VQA, page-level OCR, and reading order prediction. It specifically tests models' ability to comprehend dense, bilingual, non-Manhattan layouts, perform multi-span reasoning, and maintain global layout coherence without token reduction artifacts. Use when the user wants to benchmark on MosaicDoc, or asks about evaluating this task. Reports ANLSL.
Evaluates downstream language model capabilities across 33 question-answering tasks. It measures how effectively data pruning strategies improve general performance compared to unpruned baselines, using a normalized accuracy metric that accounts for random guessing baselines. Use when the user wants to benchmark on MosaicML evaluation gauntlet, or asks about evaluating this task. Reports average normalized accuracy.
Evaluates text-to-image generation models on their ability to produce culturally diverse, demographically accurate images with correct landmark representation. It probes alignment, visual quality, aesthetic appeal, cross-cultural knowledge, and demographic fairness across multiple languages and intersectional attributes. Use when the user wants to benchmark on MosAIG Dataset, or asks about evaluating this task. Reports CLIPScore.
Evaluates automatic speech recognition (ASR) performance on low-resource Maltese speech data. It measures the accuracy of a sequence-to-sequence model in transcribing audio into text after training on filtered open-source speech corpora. Use when the user wants to benchmark on VoxPopuli (Maltese subset), or asks about evaluating this task. Reports Word Error Rate (WER).
This benchmark evaluates video object segmentation and tracking models under highly complex, unconstrained real-world conditions. It specifically probes robustness to severe occlusions, frequent object disappearance and reappearance, adverse weather, low-light environments, camouflage, and knowledge-dependent scenarios. Use when the user wants to benchmark on MOSEv2, or asks about evaluating this task. Reports $\mathcal{J}\\&\dot{\mathcal{F}}$.
Evaluates real-time multiclass object detection capabilities for identifying individual mosquitoes, mosquito swarms, and breeding sites in natural environments. It measures how well a model can localize and classify these distinct biological and environmental targets under varying real-world conditions. Use when the user wants to benchmark on MosquitoFusion, or asks about evaluating this task. Reports mAP@50.
Evaluates multi-object tracking accuracy in a tracking-by-detection pipeline on edge hardware, measuring how well backbones support SORT-based tracking under strict computational and power constraints. Use when the user wants to benchmark on MOT15, or asks about evaluating this task. Reports MOTA.
Evaluates multi-object tracking algorithms on video sequences by measuring detection accuracy, identity consistency, and localization precision. It assesses how well trackers maintain object identities over time while correctly handling occlusions, distractors, and varying crowd densities. Use when the user wants to benchmark on MOT16, or asks about evaluating this task. Reports MOTA.
This benchmark evaluates the accuracy of malware family classification models and antivirus-based labeling tools on a large, expert-verified dataset. It probes a model's ability to correctly assign ground-truth family labels to malware samples, including handling open-set noise and alias resolution. Use when the user wants to benchmark on MOTIF, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal time series forecasting models across two scenarios: varying-history forecasting (using long and short temporal sequences) and cold-start forecasting (predicting from minimal initial observations). It probes how well models leverage static external modalities like text and metadata to improve prediction accuracy, especially for sparse or short series. Use when the user wants to benchmark on PixelRec, AmazonReview, WikiPeople, Movielens, TaobaoFashion, Tianchi, News, or as...
Evaluates a model's ability to predict future trajectories of pedestrians and other agents in crowded urban environments. It probes how well the architecture captures inter-agent dynamics and interaction patterns over short temporal windows to estimate safe crossing paths. Use when the user wants to benchmark on L-CAS, ETH-Hotel, UCY-Uni, ETH-Univ, Zara01, Zara02, or asks about evaluating this task. Reports Average Displacement Error (ADE).
Evaluates a model's ability to predict human-likeness scores for humanoid and human motion sequences based purely on kinematic data. It probes whether models can align with human perceptual judgments of motion fluency, coordination, and naturalness without relying on visual appearance cues. Use when the user wants to benchmark on HHMotion, or asks about evaluating this task. Reports Spearman's ρ.
Evaluates the effectiveness of the MotionBank dataset for downstream text-to-motion generation tasks. It measures how well rule-based, disentangled motion annotations improve single human motion synthesis and human-object interaction generation compared to baseline models. Use when the user wants to benchmark on MotionBank, HumanML3D, BEHAVE, or asks about evaluating this task. Reports R Precision.
Evaluates a model's ability to perform multi-organ and tumor segmentation on partially labeled 3D medical images. It probes the network's capacity to learn from incomplete annotations and generalize across diverse anatomical structures using a unified architecture. Use when the user wants to benchmark on MOTS (Multi-Organ and Tumor Segmentation), BCV (MICCAI 2015 Multi Atlas Labeling Beyond the Cranial Vault), BraTS (2018 Brain Tumor Segmentation Challenge), or asks about evaluating this task...
Evaluates an imitation learning agent's ability to reproduce mouse forelimb reaching kinematics and muscle activation patterns using a musculoskeletal physics model. It probes how well learned policies can match biological motion and electrophysiological signals under varying physics-aware constraints. Use when the user wants to benchmark on Mouse forelimb reaching mocap dataset, or asks about evaluating this task. Reports track replay error.
Evaluates multimodal understanding and reasoning across visual question answering, OCR, region-level VQA, and visual conversation tasks. It probes how effectively poly-visual expert ensembles fuse information from multiple encoders compared to single-expert baselines. Use when the user wants to benchmark on LLaVA-1.5 Benchmark Suite, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models' capabilities across general visual question answering, text-oriented VQA (charts, documents, diagrams), visual grounding (referring expression comprehension), and specialized medical VQA. It also assesses general multimodal reasoning and hallucination resistance. Use when the user wants to benchmark on MME, MMBench, MMBench-CN, QBench, MathVista, MathVerse, POPE, VQAv2, GQA, SQA-I, TextVQA, ChartQA, DocVQA, AI2D, RefCOCO, RefCOCO+, RefCOCOg, VQA-RAD...
Automated evaluation of text generation quality by computing semantic distance between system outputs and human references using contextualized embeddings and Earth Mover's Distance (EMD). It probes a model's ability to capture meaning-based similarity rather than surface-level n-gram overlaps across machine translation, summarization, dialogue, and image captioning tasks. Use when the user has predictions and gold and needs to compute Pearson r.
Evaluates monocular 3D human pose and trajectory estimation in global coordinates, specifically probing a model's ability to maintain physical plausibility (e.g., avoiding scene penetration, minimizing foot sliding) while tracking dynamic camera motion on complex, non-flat terrain. Use when the user wants to benchmark on MoviCam, or asks about evaluating this task. Reports MPJPE.
Predicts multi-label emotions and mental states for movie scenes and individual characters using multimodal inputs (video, dialog, character appearance). It probes long-form video understanding and the ability to integrate visual and linguistic cues for affect recognition. Use when the user wants to benchmark on MovieGraphs, or asks about evaluating this task. Reports mAP.
Evaluates the ability of abstractive summarization models to generate concise, coherent summaries of long, dispersed movie screenplay narratives. It probes long-document understanding, narrative coherence, and the model's capacity to synthesize information across thousands of tokens. Use when the user wants to benchmark on MovieSum, or asks about evaluating this task. Reports ROUGE F1 (1/2/L).
Evaluates a model's ability to generate and rank spatio-temporal proposals for moving objects in video. It probes motion-based segmentation quality, proposal coverage, and ranking accuracy on both rigid and non-rigid motion across diverse scenes. Use when the user wants to benchmark on VSB100, Moseg, or asks about evaluating this task. Reports Average best overlap, Coverage.
Evaluates a hardware-software co-design framework for recommendation systems that dynamically switches between different embedding representations (table, DHE, hybrid) across heterogeneous hardware (CPU, GPU, IPU) to optimize throughput of correct predictions and model accuracy under strict latency constraints. Use when the user wants to benchmark on Kaggle, Terabyte, or asks about evaluating this task. Reports Throughput of Correct Predictions.
Evaluates an alternating minimization model predictive control (MPC) framework for autonomous driving against a joint optimization baseline, focusing on computational efficiency, trajectory smoothness, and safety margins during critical maneuvers like overtaking, lane changes, and sudden braking. Use when the user wants to benchmark on CARSIM Simulation Benchmarks, or asks about evaluating this task. Reports iterations.
Evaluates the ability of Large Vision-Language Models (LVLMs) to suppress object and semantic hallucinations while maintaining general perception, reasoning, and generative capabilities. It probes grounding fidelity across structured yes/no queries, open-ended captioning, and fine-grained visual diagnostics. Use when the user wants to benchmark on MSCOCO, MME, LLaVA-Bench, HallusionBench, or asks about evaluating this task. Reports CHAIR_S, CHAIR_I, POPE F1.
Evaluates a modified 2D U-Net's ability to segment pediatric and adult brain tumors in MRI scans by leveraging multi-planar data augmentation to learn 3D volumetric representations. It probes the model's generalization across diverse tumor types, anatomical variations, and imaging scenarios. Use when the user wants to benchmark on Pediatrics Tumor Challenge (PED), Brain Metastasis Challenge (MET), Sub-Sahara-Africa Adult Glioma Challenge (SSA), or asks about evaluating this task. Reports Dice...
This benchmark evaluates whether vision-language models can generate scientifically grounded, figure-dependent questions rather than generic visual queries. It probes content-specific visual grounding by measuring how model outputs change when the correct figure is replaced, removed, or kept, alongside assessing the depth and diversity of the generated questions. Use when the user wants to benchmark on MQUD, or asks about evaluating this task. Reports rIG.
Evaluates large vision-language models' ability to leverage retrieved visual knowledge versus textual knowledge across perspective and transformative change scenarios. Probes robustness to noisy retrieved images and measures how effectively models utilize visually augmented information compared to human baselines. Use when the user wants to benchmark on MRAG-Bench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large language models' ability to comprehend passages and answer questions across multiple dimensions, including context understanding, external knowledge integration, and complex reasoning. It probes factual fidelity, counterfactual handling, commonsense, world knowledge, and multi-hop reasoning capabilities. Use when the user wants to benchmark on MRCEval, or asks about evaluating this task. Reports accuracy.
Evaluates reasoning-intensive multimodal retrieval across 23 expert domains using interleaved image-text queries and documents. It probes a model's ability to perform knowledge-based matching, theorem linking, and logical contradiction detection in complex, real-world scenarios. Use when the user wants to benchmark on MRMR, or asks about evaluating this task. Reports nDCG@10.
Evaluates out-of-domain generalization in extractive reading comprehension by testing models on held-out datasets from diverse domains (crowdsourced, synthetic, domain experts, Wikipedia, education, etc.) that were not seen during training. Use when the user wants to benchmark on MRQA 2019 Shared Task, or asks about evaluating this task. Reports F1.
Evaluates mono-lingual dense retrieval models across eleven typologically diverse languages by measuring their ability to rank relevant Wikipedia passages for given questions. It probes zero-shot cross-lingual generalization and the effectiveness of sparse-dense hybrid retrieval compared to strong sparse baselines. Use when the user wants to benchmark on Mr. TYDI v1.1, or asks about evaluating this task. Reports MRR@100.
Evaluates a model's ability to distinguish real images from AI-generated ones, and to identify the specific generative model that produced a synthetic image. It probes robustness to semantic alignment and fine-grained model attribution. Use when the user wants to benchmark on MS COCOAI, or asks about evaluating this task. Reports baseline_score.
Evaluates machine reading comprehension models on real-world search queries across multiple answer types (numeric, yes/no, descriptive) and tasks (answer generation, span extraction, passage ranking). Probes a model's ability to extract or generate accurate answers from noisy, multi-document web contexts and handle unanswerable questions. Use when the user wants to benchmark on MS MARCO, or asks about evaluating this task. Reports ROUGE-L.
Evaluates passage and document ranking systems on a large-scale, document-native corpus. It probes a model's ability to retrieve relevant content from millions of documents using sparse, crowd-sourced relevance judgments, while handling realistic corpus drift and query-independent passage extraction. Use when the user wants to benchmark on MS MARCO v2, or asks about evaluating this task. Reports NDCG@10.
Evaluates an LLM agent's ability to retrieve and utilize long-term memory across multiple dialogue sessions to complete goal-oriented tasks. It probes intent-aligned memory selection, slot-level tracking, and dialogue efficiency in maintaining task continuity over extended interactions. Use when the user wants to benchmark on MS-TOD, SGD, MultiWOZ 2.2, or asks about evaluating this task. Reports Success Rate (S.R.).
Evaluates audio deepfake detection models on their ability to distinguish real from synthetically generated multi-speaker conversations. It probes robustness to conversational dynamics, speech overlap, and varying acoustic conditions. Use when the user wants to benchmark on MsCADD, or asks about evaluating this task. Reports F1 score.
Evaluates image captioning quality by scoring generated captions against crowdworker references and expert THumB 1.0 scores. It measures how well automatic metrics correlate with human judgments across different captioning systems. Use when the user wants to benchmark on MSCOCO, or asks about evaluating this task. Reports RefCLIP-S.
Evaluates the adversarial robustness and cross-task generalization of 3D medical image segmentation models across diverse organs and tumors using CT and MRI modalities. It measures how well models maintain segmentation accuracy under sophisticated adversarial attacks compared to clean inference. Use when the user wants to benchmark on MSD, or asks about evaluating this task. Reports Dice score.
Evaluates the segmentation accuracy of a 3D U-Net model on brain tumor MRI volumes, and measures the computational efficiency of distributed hyperparameter tuning across multiple GPUs. Use when the user wants to benchmark on MSD Task 1, or asks about evaluating this task. Reports Dice score.
Evaluates audio embedding models and ASR pipelines across eight core auditory tasks to measure real-world generalization, compression robustness, and the performance gap between direct audio processing and text-based oracles. It quantifies how well embeddings capture semantic and acoustic information without task-specific fine-tuning. Use when the user wants to benchmark on SVQ (Simple Voice Questions), Speech-MASSIVE, FSD50K, BirdSet, or asks about evaluating this task. Reports MRR.
Evaluates a model's ability to predict the immediate next navigation step in a 3D scene given a multi-modal situation description and a textual goal. Use when the user wants to benchmark on MSNN, or asks about evaluating this task. Reports Next-step Action Accuracy.
Evaluates graduate-level materials science reasoning and factual knowledge through long-form explanatory answers and binary true/false questions. It probes model capabilities in domain-specific knowledge retrieval, multi-step scientific reasoning, and accuracy under both direct prompting and retrieval-augmented generation (RAG) settings. Use when the user wants to benchmark on MSQA, or asks about evaluating this task. Reports accuracy.
Evaluates whether fine-tuning vision-language models on policy-grounded safety reasoning improves their ability to refuse unsafe multimodal prompts while preserving general multimodal reasoning capabilities. Use when the user wants to benchmark on BeaverTails-V, MM-SafetyBench, SPA-VL Eval, MME-CoT, MM-Vet, or asks about evaluating this task. Reports safety rate.
Evaluates the safety and hazard response capabilities of vision-language models (VLMs) by testing how they handle prompts that combine text and images to elicit unsafe or hazardous outputs. Use when the user wants to benchmark on MSTS, or asks about evaluating this task. Reports unsafe_response_rate.
Evaluates large language and vision-language models' ability to comprehend complete musical scores across four hierarchical levels of reasoning. It probes bar localization, structural understanding, and complex musical reasoning, while highlighting modality gaps between textual (ABC notation) and visual (PDF/image) inputs. Use when the user wants to benchmark on MSU-Bench, or asks about evaluating this task. Reports Accuracy.