
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the ability of AI-generated video detectors to distinguish between real and synthetically generated videos across diverse generation models, tasks (T2V, I2V, V2V), and temporal/spatial artifacts. Use when the user wants to benchmark on AIGVDBench, or asks about evaluating this task. Reports accuracy.
Compute AIML-TUDA/IsomorphicPerturbationTesting via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AIML-TUDA/IsomorphicPerturbationTesting.
Compute AIML-TUDA/VerifiableRewardsForScalableLogicalReasoning via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AIML-TUDA/VerifiableRewardsForScalableLogicalReasoning.
Evaluates AI inference performance across diverse image classification model architectures on mobile and embedded devices. It measures the trade-off between inference speed and computational efficiency to compare models, frameworks, and hardware. Use when the user wants to benchmark on ImageNet 2012, or asks about evaluating this task. Reports VIPS.
Evaluates the end-to-end performance and weak scalability of heterogeneous AI-HPC systems using AutoML workloads. It measures how efficiently clusters execute dynamically scaling machine learning training and inference tasks across varying numbers of nodes. Use when the user wants to benchmark on CIFAR10, or asks about evaluating this task. Reports cumulative OPS.
Evaluates Large Audio-Language Models on foundational audio comprehension across speech, natural sounds, and music, as well as open-ended instruction-following via generative responses. It probes the model's ability to understand mixed audio, follow complex prompts, and produce accurate, contextually relevant text. Use when the user wants to benchmark on AIR-Bench, or asks about evaluating this task. Reports GPT-4 alignment strategy.
Evaluates a regression model's ability to forecast hyper-local air pollutant concentrations using fine-grained traffic intensity descriptors. It probes how well traffic patterns across different spatial rings and colors correlate with specific pollutant levels under varying training station configurations. Use when the user wants to benchmark on Mexico City Traffic & Pollution Dataset, or asks about evaluating this task. Reports RMSE.
Evaluates AI research agents across the full scientific lifecycle, including idea generation, experiment design, and iterative refinement. Agents must generate and execute code to train models on specified datasets without baseline code, testing reasoning, generalization, and solution exploration capabilities. Use when the user wants to benchmark on AIRS-Bench, or asks about evaluating this task. Reports average normalized score.
Evaluates a generative world model's ability to predict first-person future video observations under specified 6DoF aerial motion intentions. It probes spatio-temporal consistency, motion alignment, and counterfactual reasoning in 3D aerial environments. Use when the user wants to benchmark on AirScape Dataset, or asks about evaluating this task. Reports IAR.
Evaluates the ability of a multi-speaker TTS system to synthesize high-fidelity Mandarin speech that preserves speaker identity across both seen and unseen speakers. It probes zero-shot voice cloning capability and generalization to novel speakers using objective speaker verification metrics. Use when the user wants to benchmark on AISHELL-3, or asks about evaluating this task. Reports SV-EER.
Evaluates the quality, diversity, and music-motion synchronization of generated 3D dance sequences. It probes a model's ability to synthesize physically plausible, choreographically diverse, and rhythm-aligned human motion from audio input. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports FID_k.
Evaluates an agent's ability to infer and execute multi-step visual actions on Android devices from natural language instructions. It specifically probes Out-of-Distribution generalization across unseen Android OS versions, instruction language patterns (subjects/verbs), and app/web domains. Use when the user wants to benchmark on AITW, or asks about evaluating this task. Reports average score.
Proposes a diagnostic metric to quantify the discrepancy between macro-averaged benchmark scores and experienced reliability in deployment. It accounts for uneven task coverage by measuring the dispersion of error rates across domains, highlighting how gap-uniform evaluation can misstate actual user experience. Use when the user has predictions and gold and needs to compute CV_d(e_d).
Compute akki2825/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of akki2825/accents_unplugged_eval.
Evaluates voice conversion systems on speech quality and speaker similarity using subjective human ratings. It probes the ability of models to convert speech between speakers (specifically inter-gender) while preserving linguistic content and target speaker identity. Use when the user wants to benchmark on ALAGIN Japanese Speech Database Set B, or asks about evaluating this task. Reports Mean Opinion Score (MOS) for Speech Quality.
Compute albertgong1/my_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of albertgong1/my_metric.
Evaluates vision-language models' ability to actively navigate long, visually rich documents to gather evidence and answer complex queries. It probes multi-turn reasoning, retrieval accuracy, and the effectiveness of direct page-index access versus semantic search. Use when the user wants to benchmark on MMLongBench, LongDocURL, PaperTab, PaperText, FetaTab, DUDE-sub, or asks about evaluating this task. Reports GPT-4o–judged answer accuracy (Acc).
This benchmark evaluates reinforcement learning agents across 60 Atari 2600 games using a stochastic environment variant with sticky actions. It measures sample efficiency and learning stability by tracking performance at multiple frame-count thresholds (10M to 200M). Use when the user wants to benchmark on Arcade Learning Environment (ALE), or asks about evaluating this task. Reports score averages.
Evaluates pre-trained Hebrew language models on core NLP tasks including morphological analysis, named entity recognition, and sentiment analysis. It measures how well the models handle Hebrew-specific linguistic features and resource-scarce language challenges compared to existing baselines. Use when the user wants to benchmark on SPMRL Hebrew Section, UD treebanks Hebrew Section, Ben-Mordecai and Elhadad corpus, NEMO corpus, Amram et al. (2018) corpus (cleaned), or asks about evaluating thi...
Embodied instruction following in a simulated household environment, requiring an agent to execute long-horizon navigation and object manipulation tasks based on natural language commands. Use when the user wants to benchmark on ALFRED, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates cross-lingual and cross-script transfer performance for sentiment analysis and topic classification on a novel multi-layer Algerian dialect corpus. Probes how script differences (Latin/NArabizi vs. Arabic/Persian/Urdu) and typological similarity impact classification accuracy in code-switched, under-resourced vernaculars. Use when the user wants to benchmark on Algerian Dialect Corpus (NArabizi), or asks about evaluating this task. Reports Macro F1.
This benchmark evaluates a model's ability to predict human visual brain activity during object recognition. It compares model representations against fMRI and MEG neural recordings using representational similarity analysis (RSA) across spatial (EVC vs IT) and temporal (early vs late processing) dimensions. Use when the user wants to benchmark on Algonauts 2019 Challenge, or asks about evaluating this task. Reports noise-normalized variance explained.
Compute AlhitawiMohammed22/CER_Hu-Evaluation-Metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AlhitawiMohammed22/CER_Hu-Evaluation-Metrics.
Evaluates a vision-language model's cross-modal retrieval capabilities (matching images to text and vice versa) and zero-shot image classification performance without task-specific fine-tuning. It also measures transfer learning effectiveness on downstream visual benchmarks via linear probing and full fine-tuning. Use when the user wants to benchmark on Flickr30K, MSCOCO, or asks about evaluating this task. Reports R@10.
Evaluates a logistic regression model on SPECTER embeddings to distinguish AI alignment research articles from adjacent research on arXiv. Probes the model's ability to capture domain-specific semantic patterns and citation-driven textual features for automated literature filtering. Use when the user wants to benchmark on arXiv Alignment Research Corpus, or asks about evaluating this task. Reports AUC.
Evaluates information retrieval capabilities in an educational context by testing a model's ability to retrieve relevant reference pages or similar past questions given a student's query. It probes handling of noisy text (spelling/grammar errors), multimodal inputs (images, formulas), and grade-aware language complexity. Use when the user wants to benchmark on Alloprof, or asks about evaluating this task. Reports nDCG.
Evaluates a Wang-Landau sampling method combined with cluster expansion for predicting thermodynamic phase diagrams of binary alloys. It probes the method's ability to capture ordering and phase-separation tendencies, and accurately reproduce experimental phase boundaries and transition temperatures. Use when the user wants to benchmark on Cu-Au alloy, Pd-Rh alloy, or asks about evaluating this task. Reports cross-validation score.
Evaluates multimodal large language models' ability to comprehend hour-long videos across nine distinct tasks, including classification, recognition, localization, captioning, emotion recognition, and needle-in-a-haystack retrieval. It specifically probes temporal reasoning, detail extraction, and long-context retention over extended video durations. Use when the user wants to benchmark on ALLVB, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the cultural and linguistic reasoning capabilities of large multimodal models across 100 languages. It probes visual understanding and cultural knowledge through generic and culturally specific domains, testing both closed-form and open-ended question answering. Use when the user wants to benchmark on ALM-bench, or asks about evaluating this task. Reports accuracy.
Evaluates whether language model explanations (e.g., weights, qualitative descriptions) enable a second predictor model to accurately simulate and predict the behavior of a synthetic linear model across safety-relevant scenarios. The benchmark specifically probes simulatability and robustness to distributional shift between training and test variable values. Use when the user wants to benchmark on ALMANACS Synthetic Dataset, or asks about evaluating this task. Reports probability.
This evaluation probes a robot policy's ability to perform precise, temporally extended bimanual manipulation tasks using only visual and proprioceptive inputs. It specifically tests the model's robustness to compounding errors, non-Markovian dynamics, and perception challenges like transparent or low-contrast objects. Use when the user wants to benchmark on ALOHA Fine Manipulation Tasks, or asks about evaluating this task. Reports success rate.
This evaluation probes the ability of LLM-based frameworks to estimate translation quality in a reference-free setting by predicting continuous quality scores for source-target sentence pairs across multiple low-resource language directions. It specifically tests how intermediate Transformer layer representations and adaptive regression heads improve cross-lingual alignment and quality prediction compared to standard fine-tuning or zero-shot prompting. Use when the user wants to benchmark on ...
Evaluates the effectiveness of dynamic low-rank adaptation (LoRA) for fine-tuning large language models across classification, question answering, and instruction generation tasks. It measures how well rank allocation strategies preserve performance while maintaining or reducing tunable parameter counts. Use when the user wants to benchmark on SQuAD, BoolQ, COPA, ReCoRD, SST-2, RTE, QNLI, Alpaca, MT-Bench, E2E, or asks about evaluating this task. Reports accuracy, GPT-4 score.
Evaluates the alignment quality of language models by measuring their win rate against a baseline on the AlpacaEval benchmark. It specifically uses length-controlled (LC) win rates to mitigate the known bias toward longer model outputs in standard auto-annotator evaluations. Use when the user wants to benchmark on alpaca_eval, or asks about evaluating this task. Reports AlpacaEval length-controlled (LC) win rate.
Evaluates LLM response quality via pairwise win rates against a baseline, while specifically probing the metric's susceptibility to length bias, gameability via verbosity prompting, and robustness to adversarial truncation. Use when the user wants to benchmark on AlpacaEval, or asks about evaluating this task. Reports Win rate.
Evaluates active learning pipelines by comparing query strategies paired with tabular classifiers across multiple datasets. It measures how efficiently pipelines improve test performance as the labeled data budget increases, highlighting the interplay between learner choice and query strategy. Use when the user wants to benchmark on OpenML-CC18 and TabZilla Benchmark Suite, or asks about evaluating this task. Reports AUBC (Area Under the Budget Curve).
Evaluates a model's ability to generate correct, executable code for competitive programming problems under strict submission limits. It probes algorithmic reasoning, code synthesis, and the capacity to pass hidden test cases after filtering on provided examples. Use when the user wants to benchmark on CodeContests, or asks about evaluating this task. Reports solve rate.
Evaluates an LLM-based autonomous agent's ability to discover novel algorithms through iterative idea generation, code modification, and execution-based verification. It probes the model's capacity for scientific reasoning, program synthesis, and optimization under simulated peer-review feedback. Use when the user wants to benchmark on AlphaResearch Algorithm Discovery Problems, or asks about evaluating this task. Reports win_rate (excel@best > 0).
AlpsBench evaluates the full lifecycle of LLM personalization, including extracting structured memories from dialogue, dynamically updating them, retrieving relevant memories under distractors, and utilizing them to generate aligned responses across dimensions like persona awareness, preference following, and emotional intelligence. Use when the user wants to benchmark on AlpsBench, or asks about evaluating this task. Reports F1 score (exact match).
Evaluates an agentic LLM's ability to plan and execute multistep robotic manipulation tasks in simulation. It tests closed-loop reasoning via ReAct-style loops, comparing code-as-policy and tool-as-policy execution modes across linguistically diverse tasks. Use when the user wants to benchmark on ALRM Simulation Benchmark, or asks about evaluating this task. Reports task_completion.
Compute alvinasvk/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of alvinasvk/accents_unplugged_eval.
Evaluates multi-class classification performance on Alzheimer's disease MRI scans to assess a model's ability to distinguish between different stages of dementia and healthy controls under resource-constrained hardware conditions. Use when the user wants to benchmark on Alzheimer MRI 4 Classes Dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to continuously learn multiple text classification tasks without catastrophic forgetting, measuring how well it retains knowledge of previous tasks while adapting to new ones. Use when the user wants to benchmark on Standard CL benchmarks, Large number of tasks benchmark, or asks about evaluating this task. Reports average results.
Evaluates session-based recommendation and text generation capabilities across multiple languages and locales. It probes a model's ability to predict the next product in a shopping session, transfer knowledge across domain-shifted locales, and generate product titles from session context. Use when the user wants to benchmark on Amazon-M2, or asks about evaluating this task. Reports next-product prediction.
Evaluates the ability of abstractive summarization models to generate concise, informative summaries of product reviews while preserving aspect and opinion details. It measures lexical overlap with human-written reference summaries. Use when the user wants to benchmark on Amazon Reviews (Healthcare & Electronics), or asks about evaluating this task. Reports ROUGE-1.
Evaluates the ability of neural retriever-reranker pipelines to accurately retrieve relevant product entities from semi-structured e-commerce knowledge graphs using natural language queries. It probes semantic matching, cross-encoder reranking effectiveness, and the impact of graph-based augmentation on retrieval precision and recall. Use when the user wants to benchmark on Amazon STaRK SKB, or asks about evaluating this task. Reports Hit@1.
This benchmark evaluates a model's ability to identify ambiguous open-domain questions, generate multiple plausible answer spans, and produce disambiguated question rewrites that distinguish between different interpretations of the same query. Use when the user wants to benchmark on AMBIGNQ, or asks about evaluating this task. Reports F1ans.
Evaluates audio-language models' ability to recognize ambiguous emotions in speech by predicting full emotion probability distributions and dominant class labels. It specifically probes how test-time scaling (TTS) strategies and model capacity interact with varying levels of emotional ambiguity to improve or degrade recognition performance. Use when the user wants to benchmark on IEMOCAP, MSP-Podcast, CREMA-D, or asks about evaluating this task. Reports JS divergence.
Evaluates a model's ability to generate diverse, semantically valid SQL queries when a natural language question is ambiguous. It specifically probes whether the model can cover multiple valid interpretations of the same query within a limited set of top-k outputs. Use when the user wants to benchmark on AmbiQT, SPIDER, Kaggle DBQA, or asks about evaluating this task. Reports BothInTopK.
Evaluates a Text-to-SQL system's ability to generate correct SQL from ambiguous natural language queries when integrated with an interactive ambiguity resolution module. It also measures the system's precision, recall, and F1 in detecting and classifying specific types of schema-mapping and reasoning ambiguities. Use when the user wants to benchmark on AmbiSQL Constructed Dataset, or asks about evaluating this task. Reports Exact Match accuracy.