
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates large vision-language models on high-resolution remote sensing imagery by testing their ability to answer questions about color, count, position, and open-ended descriptions. It probes the model's perception and reasoning capabilities on complex, large-scale satellite/aerial images. Use when the user wants to benchmark on MME-RealWorld-RS, LRS-VQA, or asks about evaluating this task. Reports accuracy.
This evaluation probes the ability of a generative error correction model to refine audio-visual speech recognition transcripts under varying noise conditions. It measures how effectively multimodal cues (lip video and audio) combined with N-best hypotheses can reduce transcription errors compared to baseline systems. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports WER.
Evaluates sequential recommendation models on next-item prediction tasks across various domains (movies, products, games) with varying sequence lengths and sparsity. Use when the user wants to benchmark on ML-1M, Amazon-Beauty, Amazon-Video, Amazon-Sports, Steam, XLong, or asks about evaluating this task. Reports Recall@10.
Evaluates models' ability to detect and rank lexical semantic change in Spanish diachronic corpora. It probes both graded ranking of semantic shift magnitude and binary classification of sense gain/loss or change presence. Use when the user wants to benchmark on LSCDiscovery, or asks about evaluating this task. Reports Spearman rank correlation (SPR), F1 score.
Evaluates supervised open information extraction (OIE) models on extracting schema-free predicate-argument tuples from sentences. Probes the model's ability to correctly identify predicates and arguments while maintaining syntactic head alignment and argument ordering. Use when the user wants to benchmark on LSOIE, or asks about evaluating this task. Reports F1.
Evaluates narrative understanding, commonsense reasoning, and topic coherence in speech-text models by selecting the most plausible continuation from multiple candidates. The benchmark tests both speech-to-speech and text-to-text modes to assess cross-modal alignment and reasoning capabilities under compute constraints. Use when the user wants to benchmark on HellaSwag (sHellaSWAG), StoryCloze, TopicStoryCloze, or asks about evaluating this task. Reports accuracy.
Evaluates multivariate time series forecasting models on their ability to capture short-term local dependencies and long-term periodic patterns across diverse real-world datasets with varying temporal scales and frequencies. Use when the user wants to benchmark on Traffic, Solar-Energy, Electricity, Exchange-Rate, or asks about evaluating this task. Reports RSE.
Evaluates a vision transformer's ability to perform multi-label classification on chest X-ray images. It probes the model's capacity to detect multiple pathologies simultaneously and model inter-label dependencies using learnable label tokens. Use when the user wants to benchmark on NIH-CXR14, CheXpert-5, CheXpert-13, or asks about evaluating this task. Reports AUC (%).
Evaluates the vision-language understanding and domain generalization capabilities of Mixture-of-Experts (MoE) models. It tests how well a long-tailed distribution-aware router preserves routing variance for vision tokens while maintaining load balancing for language tokens, impacting both accuracy and inference efficiency. Use when the user wants to benchmark on GQA, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MM-Vet, PACS, VLCS, Office-Home, DomainNet, or asks about evaluating this task. Re...
This evaluation probes the accuracy-latency-efficiency trade-off of streaming voice agents under realistic ASR conditions. It measures how well a system maintains reasoning quality while minimizing computational overhead and response delays when processing natural speech with disfluencies, misrecognitions, and non-uniform speaking rates. Use when the user wants to benchmark on VERA (AIME and GPQA-Diamond), Spoken-MQA, BigBenchAudio, Pause-and-Repair Benchmark, or asks about evaluating this ta...
Evaluates Luxembourgish (LTZ) language understanding across eight diverse NLU tasks, including classification, sequence labeling, and textual entailment. It probes encoder models and prompted LLMs on their ability to handle low-resource language nuances, structural complexity, and label sensitivity. Use when the user wants to benchmark on ltzGLUE, or asks about evaluating this task. Reports macro-F1.
Evaluates the computational efficiency and human-readability of two General Game Playing systems (Ludii and RBG) by measuring their playout throughput and the token count required to define game rules. Use when the user wants to benchmark on Unspecified game rule suite, or asks about evaluating this task. Reports playouts.
Evaluates deep learning models on multi-vendor full-field digital mammography for breast cancer diagnosis, BI-RADS classification, and breast density prediction. It specifically probes the model's robustness to domain shifts induced by different imaging vendors and X-ray energies. Use when the user wants to benchmark on LUMINA, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to generate physically plausible, temporally coherent HDR video from standard dynamic range (SDR) inputs. It probes reconstruction fidelity in perceptually uniform HDR spaces, temporal stability across frames, and the model's capacity to recover clipped radiance details using learned visual priors. Use when the user wants to benchmark on ARRI Cinema Footage, UPIQ, or asks about evaluating this task. Reports PU21-PSNR.
This evaluation protocol assesses the quality of the Lunara dataset by measuring visual aesthetic appeal, semantic alignment between images and prompts, cross-modal retrieval accuracy, and perceptual diversity across images. It provides a structured framework to verify that the dataset prioritizes high-quality, stylistically diverse, and semantically grounded image-text pairs over noisy web-scraped alternatives. Use when the user wants to benchmark on Lunara Aesthetic Dataset, or asks about e...
Evaluates the quality, contextual variation isolation, aesthetic appeal, and identity preservation of the Lunara Aesthetic II image variation dataset compared to other web-scale datasets. Use when the user wants to benchmark on Lunara-II-Variations, Lunara-I, CC3M, LAION-2B-Aesthetic, WIT, or asks about evaluating this task. Reports LAION Aesthetics v2 score.
Evaluates a deep neural network's capability to classify network traffic packets as normal or specific attack types. It probes spatial-temporal feature extraction, handling of class imbalance, and robustness against overlapping attack signatures in intrusion detection systems. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, or asks about evaluating this task. Reports Detection Rate (DR%).
This benchmark evaluates 3D lung tumor segmentation accuracy in CT imaging, comparing traditional CNN architectures against foundation models under standard, few-shot, and prompt-based inference regimes. It probes model robustness to varying training data sizes and input prompting strategies in a medical imaging context. Use when the user wants to benchmark on NSCLC-Radiomics (Lung1), Task06 (Medical Segmentation Decathlon), or asks about evaluating this task. Reports Dice Score.
Evaluates the ability of models to perform pixel-level semantic segmentation on large-scale, diverse image collections without human annotations. It probes unsupervised representation learning, category discovery, and fine-grained mask prediction capabilities. Use when the user wants to benchmark on ImageNet-S, ImageNet-S50, ImageNet-S300, or asks about evaluating this task. Reports mIoU.
This benchmark evaluates multimodal models' ability to comprehend extreme-length videos (averaging ~70 minutes) by testing six core temporal understanding capabilities. It probes long-term memory, multi-hop reasoning, and instruction-following across diverse video categories like sports, documentaries, and TV shows. Use when the user wants to benchmark on LVBench, or asks about evaluating this task. Reports accuracy.
This evaluation probes the demographic fairness of large vision-language models (LVLMs) by measuring how accurately they classify occupations and predict demographic attributes (gender, race, age, skin tone) across different prompt formats. It specifically quantifies performance gaps between demographic groups to identify persistent biases in model predictions. Use when the user wants to benchmark on FACET, UTKFace, or asks about evaluating this task. Reports recall.
Evaluates Large Vision-Language Models on their ability to generate factually consistent outputs aligned with visual input, specifically measuring the reduction of object hallucinations in open-ended generation while preserving general multimodal reasoning and visual grounding capabilities. Use when the user wants to benchmark on POPE, CHAIR, HallusionBench, AMBER, VizWiz, MME, LLaVA-Wild, MM-Vet, or asks about evaluating this task. Reports CHAIR (object hallucination score).
This benchmark evaluates multimodal large language models' ability to perform timestamp-aware summarization of long videos. It probes temporal grounding, instruction adherence regarding length constraints, and cross-modal consistency between visual/audio content and generated text descriptions. Use when the user wants to benchmark on LVSum, or asks about evaluating this task. Reports Kendall's tau & Spearman's rho.
Compute lvwerra/accuracy_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of lvwerra/accuracy_score.
Compute lvwerra/bary_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of lvwerra/bary_score.
Compute lvwerra/test via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of lvwerra/test.
Evaluates multilingual aspect-based sentiment analysis (ABSA) models on triplet extraction (aspect term, category, sentiment) and pairwise extraction (aspect term, sentiment) across 21 languages and 7 domains. Probes cross-lingual transfer, cross-domain adaptation, and zero-shot LLM prompting capabilities. Use when the user wants to benchmark on M-ABSA, or asks about evaluating this task. Reports Micro-F1.
Probes inflexible reasoning and medical abstraction in LLMs by presenting adversarial, long-tail clinical scenarios designed to trigger the Einstellung effect. It evaluates whether models can apply deductive logic and uncertainty estimation rather than relying on rote pattern matching or memorization from pretraining data. Use when the user wants to benchmark on M-ARC, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal information retrieval models across eight heterogeneous query-to-candidate modalities (text, image, image-text pairs) using instruction-tuned and fine-tuned vision-language models. Probes zero-shot generalization, cross-modality alignment, and the impact of instruction tuning on retrieval accuracy in large-scale candidate pools. Use when the user wants to benchmark on M-BEIR, or asks about evaluating this task. Reports Recall@5.
Evaluates 3D spatial reasoning and combinatorial planning by requiring models to assemble jigsaw-style pieces into a 5x5x5 cube under physical constraints. The benchmark probes the model's ability to extract geometric patterns from rendered views and logically arrange pieces without gaps or overlaps. Use when the user wants to benchmark on M-Cube, or asks about evaluating this task. Reports binary_evaluation.
Evaluates the effectiveness and refinement efficiency of neural network architecture search and selection methods on graph datasets. It measures how well a method can find near-optimal models within a limited search budget and how quickly it reaches a target performance level across diverse graph topologies and tasks. Use when the user wants to benchmark on Graph Architecture Search Benchmark (22 datasets), or asks about evaluating this task. Reports classification accuracy / AUC-ROC.
Evaluates multi-step spatial and physical reasoning in multimodal language models by requiring them to generate or validate chain-of-thought plans for solving Portal 2-inspired puzzle maps. The benchmark probes the model's ability to integrate visual map layouts with textual instructions to produce physically sound, multi-step traversal strategies. Use when the user wants to benchmark on M-Portal, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to verify scientific claims by cross-referencing textual assertions with provided multimodal evidence (figures/diagrams). It probes cross-modal reasoning, spatial/anatomical understanding, and the generation of factually grounded explanations. Use when the user wants to benchmark on M2-Verify-Med, M2-Verify-Gen, or asks about evaluating this task. Reports Macro-F1.
Evaluates vision-language models' ability to correctly identify true statements about images while rejecting culturally plausible but visually incorrect counterfactual statements. It specifically probes grounding failures and cultural reasoning biases across multiple languages and dialects. Use when the user wants to benchmark on M²CQA, or asks about evaluating this task. Reports CFHR.
Evaluates graph neural networks on predicting material properties from 3D crystal and molecular structures. It probes model performance across diverse material types, physical properties, and realistic data partitioning strategies. Use when the user wants to benchmark on OMDB, QMOF, MP, ISMETAL, EDOS, PDOS, DF2D, Perovskites, EFORM, Phonons, Dielectric, LOG_GVRH, LOG_KVRH, OC20, QM9, or asks about evaluating this task. Reports MAE / Accuracy.
Evaluates a model's ability to generate fluent, semantically accurate, and hallucination-free natural language responses grounded in video and audio inputs. It probes multimodal fusion, knowledge grounding, and dialogue generation capabilities across diverse question types and modalities. Use when the user wants to benchmark on AVSD10, NExT-OE, MUSIC-AVQA, or asks about evaluating this task. Reports CIDEr.
Evaluates multimodal retrieval-augmented generation systems across open-domain question answering, image captioning, and fact verification. It measures how effectively a system selects and utilizes retrieved multimodal evidence to improve generation quality and factual accuracy. Use when the user wants to benchmark on M2RAG, or asks about evaluating this task. Reports CIDEr.
Evaluates few-shot voice cloning systems on their ability to preserve speaker identity and transfer speaking styles using only 100 or 5 reference samples per speaker. It probes low-data robustness, style disentanglement, and naturalness in synthetic speech generation. Use when the user wants to benchmark on M2VoC 2021 Test Set, or asks about evaluating this task. Reports MOS (Quality, Speaker Similarity, Style Similarity).
Evaluates multimodal academic lecture understanding across speech recognition, speech synthesis, and slide/script generation. It probes models' ability to handle complex academic language, rare words, multimodal alignment, and knowledge comprehension. Use when the user wants to benchmark on M3AV, or asks about evaluating this task. Reports BWER, ROUGE-1/2/L.
Evaluates cooperative autonomous driving capabilities across perception, mapping, motion forecasting, occupancy prediction, and path planning. It probes whether multi-vehicle cooperation and realistic, non-straight trajectories improve ego-vehicle performance compared to single-vehicle baselines. Use when the user wants to benchmark on M3CAD, or asks about evaluating this task. Reports AMOTA.
Evaluates vision-language models' ability to perform multi-step, multi-modal chain-of-thought reasoning across diverse domains like science, commonsense, and mathematics. It probes the model's capacity to integrate visual information with textual reasoning steps and produce accurate final answers under various prompting and fine-tuning setups. Use when the user wants to benchmark on M3CoT, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the Chain-of-Thought reasoning capabilities of multimodal large language models on medical image understanding tasks. It probes whether models can generate transparent, step-by-step diagnostic pathways that align with clinical ground truth, rather than just producing correct final answers. Use when the user wants to benchmark on M3CoTBench, or asks about evaluating this task. Reports F1.
This benchmark evaluates unsupervised domain adaptation (UDA) methods for 3D medical image segmentation across eight practical domain shifts, including inter-modality changes (MRI-CT), scanner parameters, contrast presence, and radiation dose. It measures how well models trained on a source domain can segment target domain volumes without target labels, highlighting the robustness of adaptation techniques to real-world imaging variability. Use when the user wants to benchmark on AMOS, LIDC, B...
Probes long-context financial meeting understanding across three languages (EN, ZH, JA) and 11 GICS sectors. It evaluates a model's ability to condense lengthy transcripts into structured summaries, extract relevant question-answer pairs, and localize precise answers within designated sections while ignoring noise. Use when the user wants to benchmark on M3FinMeeting, or asks about evaluating this task. Reports compression ratio.
Evaluates a vision-language model's ability to follow multi-modal instructions, answer knowledge-based visual questions, and generalize to unseen languages and video tasks. It probes cross-modal alignment, cross-lingual transfer, and the model's conversational response quality. Use when the user wants to benchmark on M^3IT, OK-VQA, A-OKVQA, ViQuAE, Flickr-8k-CN, FM-IQA, Chinese-FoodNet, MSRVTT, iVQA, ActivityNet-QA, MSRVTT-QA, MSVD-QA, or asks about evaluating this task. Reports ROUGE-L.
Evaluates the ability of multimodal retrieval models to accurately rank relevant medical documents in response to text-and-image queries. It probes domain-specific alignment, handling of complex clinical terminology, and cross-specialty generalization in safety-critical healthcare settings. Use when the user wants to benchmark on M3Retrieve, or asks about evaluating this task. Reports nNDCG@10.
Evaluates vision-language models on multilingual, multicultural, and multimodal retrieval-augmented generation tasks. It measures how different retrieval strategies, language alignment, and model scale impact accuracy on culturally diverse image-question pairs. Use when the user wants to benchmark on CVQA, WorldCuisines, or asks about evaluating this task. Reports macro-averaged accuracy.
Evaluates the effectiveness and energy efficiency of Multi-dimensional Attention (MA) modules integrated into Spiking Neural Networks (SNNs) for event-based action recognition and static image classification. Use when the user wants to benchmark on DVS128 Gesture, DVS128 Gait, ImageNet-1K, or asks about evaluating this task. Reports Top-1 Accuracy (%).
Evaluates a model's ability to rerank search results to maximize topic diversity, ensuring that top-k results cover multiple relevant subtopics for a given query rather than just maximizing single-topic relevance. Use when the user wants to benchmark on TREC 2009~2012 Web Track, DU-DIV, or asks about evaluating this task. Reports α-NDCG@10.
This benchmark evaluates a model's ability to predict conversion rates for ad clicks under multiple attribution mechanisms. It probes ranking capability by measuring how well predicted probabilities distinguish positive from negative samples, both globally and per user. The task treats conversion prediction as a weighted binary classification problem where continuous attribution weights act as sample importance weights. Use when the user wants to benchmark on MAC, or asks about evaluating thi...