
Claude Skills by qhjqhj00
github.com/qhjqhj00This benchmark evaluates a model's ability to distinguish between real human-recorded songs and AI-generated synthetic songs. It specifically probes long-range temporal dependency modeling by testing performance on both short (5s) and long (120s) audio clips, while also measuring generalization to unseen generation algorithms and singers. Use when the user wants to benchmark on SONICS, or asks about evaluating this task. Reports F1 score.
Compute sonsus/harim_plus via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sonsus/harim_plus.
Compute sorgfresser/valid_efficiency_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sorgfresser/valid_efficiency_score.
Evaluates music source separation models on their ability to isolate individual instruments (vocals, bass, drums, other) from mixed audio tracks. It specifically probes robustness to label noise and bleeding artifacts in training data, as well as standard separation performance across different leaderboards. Use when the user wants to benchmark on SDXDB23_LabelNoise, SDXDB23_Bleeding, Standard (MDXDB21), or asks about evaluating this task. Reports SDR (Signal-to-Distortion Ratio).
Evaluates a model's ability to localize sound sources in audio-visual pairs by predicting spatial response maps or bounding boxes. It measures how effectively the model aligns audio signals with visual regions containing the corresponding sound, particularly testing robustness to semantically similar but mismatched cross-modal pairs. Use when the user wants to benchmark on VGGSound, SoundNet-Flickr, VGG-SS, SoundNet-Flickr-Test, or asks about evaluating this task. Reports cIoU.
Evaluates a model's ability to infer physical properties (air column length, container dimensions, flow rate, fill time, liquid weight) and classify container shapes solely from the acoustic characteristics of pouring liquids, without visual or tactile input. Use when the user wants to benchmark on Sound of Water 50, Wilson et al. [96] dataset, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
Evaluates text-queried sound separation models on natural and mixed audio. It measures how accurately a model isolates target sound sources from background noise or other sources, and how well it handles silence when the target is absent. Use when the user wants to benchmark on AudioSet, AudioCaps, ESC-50, or asks about evaluating this task. Reports SDRi.
This evaluation probes a model's ability to spatially localize sound sources in images or video frames given an accompanying audio clip. It measures how accurately the predicted bounding box overlaps with ground-truth annotations provided by multiple human annotators. Use when the user wants to benchmark on Flickr SoundNet Testset, VGG-Sound Source (VGG-SS), or asks about evaluating this task. Reports cIoU.
This evaluation probes whether large language models exhibit source attribution bias, specifically penalizing arguments when the attributed source's expected ideological position conflicts with the argument's content (coherence bias). It measures how models adjust credibility ratings based on source-argument alignment and whether they explicitly reason about source credibility. Use when the user wants to benchmark on Source Attribution Bias Evaluation, or asks about evaluating this task. Repo...
Evaluates a model's ability to identify which sentences in a source document contribute to an abstractive summary. It probes source sentence detection capability by ranking candidate sentences based on their inferred relevance to the summary. Use when the user wants to benchmark on SourceSum, or asks about evaluating this task. Reports NDCG.
Compute the SourceAggregatedSignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SourceAggregatedSignalDistortionRatio, or asks how to score with SourceAggregatedSignalDistortionRatio.
Evaluates biomedical named entity recognition (NER) capabilities on scientific literature. It probes a model's ability to identify and classify nine distinct bioentity types (e.g., genes, cell lines, diseases) within text extracted from published biological figures and captions. Use when the user wants to benchmark on SourceData-NLP, or asks about evaluating this task. Reports F1 score.
Evaluates multilingual encoder models on South Slavic languages (Croatian and Serbian) across three diverse NLP tasks: named entity recognition, parliamentary sentiment regression, and causal commonsense reasoning. Tests whether cost-efficient additional pretraining can match dedicated monolingual encoders without full from-scratch training. Use when the user wants to benchmark on hr500k, ReLDI-NormTagNER-hr, SETimes.SR, ReLDI-NormTagNER-sr, ParlaSent (HBS), COPA (Croatian & Serbian), or asks...
Evaluates named entity recognition performance across diverse domains and languages, specifically probing a model's ability to handle out-of-vocabulary words and morphological variation through hash-based embeddings versus traditional lookup embeddings. Use when the user wants to benchmark on CoNLL 2002, WNUT 2017, AnEM, Dutch Archaeology, OntoNotes 5.0, or asks about evaluating this task. Reports F1 score.
Evaluates a multi-agent DRL framework (MAPPO) combined with whale optimization for maximizing satellite coverage and data rates in 6G sub-THz networks using reconfigurable intelligent surfaces (RIS). Use when the user wants to benchmark on Simulated LEO Satellite-RIS Network Environment, or asks about evaluating this task. Reports average data rate.
Evaluates a model's ability to estimate the 6D pose (position and orientation) of non-cooperative spacecraft from monocular images. It probes generalization from photorealistic synthetic data to real space imagery and tests robustness to perceptual aliasing and orientation ambiguity. Use when the user wants to benchmark on URSO, SPEED, or asks about evaluating this task. Reports ESA Error.
Evaluates multi-modal spacecraft perception and pose estimation across 2D/3D segmentation, object detection, monocular depth estimation, and orientation estimation. Probes zero-shot generalization to unseen spacecraft configurations and robustness to long-tail class distributions and metallic surface reflections. Use when the user wants to benchmark on SpaceSense-Bench, or asks about evaluating this task. Reports mIoU.
This benchmark evaluates dynamic spatial reasoning and perception-memory integration in embodied environments. It probes models across three levels: static spatial perception, text-conditioned temporal memory, and visual-conditioned temporal memory, testing capabilities like object recognition, visual grounding, depth estimation, trajectory tracking, and long-horizon state reconstruction. Use when the user wants to benchmark on SpaMEM, or asks about evaluating this task. Reports mIoU.
Evaluates the reliability and fairness of span-level error detection metrics for machine translation auto-evaluators. It probes whether standard micro-averaged precision/recall/F1 scores produce consistent rankings compared to a proposed partial overlap matching strategy. Use when the user wants to benchmark on MQM 2022-2024, or asks about evaluating this task. Reports micro-averaged precision/recall/F1.
This benchmark probes a model's ability to perform compositional generalization in context-dependent Text-to-SQL. It evaluates whether models can correctly combine previously seen SQL query structures with novel modification patterns (e.g., new WHERE or ORDER BY clauses) in multi-turn dialogues. Use when the user wants to benchmark on SPARC-CG, or asks about evaluating this task. Reports question match (QM).
Probes cross-domain semantic parsing in context by requiring models to generate sequential SQL queries across multiple conversational turns. It evaluates the ability to maintain state, handle thematic evolution, and generalize to unseen databases while correctly resolving contextual dependencies. Use when the user wants to benchmark on SParC, or asks about evaluating this task. Reports question match.
Evaluates zero-shot text-to-speech generation by measuring speech intelligibility and speaker similarity across Chinese and English prompts. It also probes fine-grained control over voice attributes such as gender, pitch, and speaking rate. Use when the user wants to benchmark on Seed-TTS-eval, or asks about evaluating this task. Reports CER/WER.
Evaluates the alignment, factual grounding, and rule-following capabilities of dialogue agents through human preference comparisons. It also measures resilience to adversarial probing for specific harm rules and the quality of evidence-supported responses. Use when the user wants to benchmark on ELI5 + Free Dialogue Test Set, or asks about evaluating this task. Reports Three-model preference rate.
This benchmark evaluates the spatial reasoning capabilities of Visual Foundation Models (VFMs) by testing their ability to recognize spatial relations between object triples in synthetic images. It specifically probes both egocentric (camera-perspective) and allocentric (world-perspective) spatial understanding across diverse semantic objects and environments. Use when the user wants to benchmark on SpaRRTa, or asks about evaluating this task. Reports accuracy.
Evaluates the performance and efficiency of custom sparse GPU kernels for SpMM and SDDMM operations against standard libraries like cuSPARSE on deep learning workloads. It measures computational throughput, memory usage, and end-to-end speedups across various model architectures and batch sizes. Use when the user wants to benchmark on Sparse Matrix Dataset from DNNs, or asks about evaluating this task. Reports Geometric mean speedup.
Evaluates the clustering quality and feature selection capability of sparse k-means algorithms on biological and standard machine learning benchmark datasets. It probes how well the method separates known classes and selects discriminative features compared to baseline k-means variants. Use when the user wants to benchmark on Mice protein expression dataset, UCI/Keel/ASU Benchmark Datasets, or asks about evaluating this task. Reports Normalized Mutual Information (NMI).
Evaluates novel view synthesis performance from extremely sparse inputs (3 training views). It probes a model's ability to reconstruct 3D geometry and render photorealistic images for unseen camera poses without overfitting to the limited training data. Use when the user wants to benchmark on Realistic Synthetic 360°, LLFF, or asks about evaluating this task. Reports PSNR.
This evaluation protocol assesses the mathematical reasoning capabilities of LLMs trained with sparse reinforcement learning under strict memory constraints. It measures how well models maintain accuracy on standard math benchmarks when policy rollouts are generated using compressed KV caches instead of full context. Use when the user wants to benchmark on GSM8K, MATH500, Gaokao, Minerva Math, OlympiadBench, AIME24, AMC23, or asks about evaluating this task. Reports Pass@1 / Avg@32 accuracy.
Evaluates the quality and interpretability of sparse overcomplete word vector representations by measuring their ability to capture lexical similarity and perform downstream text classification tasks compared to dense baseline vectors. Use when the user wants to benchmark on SimLex, Senti., TREC, Sports, Comp., Relig., NP, or asks about evaluating this task. Reports accuracy.
This evaluation probes the computational efficiency and throughput of GPU inference kernels under unstructured sparsity. It measures how well a sparse matrix multiplication and sparse convolution implementation scales across different matrix dimensions and sparsity levels compared to dense and existing sparse baselines. Use when the user wants to benchmark on SparseRT SpMM & Convolution Benchmark, or asks about evaluating this task. Reports speedup.
Evaluates the runtime performance and speedup of sparse deep learning operators (SpMM, SDDMM) and end-to-end models (GraphSAGE, RGCN, Transformers) on GPU hardware using composable sparse formats and transformations. It probes how format decomposition and modular scheduling primitives improve cache utilization, load balancing, and Tensor Core utilization compared to vendor libraries and existing compilers. Use when the user wants to benchmark on cora, citeseer, pubmed, ppi, ogbn-arxiv, ogbn-p...
Evaluates the ability of LLMs and hybrid QA systems to perform tree-structured, multi-hop reasoning over combined text and table data, including complex SQL operations like aggregation, grouping, ordering, and cross-modal retrieval. Use when the user wants to benchmark on SPARTA, or asks about evaluating this task. Reports F1.
Probes a model's ability to perform multi-step spatial reasoning over natural language stories. It evaluates understanding of spatial relations (e.g., near, far, containment) and tests robustness against surface-level vocabulary changes and question phrasing variations. Use when the user wants to benchmark on SPARTQA-HUMAN, or asks about evaluating this task. Reports accuracy.
Probes a model's ability to detect overlapping audio events in dynamic spatial recordings, estimate their direction of arrival and distance, and perform reasoning about moving sound sources. Use when the user wants to benchmark on STARS23, FOA-MEIR Derived, or asks about evaluating this task. Reports F-score.
Evaluates spatial reasoning capabilities in vision-language models across a 2x2 cognitive taxonomy (Intrinsic/Extrinsic × Static/Dynamic). It probes mental rotation, multi-step 3D transformations, and dynamic scene simulation using synthetically rendered 3D VQA pairs. Use when the user wants to benchmark on Spatial-DISE, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perceive, reason about, and reconstruct 3D spatial layouts, multi-view relationships, and perspective-taking from visual inputs. It probes robustness against language shortcuts and tests generalization to longer video sequences and embodied manipulation tasks. Use when the user wants to benchmark on VSI-Bench, MMSI-Bench, MindCube, ViewSpatial-Bench, SITE, MMBench-En, EmbodiedBench (spatial subset), or asks about evaluating this task. Reports accuracy.
This evaluation probes a model's ability to perform complex spatial reasoning and perspective-taking across single and multiple images. It specifically tests whether the model can correctly establish geometric reference frames, handle multi-step transformations, and generalize across different spatial logic tasks without relying on dataset-specific biases. Use when the user wants to benchmark on MMSI-Bench, MindCube-tiny, OmniSpatial, SPBench, CV-Bench, or asks about evaluating this task. Rep...
Evaluates a vision-language model's ability to perform spatial reasoning tasks, including relative positioning, counting, size comparison, and cross-dataset generalization. It probes whether models learn transferable spatial concepts rather than memorizing dataset-specific patterns or visual artifacts. Use when the user wants to benchmark on GRAID-BDD, GRAID-NuImages, BLINK, A-OKVQA, NaturalBench, RealWorldQA, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of text-to-image and large language models to accurately generate and understand spatial relationships between objects. It probes geometric scene modeling and prepositional semantics grounding by testing models on simple and complex spatial prompts using basic geometric primitives. Use when the user wants to benchmark on SpatialRelBench, or asks about evaluating this task. Reports accuracy.
Evaluates self-supervised speech representation models on downstream tasks including speaker identification, phoneme recognition, automatic speech recognition, emotion recognition, and speech localisation. It specifically probes robustness to noise and reverberation by comparing performance under clean versus noisy/reverberant training and testing conditions. Use when the user wants to benchmark on Spatial SUPERB, or asks about evaluating this task. Reports ASR WER.
This benchmark evaluates large multimodal models' ability to perform 6D spatial reasoning across multiple difficulty levels. It probes capabilities including multi-object recognition, 2D and 3D location understanding, 3D orientation interpretation, and occlusion/collision prediction. It also quantifies systematic prediction biases across visual attributes like color, shape, size, and pose. Use when the user wants to benchmark on Spatial457, or asks about evaluating this task. Reports accuracy.
Evaluates vision-language models' ability to count objects and reason about spatial relationships (depth, distance, relative position) in images. It probes segmentation capabilities, attention alignment, and robustness to linguistic variations (out-of-distribution shifts). Use when the user wants to benchmark on CLEVR_CoGenT_ValB, CVBench, Pixmo-Count, Static Spatial Reasoning (SAT), VSR, VC Bench, or asks about evaluating this task. Reports Accuracy.
Compute the SpatialCorrelationCoefficient metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpatialCorrelationCoefficient, or asks how to score with SpatialCorrelationCoefficient.
Compute the SpatialDistortionIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpatialDistortionIndex, or asks how to score with SpatialDistortionIndex.
Evaluates the ability of vision-language models to perform multi-step spatial logical reasoning by tracking object dependencies and understanding scene layouts across real-world indoor environments. Use when the user wants to benchmark on SpatiaLQA, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal LLMs on 3D spatial reasoning, depth/distance estimation, and general visual question answering. It probes the model's ability to ground objects in 3D space, understand spatial relations, and generalize to real-world VQA tasks using only RGB inputs. Use when the user wants to benchmark on SpatialThinker Evaluation Suite (12 VQA Benchmarks), or asks about evaluating this task. Reports Accuracy.
Evaluates the effectiveness of embedding-based spatial keyword retrieval models by measuring how well they rank relevant Points of Interest (POIs) based on combined location and textual query signals. It probes the model's ability to handle spatio-textual relevance without manual weighting of spatial and textual factors. Use when the user wants to benchmark on Beijing, Shanghai, Geo-Glue, or asks about evaluating this task. Reports Recall@k, NDCG@k.
Evaluates spatiotemporal forecasting models on predicting future frames or climate indices from historical observations. It probes pixel-level reconstruction accuracy, structural similarity, and the model's ability to mitigate error propagation over extended lead times. Use when the user wants to benchmark on Moving MNIST, TrafficBJ, Human 3.6, SEVIR, ICAR-ENSO, or asks about evaluating this task. Reports MSE.
Evaluates clustering algorithms on their ability to track moving and static clusters in collective animal behavior data across space and time. It probes robustness in low-data regimes and the capacity to produce stable, interpretable cluster trajectories without relying on ground-truth labels for hyperparameter tuning. Use when the user wants to benchmark on Cakmak et al. Spatiotemporal Benchmark, or asks about evaluating this task. Reports total AMI.
This benchmark evaluates multimodal large language models on compositional spatial intelligence by testing their ability to reason across 10 atomic spatial capabilities (e.g., counting, localization, spatial relations) combined into 8 complex tasks. It probes scene understanding using 2D and 3D inputs through multiple-choice questions, revealing how models handle hierarchical spatial reasoning and capability integration. Use when the user wants to benchmark on SpaCE-10, or asks about evaluati...