Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,561–4,584 of 23,503 skills
Evaluates vision-language models' ability to count objects and reason about spatial relationships (depth, distance, relative position) in images. It probes segmentation capabilities, attention alignment, and robustness to linguistic variations (out-of-distribution shifts). Use when the user wants to benchmark on CLEVR_CoGenT_ValB, CVBench, Pixmo-Count, Static Spatial Reasoning (SAT), VSR, VC Bench, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates large multimodal models' ability to perform 6D spatial reasoning across multiple difficulty levels. It probes capabilities including multi-object recognition, 2D and 3D location understanding, 3D orientation interpretation, and occlusion/collision prediction. It also quantifies systematic prediction biases across visual attributes like color, shape, size, and pose. Use when the user wants to benchmark on Spatial457, or asks about evaluating this task. Reports accuracy.
Evaluates self-supervised speech representation models on downstream tasks including speaker identification, phoneme recognition, automatic speech recognition, emotion recognition, and speech localisation. It specifically probes robustness to noise and reverberation by comparing performance under clean versus noisy/reverberant training and testing conditions. Use when the user wants to benchmark on Spatial SUPERB, or asks about evaluating this task. Reports ASR WER.
Evaluates the ability of text-to-image and large language models to accurately generate and understand spatial relationships between objects. It probes geometric scene modeling and prepositional semantics grounding by testing models on simple and complex spatial prompts using basic geometric primitives. Use when the user wants to benchmark on SpatialRelBench, or asks about evaluating this task. Reports accuracy.
Evaluates a vision-language model's ability to perform spatial reasoning tasks, including relative positioning, counting, size comparison, and cross-dataset generalization. It probes whether models learn transferable spatial concepts rather than memorizing dataset-specific patterns or visual artifacts. Use when the user wants to benchmark on GRAID-BDD, GRAID-NuImages, BLINK, A-OKVQA, NaturalBench, RealWorldQA, or asks about evaluating this task. Reports accuracy.
This evaluation probes a model's ability to perform complex spatial reasoning and perspective-taking across single and multiple images. It specifically tests whether the model can correctly establish geometric reference frames, handle multi-step transformations, and generalize across different spatial logic tasks without relying on dataset-specific biases. Use when the user wants to benchmark on MMSI-Bench, MindCube-tiny, OmniSpatial, SPBench, CV-Bench, or asks about evaluating this task. Rep...
Evaluates a model's ability to perceive, reason about, and reconstruct 3D spatial layouts, multi-view relationships, and perspective-taking from visual inputs. It probes robustness against language shortcuts and tests generalization to longer video sequences and embodied manipulation tasks. Use when the user wants to benchmark on VSI-Bench, MMSI-Bench, MindCube, ViewSpatial-Bench, SITE, MMBench-En, EmbodiedBench (spatial subset), or asks about evaluating this task. Reports accuracy.
Evaluates spatial reasoning capabilities in vision-language models across a 2x2 cognitive taxonomy (Intrinsic/Extrinsic × Static/Dynamic). It probes mental rotation, multi-step 3D transformations, and dynamic scene simulation using synthetically rendered 3D VQA pairs. Use when the user wants to benchmark on Spatial-DISE, or asks about evaluating this task. Reports accuracy.
Probes a model's ability to detect overlapping audio events in dynamic spatial recordings, estimate their direction of arrival and distance, and perform reasoning about moving sound sources. Use when the user wants to benchmark on STARS23, FOA-MEIR Derived, or asks about evaluating this task. Reports F-score.
Probes a model's ability to perform multi-step spatial reasoning over natural language stories. It evaluates understanding of spatial relations (e.g., near, far, containment) and tests robustness against surface-level vocabulary changes and question phrasing variations. Use when the user wants to benchmark on SPARTQA-HUMAN, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of LLMs and hybrid QA systems to perform tree-structured, multi-hop reasoning over combined text and table data, including complex SQL operations like aggregation, grouping, ordering, and cross-modal retrieval. Use when the user wants to benchmark on SPARTA, or asks about evaluating this task. Reports F1.
Evaluates the runtime performance and speedup of sparse deep learning operators (SpMM, SDDMM) and end-to-end models (GraphSAGE, RGCN, Transformers) on GPU hardware using composable sparse formats and transformations. It probes how format decomposition and modular scheduling primitives improve cache utilization, load balancing, and Tensor Core utilization compared to vendor libraries and existing compilers. Use when the user wants to benchmark on cora, citeseer, pubmed, ppi, ogbn-arxiv, ogbn-p...
This evaluation probes the computational efficiency and throughput of GPU inference kernels under unstructured sparsity. It measures how well a sparse matrix multiplication and sparse convolution implementation scales across different matrix dimensions and sparsity levels compared to dense and existing sparse baselines. Use when the user wants to benchmark on SparseRT SpMM & Convolution Benchmark, or asks about evaluating this task. Reports speedup.
Evaluates the quality and interpretability of sparse overcomplete word vector representations by measuring their ability to capture lexical similarity and perform downstream text classification tasks compared to dense baseline vectors. Use when the user wants to benchmark on SimLex, Senti., TREC, Sports, Comp., Relig., NP, or asks about evaluating this task. Reports accuracy.
This evaluation protocol assesses the mathematical reasoning capabilities of LLMs trained with sparse reinforcement learning under strict memory constraints. It measures how well models maintain accuracy on standard math benchmarks when policy rollouts are generated using compressed KV caches instead of full context. Use when the user wants to benchmark on GSM8K, MATH500, Gaokao, Minerva Math, OlympiadBench, AIME24, AMC23, or asks about evaluating this task. Reports Pass@1 / Avg@32 accuracy.
Evaluates novel view synthesis performance from extremely sparse inputs (3 training views). It probes a model's ability to reconstruct 3D geometry and render photorealistic images for unseen camera poses without overfitting to the limited training data. Use when the user wants to benchmark on Realistic Synthetic 360°, LLFF, or asks about evaluating this task. Reports PSNR.
Evaluates the clustering quality and feature selection capability of sparse k-means algorithms on biological and standard machine learning benchmark datasets. It probes how well the method separates known classes and selects discriminative features compared to baseline k-means variants. Use when the user wants to benchmark on Mice protein expression dataset, UCI/Keel/ASU Benchmark Datasets, or asks about evaluating this task. Reports Normalized Mutual Information (NMI).
Evaluates the performance and efficiency of custom sparse GPU kernels for SpMM and SDDMM operations against standard libraries like cuSPARSE on deep learning workloads. It measures computational throughput, memory usage, and end-to-end speedups across various model architectures and batch sizes. Use when the user wants to benchmark on Sparse Matrix Dataset from DNNs, or asks about evaluating this task. Reports Geometric mean speedup.
This benchmark evaluates the spatial reasoning capabilities of Visual Foundation Models (VFMs) by testing their ability to recognize spatial relations between object triples in synthetic images. It specifically probes both egocentric (camera-perspective) and allocentric (world-perspective) spatial understanding across diverse semantic objects and environments. Use when the user wants to benchmark on SpaRRTa, or asks about evaluating this task. Reports accuracy.
Evaluates the alignment, factual grounding, and rule-following capabilities of dialogue agents through human preference comparisons. It also measures resilience to adversarial probing for specific harm rules and the quality of evidence-supported responses. Use when the user wants to benchmark on ELI5 + Free Dialogue Test Set, or asks about evaluating this task. Reports Three-model preference rate.
Evaluates zero-shot text-to-speech generation by measuring speech intelligibility and speaker similarity across Chinese and English prompts. It also probes fine-grained control over voice attributes such as gender, pitch, and speaking rate. Use when the user wants to benchmark on Seed-TTS-eval, or asks about evaluating this task. Reports CER/WER.
Probes cross-domain semantic parsing in context by requiring models to generate sequential SQL queries across multiple conversational turns. It evaluates the ability to maintain state, handle thematic evolution, and generalize to unseen databases while correctly resolving contextual dependencies. Use when the user wants to benchmark on SParC, or asks about evaluating this task. Reports question match.
This benchmark probes a model's ability to perform compositional generalization in context-dependent Text-to-SQL. It evaluates whether models can correctly combine previously seen SQL query structures with novel modification patterns (e.g., new WHERE or ORDER BY clauses) in multi-turn dialogues. Use when the user wants to benchmark on SPARC-CG, or asks about evaluating this task. Reports question match (QM).
Evaluates the reliability and fairness of span-level error detection metrics for machine translation auto-evaluators. It probes whether standard micro-averaged precision/recall/F1 scores produce consistent rankings compared to a proposed partial overlap matching strategy. Use when the user wants to benchmark on MQM 2022-2024, or asks about evaluating this task. Reports micro-averaged precision/recall/F1.