Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,657–4,680 of 23,504 skills
Evaluates a generative model's ability to produce coherent stroke-based vector sketches through reconstruction, latent space interpolation, and completion of incomplete drawings. It probes how well the model captures conceptual features and organizes them in a continuous latent manifold. Use when the user wants to benchmark on QuickDraw, or asks about evaluating this task. Reports LR.
Evaluates early retrieval performance using partial or low-quality sketches combined with text, measuring how quickly the correct target image appears in the ranked list as the sketch is drawn. It probes robustness to incomplete visual inputs and multimodal fusion. Use when the user wants to benchmark on FS2K-SDE1, FS2K-SDE2, User-SDE, or asks about evaluating this task. Reports m@A, m@B.
Probes a model's ability to detect abnormal events in pedestrian videos by analyzing skeleton sequences. It measures how well the model distinguishes normal from anomalous motion patterns through reconstruction or prediction errors, evaluated via frame-level anomaly scoring. Use when the user wants to benchmark on Avenue, HR-Avenue, HR-STC, UBnormal, HR-UBnormal, or asks about evaluating this task. Reports AUC.
Evaluates LLM-based natural language question answering over heterogeneous data sources by testing the model's ability to generate SQL queries that correctly invoke both database tables and external APIs. It probes complex query planning, API sequencing, and routing across mixed data modalities. Use when the user wants to benchmark on Spider (modified with API-replaced tables), or asks about evaluating this task. Reports execution accuracy.
Evaluates simultaneous 3D scene reconstruction and multi-task scene understanding on sparse multi-view images. It probes geometric accuracy, novel view synthesis quality, and cross-view consistent segmentation across semantic, instance, panoptic, and text-referred tasks. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports mIoU.
Evaluates the policy success rate and human workload reduction of a human-in-the-loop robot learning framework over multiple deployment rounds on contact-rich manipulation tasks in simulation and real-world settings. Use when the user wants to benchmark on Sirius Robot Manipulation Tasks, or asks about evaluating this task. Reports success rate.
Evaluates multimodal large language models on their ability to understand and assess the quality of scientific images. It probes two distinct capabilities: factual reasoning about scientific content (SIQA-U) and alignment with human expert judgments on perceptual and knowledge dimensions (SIQA-S). Use when the user wants to benchmark on SIQA, or asks about evaluating this task. Reports accuracy, SRCC.
Evaluates a neural network's ability to interpolate sparse optical flow matches into dense flow maps. It probes the model's capacity to handle missing pixels, occlusions, and motion boundaries while preserving flow accuracy across diverse scenes. Use when the user wants to benchmark on MPI Sintel, KITTI 2012, KITTI 2015, or asks about evaluating this task. Reports EPE.
Evaluates language models' ability to process and generate text in two distinct Sinhala writing systems: standard Unicode and Romanized transliteration. It probes script-specific perplexity, coherence, and grammatical correctness, highlighting morphological understanding and training data biases. Use when the user wants to benchmark on Sinhala Unicode & Romanized, or asks about evaluating this task. Reports perplexity.
Evaluates singing voice enhancement models on real-world acoustic scenarios, measuring their ability to improve perceptual quality and content intelligibility of degraded singing vocals without degrading speech capabilities. Use when the user wants to benchmark on SingVERSE, or asks about evaluating this task. Reports perceptual quality.
Evaluates clustering algorithms on single-cell RNA sequencing data to assess their ability to separate distinct cell types and capture transitional states. It combines quantitative clustering metrics with visual validation using marker gene expression profiles. Use when the user wants to benchmark on Breast cancer dataset, Embryo neurone dataset, or asks about evaluating this task. Reports misclustering rate.
This benchmark evaluates a model's ability to extract singing vocals from mixed audio tracks. It probes the effectiveness of adversarial semi-supervised learning in separating sources without relying on perfectly paired mixture-source training data. Use when the user wants to benchmark on DSD100, iKala, MedleyDB, CCMixter, or asks about evaluating this task. Reports SDR.
Probes a model's ability to convert speech to high-quality singing while preserving the target speaker's timbre, using only normal speech samples. Evaluates both audio naturalness and speaker similarity through subjective listening tests. Use when the user wants to benchmark on Tencent multi-speaker speech corpus (TSP), Tencent singing corpus (TSG), Separate singing corpus (test), or asks about evaluating this task. Reports MOS (Naturalness).
Evaluates the ability of audio embedding and classification models to correctly identify the vocalist performing a song track. It specifically probes robustness to synthetic/deepfake voices and generalization across different music datasets and genre contexts. Use when the user wants to benchmark on Train, Validation, Closed, FMA, MTG, Cloned, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates multimodal large language models' ability to detect offensive memes containing social biases within a Singaporean cultural and linguistic context. It tests both standalone VLM reasoning and a multi-step pipeline combining OCR and translation, while comparing different fine-tuning strategies and data compositions. Use when the user wants to benchmark on Singapore Offensive Memes Dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates text-independent speaker identification and verification on raw audio waveforms, testing the model's ability to extract speaker-specific features and generalize across different corpus sizes and utterance lengths. Use when the user wants to benchmark on TIMIT, Librispeech, or asks about evaluating this task. Reports accuracy.
Evaluates simultaneous speech-to-speech translation models on translation accuracy, latency, speaker voice preservation, and audio naturalness across multiple languages. It probes the model's ability to generate high-quality target speech in real-time without relying on word-level aligned training data. Use when the user wants to benchmark on Audio-NTREX-4L, Europarl-ST, or asks about evaluating this task. Reports ASR-BLEU.
Probes an agent's ability to abstain from tool use when no appropriate tools are available, specifically measuring its tendency to hallucinate tool invocations in two controlled scenarios: when no tools are provided and when irrelevant distractor tools are present. Use when the user wants to benchmark on SimpleToolHalluBench, or asks about evaluating this task. Reports R_NTA.
Evaluates the lexical, semantic, and syntactic diversity of a synthetic story dataset compared to a baseline. It measures n-gram distribution, compression ratio, Self-BLEU, n-gram diversity scores, POS template rates, and model-judged semantic variation to assess how well the dataset avoids formulaic phrasing and redundancy. Use when the user wants to benchmark on SimpleStories, or asks about evaluating this task. Reports compression_ratio.
Evaluates a robot's ability to execute manipulation tasks from visual inputs and textual instructions, testing both atomic skill execution and high-level instruction generalization in simulated and real-world settings. Use when the user wants to benchmark on SimplerEnv, SimplerEnv-Instruct, or asks about evaluating this task. Reports visual matching (VM).
This benchmark evaluates a model's ability to retrieve videos based on semantic motion similarity, disentangling dynamic behavior from static appearance, camera viewpoint, and scene context. It tests robustness to appearance variations in controlled synthetic settings and unsynchronized, in-the-wild video pairs. Use when the user wants to benchmark on SimMotion-Synthetic, SimMotion-Real-1K, Jester, or asks about evaluating this task. Reports Retrieval accuracy.
Evaluates a model's ability to generalize across unseen domains in multi-modal action recognition. It probes feature disentanglement, cross-modal translation for missing modalities, and robustness to domain shifts in video, audio, and optical flow inputs. Use when the user wants to benchmark on EPIC-Kitchens, HAC, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates multimodal task-oriented dialogue capabilities, specifically focusing on dialogue state tracking, disambiguation, coreference resolution, and response generation using visual scene representations. Use when the user wants to benchmark on SIMMC 2.0, or asks about evaluating this task. Reports Intent-F1.
Evaluates the sample efficiency, optimization stability, and scaling capabilities of deep reinforcement learning algorithms across diverse continuous control environments. Use when the user wants to benchmark on MuJoCo, DMC Suite, MyoSuite, HumanoidBench, or asks about evaluating this task. Reports Normalized Return.