Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,513–4,536 of 23,503 skills
Evaluates a model's ability to translate natural language questions into correct SQL queries across diverse database domains. It probes schema linking, lexical matching, and complex query synthesis including joins, aggregations, and subqueries. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports exact matching accuracy.
Evaluates the text-to-SQL and dialog state tracking capabilities of language models by measuring how accurately they generate syntactically and semantically valid SQL queries from natural language questions, with and without constrained auto-regressive decoding. Use when the user wants to benchmark on Spider, CoSQL, or asks about evaluating this task. Reports exact-set-match accuracy.
Evaluates language models' ability to perform real-world enterprise text-to-SQL workflows. It probes agentic reasoning, multi-step data transformation, database schema navigation, and SQL dialect adaptation across complex, long-context tasks. Use when the user wants to benchmark on Spider 2.0, Spider 2.0-lite, Spider 2.0-snow, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates the semantic propositional content of image captions by transforming them into scene graphs that encode objects, attributes, and relations. It computes an F-score over these logical propositions to measure how well a generated caption captures the underlying meaning of an image compared to human references. Use when the user has predictions and gold and needs to compute SPICE.
This evaluation probes a model's ability to solve challenging mathematical and general reasoning tasks, both from standard benchmarks and document-grounded self-play generated questions. It measures how well the model can extract information, perform multi-step logical deduction, and produce verifiable answers across diverse academic and competition-level datasets. Use when the user wants to benchmark on MATH-500, OlympiadBench, Minerva Math, GSM8K, AMC, AIME'24, AIME'25, SuperGPQA, GPQA-Diam...
Evaluates the energy efficiency and computational capability of the SpiNNaker2 processing element architecture across a suite of neuromorphic and deep learning workloads, including classical spiking neural networks, hybrid SNN/DNN frameworks, and standard DNN layers. Use when the user wants to benchmark on SpiNNaker2 Benchmark Suite, or asks about evaluating this task. Reports energy_efficiency.
Probes visual perception and reasoning capabilities of vision-language models across 25 distinct task types, including symmetry, spatial transformations, chart interpretation, and sequence prediction. Uses a synthetic environment with verifiable ground truth to measure model accuracy against human baselines. Use when the user wants to benchmark on Sphinx, or asks about evaluating this task. Reports accuracy.
This evaluation probes a differentiable spherical Voronoi partition for modeling view-dependent appearance and reflections in 3D Gaussian Splatting. It measures novel-view synthesis reconstruction fidelity and rendering efficiency against established radiance field baselines across synthetic and real-world scenes. Use when the user wants to benchmark on Mip-NeRF360, DeepBlending, Tanks&Temples, NeRF-Synthetic, Ref-NeRF, GlossySynthetic, Ref-Real, or asks about evaluating this task. Reports PSNR.
This evaluation probes an algorithm's ability to detect faint exoplanet signals buried in structured stellar speckle noise and accurately characterize their physical properties in direct imaging observations. It measures detection sensitivity across varying false alarm rates and quantifies regression accuracy for astrophysical parameters like flux and sub-pixel position. Use when the user wants to benchmark on SPHERE, or asks about evaluating this task. Reports ARE.
Evaluates spatial reasoning and visual understanding in vision-language models across a hierarchy of tasks, including single-skill perception (position, counting, distance, size), multi-skill integration, and complex physical-world reasoning (object occlusion and manipulation). It specifically probes egocentric vs. allocentric perspective taking and susceptibility to object hallucination. Use when the user wants to benchmark on SPHERE, or asks about evaluating this task. Reports accuracy.
Evaluates end-to-end speech-to-text models on financial domain audio, specifically testing their ability to produce fully formatted orthographic transcriptions including punctuation, capitalization, number denormalization, and disfluency handling. The benchmark measures how well acoustic architectures can learn text formatting directly from audio signals without relying on post-processing pipelines. Use when the user wants to benchmark on SPGISpeech, or asks about evaluating this task. Report...
Evaluates end-to-end speaker-tagged automatic speech recognition (ASR) and speaker diarization on financial domain audio. It probes a model's ability to accurately transcribe speech while correctly assigning speaker identities to utterance segments in multi-speaker conversations. Use when the user wants to benchmark on SPGISpeech 2.0, or asks about evaluating this task. Reports cpWER.
Evaluates a model's ability to jointly identify named entity spans with their types and extract relational tuples between them from unstructured text. It probes span-based representation learning, localized context modeling, and joint classification without relying on sequential tagging schemes like BIO. Use when the user wants to benchmark on CoNLL04, SciERC, ADE, or asks about evaluating this task. Reports F1 score (micro/macro-averaged).
Evaluates the accuracy and throughput of speculative decoding methods across diverse semantic domains and varying input sequence lengths. It probes how draft length, batch size, vocabulary pruning, and inference frameworks impact real-world serving efficiency compared to baseline autoregressive generation. Use when the user wants to benchmark on SPEED-Bench, or asks about evaluating this task. Reports AL.
Probes speech reasoning capabilities in large audio-language models across factual, procedural, and normative dimensions. It tests whether models can perform multi-step inference, maintain logical coherence, and make normative judgments when processing spoken input under varying prosodic and emotional conditions. Use when the user wants to benchmark on SpeechR, or asks about evaluating this task. Reports Accuracy.
Evaluates large audio-language models (LALMs) on their ability to generate speech with fine-grained paralinguistic features, including dynamic intra-utterance variation and context-aware adaptation. It probes how well models interpret and modulate tone, pitch, emotion, and non-linguistic vocalizations in response to textual instructions and contextual cues. Use when the user wants to benchmark on SpeechParaling-Bench, or asks about evaluating this task. Reports Judge Score (0-100).
Binary classification of spoken multi-speaker dialogues to detect the presence of mental manipulation tactics. It probes an audio-language model's ability to identify subtle manipulative cues in synthetic speech without relying on text transcripts. Use when the user wants to benchmark on SpeechMentalManip, or asks about evaluating this task. Reports accuracy.
Evaluates speech language models on medical consultation tasks, covering single-turn medical knowledge Q&A, multi-turn diagnostic conversations, real-world clinical robustness, and speech output quality. Use when the user wants to benchmark on SpeechMedBench, CMB, CME, MedDG, AIHospital, MedSafetyBench, Wild, or asks about evaluating this task. Reports CMB, CME, MedDG, AIHospital.
Evaluates a unified speech-language model's ability to perform end-to-end automatic speech recognition (ASR), named entity recognition (NER), and sentiment analysis (SA) on low-resource speech datasets. It tests parameter-efficient adapter-based alignment of speech encoder features to a language model, along with classifier regularization and LoRA fine-tuning. Use when the user wants to benchmark on LibriSpeech, SLUE-VoxPopuli, SLUE-VoxCeleb, or asks about evaluating this task. Reports WER, F...
This benchmark evaluates how well automated metrics and large audio-language models can judge speech naturalness and detect deepfakes compared to human preferences. It probes the alignment of computational quality scores with human perceptual judgments across multiple languages and TTS models. Use when the user wants to benchmark on SpeechJudge-Eval, or asks about evaluating this task. Reports AudioLLM Pairwise Accuracy.
Evaluates speech instruction-following capabilities across closed-ended, open-ended, and adjustment tasks under varying acoustic conditions (background noise, accents, disfluencies) in English and Chinese. Use when the user wants to benchmark on SpeechInstructBench, or asks about evaluating this task. Reports instruction-level accuracy (I).
Evaluates whether self-supervised speech models capture linguistic knowledge by probing them on a speech-adapted version of the GLUE benchmark. Tasks include grammaticality judgment, sentence similarity, paraphrase identification, and natural language inference, all converted from text to speech via TTS. Use when the user wants to benchmark on SpeechGLUE, or asks about evaluating this task. Reports Accuracy (Acc).
This benchmark evaluates speech deepfake detection models on their ability to generalize across diverse synthesis methods (TTS, voice conversion, neural vocoders), multiple languages, and unseen speaker identities. It probes whether models learn inherent spoofing artifacts or merely memorize specific speaker voices or generation techniques. Use when the user wants to benchmark on SpeechFake, or asks about evaluating this task. Reports EER.
Evaluates speech separation models' ability to isolate individual speaker signals from multi-speaker mixtures under various acoustic conditions, including moving sources, environmental noise, and musical noise. It measures both objective signal quality and subjective perceptual metrics to assess generalization from synthetic to real-world dynamic scenarios. Use when the user wants to benchmark on SonicSet, RealSEP, HumanSEP, LRS2-2Mix, Libri2Mix, or asks about evaluating this task. Reports SI...