Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,481–6,504 of 20,821 skills
Evaluates automatic speech recognition (ASR) models on long-form audio, measuring accuracy in predicting word and character sequences. It specifically probes the model's ability to handle full-text formatting, including punctuation and casing, and tests performance across different training data scales. Use when the user wants to benchmark on Libriheavy, or asks about evaluating this task. Reports WER.
Evaluates continuous speech separation and speaker diarization in reverberant, multi-microphone environments with varying speaker overlap ratios. It probes the system's ability to separate overlapping speech, estimate speaker locations (DOA), cluster them across time blocks, and produce accurate diarization and speech recognition outputs. Use when the user wants to benchmark on LibriCSS, or asks about evaluating this task. Reports DER.
This benchmark evaluates non-invasive brain-computer interface (BCI) capabilities by testing neural speech decoding from magnetoencephalography (MEG) recordings. It probes a model's ability to detect speech presence, classify phonemes, and identify words from high-fidelity, within-subject neural data aligned with naturalistic audio stimuli. Use when the user wants to benchmark on LibriBrain, or asks about evaluating this task. Reports Balanced Accuracy.
Evaluates the ability of end-to-end speech separation models to isolate target speakers from noisy multi-speaker mixtures. It probes noise-robustness and speaker separation capability under realistic background noise conditions. Use when the user wants to benchmark on Libri2Mix-noisy, Libri3Mix-noisy, or asks about evaluating this task. Reports SI-SNRi (dB).
Evaluates the safety and capability of large language models across a broad set of safety tasks, measuring how well models handle direct risky prompts, adversarial attacks, and benign prompts without over-refusal or unsafe generation. Use when the user wants to benchmark on Libra-Eval, or asks about evaluating this task. Reports task_score.
Evaluates the robustness of vision-language-action (VLA) models under realistic perturbations across seven dimensions (camera, robot, language, light, background, noise, layout). It probes visual shift tolerance, kinematic reasoning, and linguistic robustness by measuring success rates on a curated set of non-trivial tasks. Use when the user wants to benchmark on LIBERO-Plus, or asks about evaluating this task. Reports success rate.
Evaluates a Vision-Language-Action model's ability to perform precise robotic manipulation and maintain robustness under environmental perturbations. It probes spatial understanding, object manipulation, instruction following, and long-horizon task execution in both standard and perturbed simulation environments. Use when the user wants to benchmark on LIBERO, LIBERO-Plus, or asks about evaluating this task. Reports success_rate.
Evaluates a robot policy's ability to sequentially learn multiple manipulation tasks while transferring knowledge and minimizing catastrophic forgetting. It measures forward transfer speed, backward transfer (forgetting), and overall performance across a curriculum of procedurally generated tasks. Use when the user wants to benchmark on LIBERO-LONG, LIBERO-SPATIAL, LIBERO-OBJECT, LIBERO-GOAL, or asks about evaluating this task. Reports FWT.
Evaluates whether Vision-Language-Action (VLA) models can follow counterfactual language instructions in robotic manipulation tasks. It specifically probes for 'vision shortcuts' where models default to well-learned visual behaviors instead of adhering to the given text commands. Use when the user wants to benchmark on LIBERO-CF, or asks about evaluating this task. Reports grounding rate.
This evaluation probes a model's ability to perform binary fact-checking on short political claims by mapping multi-class truthfulness labels to positive/negative categories. It specifically tests how well the system handles compositional reasoning and uncertainty, requiring it to output definitive verdicts or abstain. Use when the user wants to benchmark on LIAR, or asks about evaluating this task. Reports accuracy.
Evaluates a malware detection model's ability to adapt to natural concept drift over time using a rolling monthly update setup on real-world Windows malware binaries. Use when the user wants to benchmark on MB-24+, or asks about evaluating this task. Reports accuracy.
Evaluates multilingual legal NLP models across text classification and named entity recognition tasks. It probes the ability of models to handle domain-specific legal jargon, long-form documents, and cross-lingual generalization across 24 languages. Use when the user wants to benchmark on LEXTREME, or asks about evaluating this task. Reports macro-F1.
This benchmark evaluates the ability of sequence-to-sequence models to generate accurate, concise, and faithful summaries of long legal documents across multiple jurisdictions. It probes domain-specific summarization capabilities, testing how well models handle varying input lengths, compression ratios, and the balance between extractive and abstractive generation in legal English. Use when the user wants to benchmark on BillSum, EurLexSum, GovReport, MultiLexSum-Long, MultiLexSum-Short, Mult...
Probes large language models' ability to extract structured legal relations (relation types and factual arguments) from Chinese civil court judgments. It evaluates both zero-shot prompting and fine-tuning capabilities, while also measuring performance on long-tail relation types and downstream legal reasoning tasks. Use when the user wants to benchmark on LexRel, or asks about evaluating this task. Reports micro-F1.
Evaluates how well a model learns word meanings and general language modeling performance when trained with lexicon-level contrastive visual grounding. Probes concrete vs. abstract word acquisition, verb relation learning, and next-token prediction accuracy on held-out text. Use when the user wants to benchmark on Word Relatedness, Semantic Feature Prediction, Context Understanding, Lexical Relation Prediction, SimVerb-3500, or asks about evaluating this task. Reports Perplexity.
Evaluates the robustness and classification accuracy of pre-trained language models when augmented with rule-based lexical simplification as auxiliary inputs. It probes whether lemmatization and rare-word replacement preserve semantic meaning while mitigating lexical diversity effects on downstream NLU tasks. Use when the user wants to benchmark on SST-2, CR, SUBJ, MR, AG, or asks about evaluating this task. Reports accuracy.
Probes large language models' ability to perform structured, multi-step legal reasoning on real-world law exam questions. It evaluates both open-ended legal analysis and multiple-choice selection across diverse jurisdictions and legal domains. Use when the user wants to benchmark on LEXam, or asks about evaluating this task. Reports accuracy.
Evaluates text-to-image generation models on their ability to accurately render specified text within images, control visual attributes (color, position, font), and maintain aesthetic quality. It measures OCR fidelity, attribute controllability, and human-perceived aesthetics. Use when the user wants to benchmark on LeX-Bench, SimpleBench, CreateBench, AnyText-Benchmark, or asks about evaluating this task. Reports PNED.
Evaluates whether a latent world model captures physical structure and dynamics by probing latent representations for physical quantities and measuring predictive surprise under physical versus visual perturbations. Use when the user wants to benchmark on TwoRoom, PushT, OGBench-Cube, Reacher, or asks about evaluating this task. Reports MSE.
Evaluates a video-language model's capacity for structured video understanding across six professional dimensions (subject, aesthetics, camera language, editing, narrative, dissemination) while maintaining general multimodal capabilities. It probes timeline-grounded reasoning, temporal localization, and document/OCR comprehension. Use when the user wants to benchmark on FeedBench, Open Benchmarks (Video-MME, MVBench, MMBench-EN, etc.), or asks about evaluating this task. Reports accuracy.
Evaluates the ability of multilingual embedding models to retrieve relevant legislative documents given structured metadata queries. It probes cross-lingual semantic alignment and domain-adaptive retrieval performance across varying language resource levels. Use when the user wants to benchmark on LEMUR, or asks about evaluating this task. Reports Acc@k.
Evaluates how neural network hyperparameters (activation functions, depth, learning rate) affect output complexity and robustness to input perturbations. Use when the user has predictions and gold and needs to compute Lempel-Ziv Complexity.
Evaluates the accuracy and robustness of crystal structure fingerprinting and hashing algorithms for de-duplicating quantum chemistry materials databases. It probes sensitivity to structural perturbations (atomic noise, lattice strain, translations) and performance on disordered crystal systems. Use when the user wants to benchmark on LeMat-Bulk, or asks about evaluating this task. Reports success rate.
This benchmark evaluates multilingual text-to-speech synthesis and text-based speech editing capabilities. It probes pronunciation stability, cross-lingual generalization, and the perceptual naturalness of localized audio edits across multiple languages. Use when the user wants to benchmark on LEMAS-Dataset, or asks about evaluating this task. Reports WER.