Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 2,857–2,880 of 22,847 skills
This benchmark evaluates multimodal large language models' ability to perform multi-step, context-sensitive reasoning over symbolic musical notation. It probes capabilities in harmonic analysis, rhythmic interpretation, structural form recognition, and expressive markings through multiple-choice questions derived from real-world compositions and forum queries. Use when the user wants to benchmark on WildScore, or asks about evaluating this task. Reports accuracy.
Evaluates language models' scientific reasoning capabilities by testing their ability to answer domain-specific multiple-choice questions derived from peer-reviewed literature and established scientific benchmarks. Use when the user wants to benchmark on WildSci-Val, GPQA-Aug, SuperGPQA, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
Probes the capability of models to perform semantic segmentation on unstructured, large-scale natural environments using both 2D images and 3D LiDAR point clouds. It evaluates robustness to semantic ambiguity, clutter, and temporal environmental shifts in outdoor traversals. Use when the user wants to benchmark on WildScenes, or asks about evaluating this task. Reports mIoU.
Evaluates machine learning models' robustness to real-world distribution shifts, specifically domain generalization and subpopulation shifts. It measures how much model performance degrades when tested on out-of-distribution (OOD) data compared to in-distribution (ID) data, highlighting gaps in generalization for real-world deployment. Use when the user wants to benchmark on WILDS, or asks about evaluating this task. Reports ID and OOD performance.
Evaluates novel view synthesis and motion mask estimation in dynamic environments where both camera and objects move. It probes a model's ability to remove transient objects, complete occluded backgrounds, and preserve scene geometry from sparse input views without 3D supervision or ground-truth poses. Use when the user wants to benchmark on D-RE10K-Mask, D-RE10K-iPhone, or asks about evaluating this task. Reports PSNR.
Evaluates the safety and robustness of language models against adversarial jailbreak attacks. It probes whether models can correctly refuse harmful requests while avoiding over-refusal on benign prompts, specifically under stealthy, adversarially composed prompts. Use when the user wants to benchmark on WILDJAILBREAK, or asks about evaluating this task. Reports Attack success rate (ASR).
This evaluation protocol assesses the safety moderation capabilities of LLMs and dedicated moderation models. It probes their ability to detect harmful content in user prompts, classify harmful or safe model responses, and identify whether a model appropriately refuses unsafe requests across multiple risk categories. Use when the user wants to benchmark on ToxicChat, OpenAI Mod, AegisSafetyTest, SimpleSafetyTests, Harmbench Prompt, Harmbench Resp, BeaverTails, SafeRLHF, XSTest-Resp, WildGuard...
This benchmark evaluates a model's ability to forecast the final spatial extent of a wildfire using multi-day spatio-temporal environmental and dynamic features. It probes the model's capacity to capture complex temporal dependencies and spatial patterns in binary segmentation tasks under significant class imbalance. Use when the user wants to benchmark on Mediterranean Wildfire Dataset (2006-2022), or asks about evaluating this task. Reports Dice Score.
This benchmark evaluates automatic speech recognition (ASR) capabilities on Mandarin speech produced by elderly individuals. It probes a model's robustness to real-world acoustic degradation, articulation variability, tremors, and diverse accent strengths under uncontrolled recording conditions. Use when the user wants to benchmark on WildElder, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates open-vocabulary monocular 3D object detection across diverse real-world and synthetic scenes. It probes the model's ability to localize and regress 3D bounding boxes using text or geometric prompts, measuring generalization to unseen categories and datasets with and without depth cues. Use when the user wants to benchmark on WildDet3D-Bench, Omni3D, Argoverse 2, ScanNet, Stereo4D, or asks about evaluating this task. Reports AP_3D.
Evaluates object detection algorithms on drone-captured images of wild berries in cluttered, dynamic forest environments. It probes localization and classification capabilities under severe lighting variations, occlusion, and cross-domain transfer settings (different areas, cameras, and datasets). Use when the user wants to benchmark on WildBe, or asks about evaluating this task. Reports Average Precision (AP).
This benchmark probes the robustness of automatic speech recognition (ASR) systems under realistic, out-of-distribution conditions. It specifically evaluates performance degradation across environmental noise, demographic shifts (accent, age, child speech), and linguistic diversity (short, incomplete, code-switched utterances), while also measuring semantic hallucination rates beyond standard lexical error metrics. Use when the user wants to benchmark on WildASR, or asks about evaluating this...
Evaluates the out-of-distribution (OOD) generalization capability of tabular regression models by measuring performance gaps between in-distribution and out-of-distribution test sets. It probes whether advanced OOD training strategies or complex architectures can reliably outperform simple Empirical Risk Minimization (ERM) on unseen data distributions. Use when the user wants to benchmark on VPower_S, VPower_R, Weather, or asks about evaluating this task. Reports MAE.
Evaluates a model's ability to perform complex tabular reasoning to answer open-ended questions based on a provided table. It probes the model's capacity to extract, aggregate, and filter information from structured data to produce short text span answers. Use when the user wants to benchmark on WikiTQ, or asks about evaluating this task. Reports denotation accuracy.
Evaluates a model's ability to perform semantic parsing over semi-structured tabular data by generating database queries or answers from natural language questions. It measures how well the model aligns textual utterances with table schemas and content to retrieve correct results. Use when the user wants to benchmark on WIKITABLEQUESTIONS, or asks about evaluating this task. Reports execution accuracy.
Evaluates a model's ability to translate natural language questions into executable SQL queries against a given database schema. It probes semantic parsing, table understanding, and the generation of syntactically and semantically correct structured queries. Use when the user wants to benchmark on WikiSQL, or asks about evaluating this task. Reports Acc_ex.
Evaluates a model's ability to detect malicious Wikipedia editors (vandals) using only benign user data for training. It probes one-class anomaly detection and sequential behavior modeling by measuring how well the system distinguishes benign from malicious users based on edit sequences. Use when the user wants to benchmark on UMDWikipedia, or asks about evaluating this task. Reports F1.
Assesses the quality of automatically mined parallel sentence pairs by training neural machine translation models and measuring their downstream translation accuracy. Use when the user wants to benchmark on WikiMatrix, or asks about evaluating this task. Reports BLEU.
This benchmark evaluates cross-lingual abstractive summarization, specifically the ability of models to generate coherent English summaries from articles written in other languages. It probes how well systems can handle translation and summarization jointly, either through direct cross-lingual fine-tuning or two-step pipeline approaches. Use when the user wants to benchmark on WikiLingua, or asks about evaluating this task. Reports ROUGE-L F1.
Evaluates text summarization systems on procedural, step-by-step articles written by non-journalists. It probes the model's ability to handle long sequences, non-inverted-pyramid structures, and high-abstraction content compared to standard news datasets. Use when the user wants to benchmark on WikiHow, or asks about evaluating this task. Reports ROUGE-L.
Evaluates a model's ability to disambiguate named entities in text by matching them to correct Wikidata entries using graph-based representations. It probes how well different neural architectures leverage graph triplet information versus full graph topology for entity resolution. Use when the user wants to benchmark on Wikidata-Disamb, or asks about evaluating this task. Reports F1.
Evaluates the factual accuracy, conversational quality, and latency of knowledge-grounded chatbots in simulated multi-turn dialogues across head, tail, and recent knowledge domains. Use when the user wants to benchmark on Simulated Dialogues (WikiChat), or asks about evaluating this task. Reports factual_accuracy.
Evaluates abstractive multi-document summarization models on their ability to generate coherent, content-adequate summaries across three domains (Company, Film, Animal). The protocol measures lexical and sentence-level overlap between generated summaries and reference summaries using ROUGE metrics, while also contextualizing scores against a baseline overlap between input documents and summaries. Use when the user wants to benchmark on WIKICATSUM, or asks about evaluating this task. Reports R...
Evaluates multi-domain aspect-based summarization, requiring models to first discover relevant aspects (Wikipedia section titles) from cited references and then generate domain-specific summaries. It probes content selection, cross-document pronoun resolution, and temporal ordering in multi-source generation. Use when the user wants to benchmark on WikiAsp, or asks about evaluating this task. Reports R-2.