Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 2,785–2,808 of 22,846 skills
Evaluates LLM-based agents' ability to comprehend multilingual shopping instructions and successfully navigate interactive web environments across 14 languages. Use when the user wants to benchmark on X-WebAgentBench, or asks about evaluating this task. Reports Task Score.
Probes the spatial feature flow and patch-level influence in Vision Mamba models using classical control theory. It quantifies how input image patches drive hidden state dynamics across hierarchical layers, revealing domain-specific diagnostic feature extraction patterns. Use when the user wants to benchmark on CMMD, DermaMNIST, BloodMNIST, or asks about evaluating this task. Reports influence score.
This benchmark evaluates multilingual topic classification on social media tweets across four languages (English, Spanish, Japanese, Greek). It probes models' ability to generalize across languages and training regimes, including zero-shot, few-shot, monolingual, cross-lingual, and multilingual fine-tuning settings. Use when the user wants to benchmark on X-Topic, or asks about evaluating this task. Reports macro-F1.
Evaluates multi-modal large language models on progressive clinical reasoning in ophthalmic diagnosis. It tests the model's ability to perform a six-stage diagnostic chain (from image quality assessment to clinical decision-making) while integrating cross-modality imaging data and calibrating its uncertainty. Use when the user wants to benchmark on X-PCR, or asks about evaluating this task. Reports Stage-Wise Accuracy (SWA).
Evaluates the text rendering, text-to-image generation, and image understanding capabilities of a discrete autoregressive image generation model trained with reinforcement learning. It probes the model's ability to follow complex instructions, render long texts accurately, and generate high-fidelity images without relying on classifier-free guidance. Use when the user wants to benchmark on OneIG-Bench, LongText-Bench, DPG-Bench, GenEval, POPE, GQA, MMBench, SEEDBench-Img, DocVQA, OCRBench, or...
Evaluates an end-to-end navigation model's ability to predict robot dynamics and successfully navigate through structured and cluttered warehouse environments. It probes both open-loop trajectory and speed prediction accuracy, as well as closed-loop mission success, navigation efficiency, and motion smoothness in seen and out-of-distribution settings. Use when the user wants to benchmark on X-Mobility Warehouse Dataset, or asks about evaluating this task. Reports mission success rate (SR).
Evaluates large language models' ability to understand and classify disruptive weather impacts from historical and modern newspaper articles, and to answer related questions by ranking relevant information. It specifically probes models' capacity to handle climate-related polysemy, extract nuanced societal responses, and correctly identify passages without weather impacts. Use when the user wants to benchmark on WXImpactBench, or asks about evaluating this task. Reports F1-score.
This evaluation probes a model's ability to perform single-channel speech separation by learning discriminative time-frequency embeddings that group mixture components into distinct speaker clusters. It specifically tests generalization to unseen speakers and scaling to three-speaker mixtures without retraining. Use when the user wants to benchmark on WSJ0-based Speech Mixtures, or asks about evaluating this task. Reports SDR improvement (dB).
Evaluates speech recognition models on their ability to accurately transcribe spoken audio into text (WER) and characters (LER). It probes the effectiveness of unsupervised pre-training on raw audio for downstream acoustic modeling and decoding. Use when the user wants to benchmark on TIMIT, WSJ, or asks about evaluating this task. Reports WER.
Evaluates Word Sense Induction (WSI) by clustering contextualized word embeddings to predict sense assignments. It probes a model's ability to capture lexical polysemy and contextual meaning without supervised sense labels, using natural corpus distributions rather than artificially skewed benchmarks. Use when the user wants to benchmark on SemCor, or asks about evaluating this task. Reports F-B^3.
Evaluates whole-slide image (WSI) classification performance using self-supervised patch representations and feature-space data augmentation. It probes how well distribution-guided representation learning captures discriminative histopathological patterns for diagnostic subtyping. Use when the user wants to benchmark on USTC-EGFR, TCGA-EGFR, TCGA-LUNG-3K, or asks about evaluating this task. Reports micro-average area under the curve (AUC).
Evaluates a model's ability to disambiguate word senses for both common nouns and proper nouns exhibiting regular polysemy. It probes contextual understanding and the capacity to leverage structured sense glosses and dot-object type classes to select the correct meaning from a candidate inventory. Use when the user wants to benchmark on WSD dataset (CWN 2.0), RP dataset (Revised Mandarin Chinese Dictionary), or asks about evaluating this task. Reports accuracy.
Evaluates the ability of unsupervised and supervised metrics to identify word-level translation errors by comparing their continuous scores against human-annotated error spans and multi-annotator agreement rates. Use when the user has predictions and gold and needs to compute Average Precision (AP).
Evaluates a model's ability to perform sequential recommendation by predicting the next item a user will interact with based on their chronological interaction history. It probes the model's capacity to capture temporal dynamics and collaborative filtering signals while ranking items against a full candidate set. Use when the user wants to benchmark on MovieLens-1M*, Amazon-Beauty, Amazon-Sports, LastFM (HetRec 2011), or asks about evaluating this task. Reports HR@10.
Evaluates embodied world models on conditional video generation from an initial image and text instruction. It probes instruction understanding, long-horizon planning, physical/causal reasoning, and temporal consistency in robotic interaction scenarios. Use when the user wants to benchmark on WoWBench, or asks about evaluating this task. Reports Planning Score ($S_{plan}$), Overall Benchmark Score.
Evaluates multimodal large language models' ability to perform real-world omni-modal understanding by jointly processing tightly coupled audio and video inputs. It probes complex temporal reasoning, cross-modal integration, and fine-grained perception across diverse everyday scenarios. Use when the user wants to benchmark on WorldSense, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal video understanding and long-chain reasoning by requiring models to integrate visual, auditory, and external world knowledge to answer open-ended and multiple-choice questions. Use when the user wants to benchmark on WorldQA, or asks about evaluating this task. Reports GPT-4 open-ended score.
This benchmark evaluates interactive Image-to-Video world models by measuring their ability to generate temporally coherent videos in response to standardized action commands. It probes three core capabilities: visual fidelity, precise camera/object control alignment, and long-horizon world consistency across different perspectives and visual styles. Use when the user wants to benchmark on WorldMark Image Suite, or asks about evaluating this task. Reports Aesthetic Quality.
Evaluates driving world models across five dimensions: generation quality, 3D/4D reconstruction coherence, action-following capability in closed-loop simulation, downstream perception task utility, and alignment with human preference. It probes geometric consistency, physical plausibility, and functional reliability of synthesized driving scenes. Use when the user wants to benchmark on WorldLens, or asks about evaluating this task. Reports Route Completion (%).
Evaluates an agent's ability to automate desktop and web GUI tasks from arbitrary starting states. It probes robustness to dynamic initial conditions, contextual variations, and multi-step interaction planning in real-world software environments. Use when the user wants to benchmark on WorldGUI, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates AI models on work-domain recommendation and NLP tasks, primarily focusing on ranking and retrieval scenarios such as occupation-to-skill matching, candidate recommendation, and skill/job normalization. It tests cross-lingual and multilingual retrieval capabilities over standardized occupational ontologies like ESCO. Use when the user wants to benchmark on ESCO Occupation-to-Skill, ESCO Skill-to-Occupation, Job Title Sim., SkillMatch-1K, Query-Candidate, Project-Candidate, JobBERT, M...
Evaluates the latency performance of AI workload allocation strategies across hierarchical cloud/edge/device computing environments for latency-sensitive medical ICU applications. It measures how effectively dynamic routing minimizes end-to-end response time when processing and transmission delays are factored in. Use when the user wants to benchmark on Edge AIBench ICU Applications (MIMIC-III derived), or asks about evaluating this task. Reports response time.
Evaluates whether automatically generated scientific workflow benchmarks accurately replicate the execution time and performance characteristics of real scientific workflows under varying hardware architectures and external memory loads. Use when the user wants to benchmark on Montage, 1000Genome, or asks about evaluating this task. Reports execution_time_ratio.
Evaluates web agents' ability to perform complex, knowledge-worker tasks on enterprise UIs (ServiceNow) and standard web benchmarks. It probes multimodal browser observation processing, large DOM navigation, and action execution in interactive environments. Use when the user wants to benchmark on WorkArena, MiniWoB, WebGum Subset, or asks about evaluating this task. Reports success rate.