Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,846
skills in category
952
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 2,7852,808 of 22,846 skills

X Webagentbench EvalA

Evaluates LLM-based agents' ability to comprehend multilingual shopping instructions and successfully navigate interactive web environments across 14 languages. Use when the user wants to benchmark on X-WebAgentBench, or asks about evaluating this task. Reports Task Score.

researchpythongo
0
3
X Vmamba Controllability EvalA

Probes the spatial feature flow and patch-level influence in Vision Mamba models using classical control theory. It quantifies how input image patches drive hidden state dynamics across hierarchical layers, revealing domain-specific diagnostic feature extraction patterns. Use when the user wants to benchmark on CMMD, DermaMNIST, BloodMNIST, or asks about evaluating this task. Reports influence score.

researchpythongo
0
3
X Topic EvalA

This benchmark evaluates multilingual topic classification on social media tweets across four languages (English, Spanish, Japanese, Greek). It probes models' ability to generalize across languages and training regimes, including zero-shot, few-shot, monolingual, cross-lingual, and multilingual fine-tuning settings. Use when the user wants to benchmark on X-Topic, or asks about evaluating this task. Reports macro-F1.

researchpythonperformance
0
3
X Pcr EvalA

Evaluates multi-modal large language models on progressive clinical reasoning in ophthalmic diagnosis. It tests the model's ability to perform a six-stage diagnostic chain (from image quality assessment to clinical decision-making) while integrating cross-modality imaging data and calibrating its uncertainty. Use when the user wants to benchmark on X-PCR, or asks about evaluating this task. Reports Stage-Wise Accuracy (SWA).

researchpythongo
0
3
X Omni EvalA

Evaluates the text rendering, text-to-image generation, and image understanding capabilities of a discrete autoregressive image generation model trained with reinforcement learning. It probes the model's ability to follow complex instructions, render long texts accurately, and generate high-fidelity images without relying on classifier-free guidance. Use when the user wants to benchmark on OneIG-Bench, LongText-Bench, DPG-Bench, GenEval, POPE, GQA, MMBench, SEEDBench-Img, DocVQA, OCRBench, or...

researchpythongo
0
3
X Mobility Nav EvalA

Evaluates an end-to-end navigation model's ability to predict robot dynamics and successfully navigate through structured and cluttered warehouse environments. It probes both open-loop trajectory and speed prediction accuracy, as well as closed-loop mission success, navigation efficiency, and motion smoothness in seen and out-of-distribution settings. Use when the user wants to benchmark on X-Mobility Warehouse Dataset, or asks about evaluating this task. Reports mission success rate (SR).

researchpythongo
0
3
Wximpactbench EvalA

Evaluates large language models' ability to understand and classify disruptive weather impacts from historical and modern newspaper articles, and to answer related questions by ranking relevant information. It specifically probes models' capacity to handle climate-related polysemy, extract nuanced societal responses, and correctly identify passages without weather impacts. Use when the user wants to benchmark on WXImpactBench, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Wsj0 Speech Separation EvalA

This evaluation probes a model's ability to perform single-channel speech separation by learning discriminative time-frequency embeddings that group mixture components into distinct speaker clusters. It specifically tests generalization to unseen speakers and scaling to three-speaker mixtures without retraining. Use when the user wants to benchmark on WSJ0-based Speech Mixtures, or asks about evaluating this task. Reports SDR improvement (dB).

researchpythonperformance
0
3
Wsj Timit Speech Recognition EvalA

Evaluates speech recognition models on their ability to accurately transcribe spoken audio into text (WER) and characters (LER). It probes the effectiveness of unsupervised pre-training on raw audio for downstream acoustic modeling and decoding. Use when the user wants to benchmark on TIMIT, WSJ, or asks about evaluating this task. Reports WER.

researchpythonexpress
0
3
Wsi Semcor EvalA

Evaluates Word Sense Induction (WSI) by clustering contextualized word embeddings to predict sense assignments. It probes a model's ability to capture lexical polysemy and contextual meaning without supervised sense labels, using natural corpus distributions rather than artificially skewed benchmarks. Use when the user wants to benchmark on SemCor, or asks about evaluating this task. Reports F-B^3.

researchpythongo
0
3
Wsi Classification EvalA

Evaluates whole-slide image (WSI) classification performance using self-supervised patch representations and feature-space data augmentation. It probes how well distribution-guided representation learning captures discriminative histopathological patterns for diagnostic subtyping. Use when the user wants to benchmark on USTC-EGFR, TCGA-EGFR, TCGA-LUNG-3K, or asks about evaluating this task. Reports micro-average area under the curve (AUC).

researchpythonapi
0
3
Wsd Rp Accuracy EvalA

Evaluates a model's ability to disambiguate word senses for both common nouns and proper nouns exhibiting regular polysemy. It probes contextual understanding and the capacity to leverage structured sense glosses and dot-object type classes to select the correct meaning from a candidate inventory. Use when the user wants to benchmark on WSD dataset (CWN 2.0), RP dataset (Revised Mandarin Chinese Dictionary), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Wqe Metric EvalA

Evaluates the ability of unsupervised and supervised metrics to identify word-level translation errors by comparing their continuous scores against human-annotated error spans and multi-annotator agreement rates. Use when the user has predictions and gold and needs to compute Average Precision (AP).

researchpythongo
0
3
Wpgrec EvalA

Evaluates a model's ability to perform sequential recommendation by predicting the next item a user will interact with based on their chronological interaction history. It probes the model's capacity to capture temporal dynamics and collaborative filtering signals while ranking items against a full candidate set. Use when the user wants to benchmark on MovieLens-1M*, Amazon-Beauty, Amazon-Sports, LastFM (HetRec 2011), or asks about evaluating this task. Reports HR@10.

researchpython
0
3
Wowbench EvalA

Evaluates embodied world models on conditional video generation from an initial image and text instruction. It probes instruction understanding, long-horizon planning, physical/causal reasoning, and temporal consistency in robotic interaction scenarios. Use when the user wants to benchmark on WoWBench, or asks about evaluating this task. Reports Planning Score ($S_{plan}$), Overall Benchmark Score.

researchpythonnode
0
3
Worldsense EvalA

Evaluates multimodal large language models' ability to perform real-world omni-modal understanding by jointly processing tightly coupled audio and video inputs. It probes complex temporal reasoning, cross-modal integration, and fine-grained perception across diverse everyday scenarios. Use when the user wants to benchmark on WorldSense, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Worldqa EvalA

Evaluates multimodal video understanding and long-chain reasoning by requiring models to integrate visual, auditory, and external world knowledge to answer open-ended and multiple-choice questions. Use when the user wants to benchmark on WorldQA, or asks about evaluating this task. Reports GPT-4 open-ended score.

researchpythongo
0
3
Worldmark EvalA

This benchmark evaluates interactive Image-to-Video world models by measuring their ability to generate temporally coherent videos in response to standardized action commands. It probes three core capabilities: visual fidelity, precise camera/object control alignment, and long-horizon world consistency across different perspectives and visual styles. Use when the user wants to benchmark on WorldMark Image Suite, or asks about evaluating this task. Reports Aesthetic Quality.

researchpythonangular
0
3
Worldlens EvalA

Evaluates driving world models across five dimensions: generation quality, 3D/4D reconstruction coherence, action-following capability in closed-loop simulation, downstream perception task utility, and alignment with human preference. It probes geometric consistency, physical plausibility, and functional reliability of synthesized driving scenes. Use when the user wants to benchmark on WorldLens, or asks about evaluating this task. Reports Route Completion (%).

researchpythongo
0
3
Worldgui EvalA

Evaluates an agent's ability to automate desktop and web GUI tasks from arbitrary starting states. It probes robustness to dynamic initial conditions, contextual variations, and multi-step interaction planning in real-world software environments. Use when the user wants to benchmark on WorldGUI, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Workrb EvalA

Evaluates AI models on work-domain recommendation and NLP tasks, primarily focusing on ranking and retrieval scenarios such as occupation-to-skill matching, candidate recommendation, and skill/job normalization. It tests cross-lingual and multilingual retrieval capabilities over standardized occupational ontologies like ESCO. Use when the user wants to benchmark on ESCO Occupation-to-Skill, ESCO Skill-to-Occupation, Job Title Sim., SkillMatch-1K, Query-Candidate, Project-Candidate, JobBERT, M...

researchpythongo
0
3
Workload Allocation EvalA

Evaluates the latency performance of AI workload allocation strategies across hierarchical cloud/edge/device computing environments for latency-sensitive medical ICU applications. It measures how effectively dynamic routing minimizes end-to-end response time when processing and transmission delays are factored in. Use when the user wants to benchmark on Edge AIBench ICU Applications (MIMIC-III derived), or asks about evaluating this task. Reports response time.

researchpythongo
0
3
Workflow Benchmark Accuracy EvalA

Evaluates whether automatically generated scientific workflow benchmarks accurately replicate the execution time and performance characteristics of real scientific workflows under varying hardware architectures and external memory loads. Use when the user wants to benchmark on Montage, 1000Genome, or asks about evaluating this task. Reports execution_time_ratio.

researchpythonnode
0
3
Workarena EvalA

Evaluates web agents' ability to perform complex, knowledge-worker tasks on enterprise UIs (ServiceNow) and standard web benchmarks. It probes multimodal browser observation processing, large DOM navigation, and action execution in interactive environments. Use when the user wants to benchmark on WorkArena, MiniWoB, WebGum Subset, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3