Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,870
skills in category
953
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,0253,048 of 22,870 skills

Wenetspeech EvalA

Evaluates Mandarin speech recognition systems across diverse, real-world domains (internet, meetings) and mixed Mandarin-English content. It benchmarks the robustness of ASR models against noisy, production-level audio and varying data scales. Use when the user wants to benchmark on WenetSpeech, or asks about evaluating this task. Reports MER%.

researchpythonshell
0
3
Weedsense EvalA

Evaluates a model's ability to jointly perform fine-grained weed species segmentation, continuous plant height regression, and discrete temporal growth stage classification from single RGB images. Use when the user wants to benchmark on WeedSense, or asks about evaluating this task. Reports mIoU, MAE, Accuracy.

researchpythongo
0
3
Webwalker EvalA

Evaluates LLM-based agents' ability to systematically navigate multi-layered, real-world websites to extract buried information. It probes long-range reasoning, memory management, and step-by-step click-based navigation under strict action limits. Use when the user wants to benchmark on WebWalkerQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Webvoyager EvalA

Measures an autonomous web agent's end-to-end task completion capability across dynamic, real-world websites. It evaluates multi-step navigation, form filling, and robustness in open-web environments. Use when the user wants to benchmark on WebVoyager, or asks about evaluating this task. Reports Success Rate.

researchpython
0
3
Webuot 1m EvalA

Evaluates the robustness and accuracy of deep object trackers in challenging underwater environments. It probes cross-domain adaptation from open-air to underwater domains, as well as within-domain fine-tuning capabilities, while also assessing performance under varying frame rates and complex visual conditions like occlusion and low visibility. Use when the user wants to benchmark on WebUOT-1M, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Webuav 3m EvalA

Evaluates the robustness and accuracy of deep visual tracking algorithms on large-scale, real-world UAV video sequences. It probes how well trackers handle diverse environmental conditions, motion dynamics, and target appearance changes without parameter tuning. Use when the user wants to benchmark on WebUAV-3M, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Webmainbench EvalA

Evaluates a model's ability to extract main content from HTML web pages by classifying semantic blocks and generating clean text or Markdown. It probes robustness across varying difficulty levels and rich content types such as tables, code, and equations. Use when the user wants to benchmark on WebMainBench, WCEB, or asks about evaluating this task. Reports ROUGE-N F1.

researchpythongo
0
3
Webgym EvalA

This benchmark evaluates visual web agents on their ability to navigate complex, multi-step tasks across diverse websites and domains. It specifically probes long-horizon interaction, information extraction, and out-of-distribution generalization to unseen websites. Use when the user wants to benchmark on WebGym, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Webgen Bench EvalA

Evaluates a model's ability to generate functional and visually accurate website codebases from natural language instructions. It measures both functional correctness via automated GUI-agent testing and visual fidelity via VLM-based appearance scoring. Use when the user wants to benchmark on WebGen-Bench, or asks about evaluating this task. Reports Accuracy.

researchjavascriptpython
0
3
Webforge Bench EvalA

Evaluates browser agents' ability to complete interactive web tasks across varying difficulty levels and domains. It probes navigation, visual understanding, multi-step reasoning, and interaction capabilities in realistic, self-contained web environments. Use when the user wants to benchmark on WebForge-Bench, or asks about evaluating this task. Reports accuracy (%).

researchpythongo
0
3
Webfg Webinat EvalA

Evaluates fine-grained image classification performance on large-scale webly supervised datasets characterized by label noise and extreme class imbalance. It probes a model's ability to learn robust visual features and correct noisy labels through mutual peer learning or standard fine-tuning. Use when the user wants to benchmark on WebFG-496, WebiNat-5089, or asks about evaluating this task. Reports classification accuracy (%).

researchpythongo
0
3
Webfaq Retrieval EvalA

Evaluates multilingual dense retrieval models on natural Q&A pairs by measuring ranking quality against gold answers. It tests the model's ability to retrieve relevant FAQ documents across multiple languages and assesses zero-shot generalization to other Wikipedia-based benchmarks. Use when the user wants to benchmark on WebFAQ, Mr. TyDi, MIRACL (Hard Negatives), or asks about evaluating this task. Reports NDCG@10.

researchpythongo
0
3
Webcoderbench EvalA

Evaluates LLMs' ability to generate complete web applications from real-world user requirements. It probes multi-modal understanding, code generation quality, and strict adherence to ground-truth checklists across functionality, visual design, and content dimensions. Use when the user wants to benchmark on WebCoderBench, or asks about evaluating this task. Reports checklist-based evaluation.

researchpythongo
0
3
Webcode2m EvalA

This benchmark evaluates multimodal models' ability to translate webpage design screenshots into functional HTML/CSS code. It probes visual fidelity, structural hierarchy recall, and the capacity to generate long, complex, real-world front-end code from visual inputs. Use when the user wants to benchmark on WebCode2M, or asks about evaluating this task. Reports TreeBLEU.

researchpythongo
0
3
Webarena Human Trajectory EvalA

Evaluates the planning and execution quality of LLM-based web agents by comparing their action trajectories against human-demonstrated gold standards. It measures recovery from deviations, action repetitiveness, step fulfillment, partial task completion, and alignment between planned and executed actions. Use when the user wants to benchmark on WebArena Human Trajectory Dataset, or asks about evaluating this task. Reports Recovery Rate.

researchpythongo
0
3
Weavebench EvalA

Evaluates multi-turn, context-aware image comprehension and generation in an interleaved setting. It probes a model's ability to maintain visual consistency, follow iterative editing instructions, and integrate historical context across multiple turns. Use when the user wants to benchmark on WEAVEBench, or asks about evaluating this task. Reports WEAVEBench.

researchpythongo
0
3
Weatherqa EvalA

Evaluates multimodal reasoning capabilities in the meteorological domain, specifically testing a model's ability to interpret weather maps and answer domain-specific multiple-choice questions. It also measures cross-task generalization and the logical consistency of the model's reasoning chains versus final answers. Use when the user wants to benchmark on WeatherQA, ScienceQA, or asks about evaluating this task. Reports Multiple-choice accuracy.

researchpythongo
0
3
Weatherdiffusion EvalA

Evaluates the capability of diffusion models to perform controllable weather editing in intrinsic space. It probes whether the model can preserve geometric and material consistency while synthesizing realistic weather effects like rain, snow, and fog. Use when the user wants to benchmark on WeatherSynthetic, ACDC, TransWeather, Waymo, or asks about evaluating this task. Reports PickScore.

researchpythongo
0
3
Weatherbench2 EvalA

Evaluates the capability of generative weather forecasting models to predict global atmospheric and surface conditions from mid-range to sub-seasonal horizons (up to 30 days). It probes deterministic accuracy and probabilistic ensemble calibration against established meteorological baselines and climatology. Use when the user wants to benchmark on WeatherBench-2, or asks about evaluating this task. Reports Latitude-weighted RMSE.

researchpythonperformance
0
3
Weatherbench Probability EvalA

Evaluates the accuracy and reliability of probabilistic medium-range weather forecasting models against operational ensemble baselines. It probes how well deep learning methods capture uncertainty, calibration, and sharpness for key atmospheric variables. Use when the user wants to benchmark on WeatherBench Probability, or asks about evaluating this task. Reports CRPS.

researchpythongo
0
3
Weather2k EvalA

Evaluates the ability of deep learning models to forecast multivariate and univariate meteorological factors (temperature, visibility, humidity) using historical time-series and spatio-temporal data from ground weather stations. Use when the user wants to benchmark on Weather2K, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3
Weather Robustness Od EvalA

Evaluates object detection robustness to adverse weather by measuring performance degradation when models trained on clear-weather datasets are tested on weather-corrupted images. It specifically quantifies dataset bias by comparing baseline in-distribution performance against out-of-distribution performance on the DAWN dataset. Use when the user wants to benchmark on DAWN, Pascal VOC 2012, Microsoft COCO 2017, or asks about evaluating this task. Reports mAP.

researchpythontesting
0
3
Weather 5k EvalA

Evaluates the capability of data-driven time-series forecasting models to predict global meteorological variables over short to long horizons, and assesses their robustness in forecasting extreme weather events compared to numerical weather prediction baselines. Use when the user wants to benchmark on WEATHER-5K, or asks about evaluating this task. Reports MAE.

researchpythongit
0
3
Wearvox EvalA

Evaluates speech-language models on egocentric, multi-channel wearable audio tasks, probing their ability to handle noisy real-world acoustic conditions, reject side-talk, execute tool calls, answer questions with or without context, and translate speech in conversational settings. Use when the user wants to benchmark on WearVox, or asks about evaluating this task. Reports Turn-basedMicro-avg.

researchpythongo
0
3