Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,025–3,048 of 22,870 skills
Evaluates Mandarin speech recognition systems across diverse, real-world domains (internet, meetings) and mixed Mandarin-English content. It benchmarks the robustness of ASR models against noisy, production-level audio and varying data scales. Use when the user wants to benchmark on WenetSpeech, or asks about evaluating this task. Reports MER%.
Evaluates a model's ability to jointly perform fine-grained weed species segmentation, continuous plant height regression, and discrete temporal growth stage classification from single RGB images. Use when the user wants to benchmark on WeedSense, or asks about evaluating this task. Reports mIoU, MAE, Accuracy.
Evaluates LLM-based agents' ability to systematically navigate multi-layered, real-world websites to extract buried information. It probes long-range reasoning, memory management, and step-by-step click-based navigation under strict action limits. Use when the user wants to benchmark on WebWalkerQA, or asks about evaluating this task. Reports accuracy.
Measures an autonomous web agent's end-to-end task completion capability across dynamic, real-world websites. It evaluates multi-step navigation, form filling, and robustness in open-web environments. Use when the user wants to benchmark on WebVoyager, or asks about evaluating this task. Reports Success Rate.
Evaluates the robustness and accuracy of deep object trackers in challenging underwater environments. It probes cross-domain adaptation from open-air to underwater domains, as well as within-domain fine-tuning capabilities, while also assessing performance under varying frame rates and complex visual conditions like occlusion and low visibility. Use when the user wants to benchmark on WebUOT-1M, or asks about evaluating this task. Reports AUC.
Evaluates the robustness and accuracy of deep visual tracking algorithms on large-scale, real-world UAV video sequences. It probes how well trackers handle diverse environmental conditions, motion dynamics, and target appearance changes without parameter tuning. Use when the user wants to benchmark on WebUAV-3M, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to extract main content from HTML web pages by classifying semantic blocks and generating clean text or Markdown. It probes robustness across varying difficulty levels and rich content types such as tables, code, and equations. Use when the user wants to benchmark on WebMainBench, WCEB, or asks about evaluating this task. Reports ROUGE-N F1.
This benchmark evaluates visual web agents on their ability to navigate complex, multi-step tasks across diverse websites and domains. It specifically probes long-horizon interaction, information extraction, and out-of-distribution generalization to unseen websites. Use when the user wants to benchmark on WebGym, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to generate functional and visually accurate website codebases from natural language instructions. It measures both functional correctness via automated GUI-agent testing and visual fidelity via VLM-based appearance scoring. Use when the user wants to benchmark on WebGen-Bench, or asks about evaluating this task. Reports Accuracy.
Evaluates browser agents' ability to complete interactive web tasks across varying difficulty levels and domains. It probes navigation, visual understanding, multi-step reasoning, and interaction capabilities in realistic, self-contained web environments. Use when the user wants to benchmark on WebForge-Bench, or asks about evaluating this task. Reports accuracy (%).
Evaluates fine-grained image classification performance on large-scale webly supervised datasets characterized by label noise and extreme class imbalance. It probes a model's ability to learn robust visual features and correct noisy labels through mutual peer learning or standard fine-tuning. Use when the user wants to benchmark on WebFG-496, WebiNat-5089, or asks about evaluating this task. Reports classification accuracy (%).
Evaluates multilingual dense retrieval models on natural Q&A pairs by measuring ranking quality against gold answers. It tests the model's ability to retrieve relevant FAQ documents across multiple languages and assesses zero-shot generalization to other Wikipedia-based benchmarks. Use when the user wants to benchmark on WebFAQ, Mr. TyDi, MIRACL (Hard Negatives), or asks about evaluating this task. Reports NDCG@10.
Evaluates LLMs' ability to generate complete web applications from real-world user requirements. It probes multi-modal understanding, code generation quality, and strict adherence to ground-truth checklists across functionality, visual design, and content dimensions. Use when the user wants to benchmark on WebCoderBench, or asks about evaluating this task. Reports checklist-based evaluation.
This benchmark evaluates multimodal models' ability to translate webpage design screenshots into functional HTML/CSS code. It probes visual fidelity, structural hierarchy recall, and the capacity to generate long, complex, real-world front-end code from visual inputs. Use when the user wants to benchmark on WebCode2M, or asks about evaluating this task. Reports TreeBLEU.
Evaluates the planning and execution quality of LLM-based web agents by comparing their action trajectories against human-demonstrated gold standards. It measures recovery from deviations, action repetitiveness, step fulfillment, partial task completion, and alignment between planned and executed actions. Use when the user wants to benchmark on WebArena Human Trajectory Dataset, or asks about evaluating this task. Reports Recovery Rate.
Evaluates multi-turn, context-aware image comprehension and generation in an interleaved setting. It probes a model's ability to maintain visual consistency, follow iterative editing instructions, and integrate historical context across multiple turns. Use when the user wants to benchmark on WEAVEBench, or asks about evaluating this task. Reports WEAVEBench.
Evaluates multimodal reasoning capabilities in the meteorological domain, specifically testing a model's ability to interpret weather maps and answer domain-specific multiple-choice questions. It also measures cross-task generalization and the logical consistency of the model's reasoning chains versus final answers. Use when the user wants to benchmark on WeatherQA, ScienceQA, or asks about evaluating this task. Reports Multiple-choice accuracy.
Evaluates the capability of diffusion models to perform controllable weather editing in intrinsic space. It probes whether the model can preserve geometric and material consistency while synthesizing realistic weather effects like rain, snow, and fog. Use when the user wants to benchmark on WeatherSynthetic, ACDC, TransWeather, Waymo, or asks about evaluating this task. Reports PickScore.
Evaluates the capability of generative weather forecasting models to predict global atmospheric and surface conditions from mid-range to sub-seasonal horizons (up to 30 days). It probes deterministic accuracy and probabilistic ensemble calibration against established meteorological baselines and climatology. Use when the user wants to benchmark on WeatherBench-2, or asks about evaluating this task. Reports Latitude-weighted RMSE.
Evaluates the accuracy and reliability of probabilistic medium-range weather forecasting models against operational ensemble baselines. It probes how well deep learning methods capture uncertainty, calibration, and sharpness for key atmospheric variables. Use when the user wants to benchmark on WeatherBench Probability, or asks about evaluating this task. Reports CRPS.
Evaluates the ability of deep learning models to forecast multivariate and univariate meteorological factors (temperature, visibility, humidity) using historical time-series and spatio-temporal data from ground weather stations. Use when the user wants to benchmark on Weather2K, or asks about evaluating this task. Reports MAE.
Evaluates object detection robustness to adverse weather by measuring performance degradation when models trained on clear-weather datasets are tested on weather-corrupted images. It specifically quantifies dataset bias by comparing baseline in-distribution performance against out-of-distribution performance on the DAWN dataset. Use when the user wants to benchmark on DAWN, Pascal VOC 2012, Microsoft COCO 2017, or asks about evaluating this task. Reports mAP.
Evaluates the capability of data-driven time-series forecasting models to predict global meteorological variables over short to long horizons, and assesses their robustness in forecasting extreme weather events compared to numerical weather prediction baselines. Use when the user wants to benchmark on WEATHER-5K, or asks about evaluating this task. Reports MAE.
Evaluates speech-language models on egocentric, multi-channel wearable audio tasks, probing their ability to handle noisy real-world acoustic conditions, reject side-talk, execute tool calls, answer questions with or without context, and translate speech in conversational settings. Use when the user wants to benchmark on WearVox, or asks about evaluating this task. Reports Turn-basedMicro-avg.