Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,049–3,072 of 22,870 skills
This evaluation protocol assesses a multimodal fusion model's ability to classify emotional and stress states from wearable physiological signals. It probes the model's robustness across different data collection settings, subject variability, and varying label granularities (binary vs. multi-class affect). Use when the user wants to benchmark on WESAD, SWELL-KW, CASE, or asks about evaluating this task. Reports accuracy, macro-F1.
This evaluation probes a model's ability to perform named entity recognition under weak supervision, where training labels are noisy and derived from multiple distant supervision sources. It measures how well the model can denoise these labels and generalize entity patterns across general, biomedical, and review domains. Use when the user wants to benchmark on CoNLL 2003, LaptopReview, NCBI-Disease, BC5CDR, or asks about evaluating this task. Reports entity-level F1.
Evaluates inertial-based activity recognition models trained on weakly-supervised labels generated via vision foundation model clustering, benchmarked against fully-supervised and few-shot baselines. Use when the user wants to benchmark on WEAR, Wetlab, ActionSense, or asks about evaluating this task. Reports Acc.
Evaluates the visual mathematical reasoning capabilities of Large Multimodal Models (LMMs). It probes their ability to decompose composite problems, apply hierarchical knowledge concepts, and reason through multi-step visual math tasks without relying on rote memorization. Use when the user wants to benchmark on We-Math testmini, or asks about evaluating this task. Reports accuracy.
Evaluates the quality and diversity of multi-objective optimization algorithms for water distribution system design by comparing generated Pareto fronts against established benchmark fronts. It measures coverage of known solutions, discovery of novel non-dominated designs, and computational efficiency. Use when the user wants to benchmark on HAN, NYT, BLA, and GOY networks, or asks about evaluating this task. Reports N_A^u, N_B^u, N_A^a, N_B^a, N_c, N_{FE}.
Evaluates multi-document summarization systems on news event clusters by measuring how well generated summaries match human-written reference summaries. It probes the model's ability to extract or generate concise, informative summaries from highly redundant, large-scale document collections. Use when the user wants to benchmark on WCEP, or asks about evaluating this task. Reports ROUGE F1-score.
This benchmark evaluates 3D and 2D object detection, as well as multi-object tracking, for autonomous driving perception. It probes a model's ability to accurately localize vehicles and pedestrians using synchronized LiDAR and camera data, while measuring robustness to geographic domain shifts and varying training data scales. Use when the user wants to benchmark on Waymo Open Dataset, or asks about evaluating this task. Reports APH.
Evaluates a video generation model's capacity to synthesize physically coherent motion, high-fidelity visuals, and strict adherence to textual or image prompts across diverse scenarios. Use when the user wants to benchmark on Waver-Bench 1.0, Hermes Motion Testset, or asks about evaluating this task. Reports Human Preference Win Rate.
Evaluates the capability of autoregressive generative models to synthesize high-quality raw audio waveforms for speech and music, and to perform discriminative tasks like speech recognition directly on raw audio without intermediate feature extraction. Use when the user wants to benchmark on VCTK (CSTR Voice Cloning Toolkit), Google TTS (NA English & Mandarin), MagnaTagATune, YouTube Piano, TIMIT, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
Evaluates end-to-end time-domain audio source separation models on singing voice and multi-instrument separation tasks. Probes the model's ability to isolate specific audio sources from mixed recordings using raw waveform inputs. Use when the user wants to benchmark on MUSDB, CCMixter, or asks about evaluating this task. Reports MSE.
Evaluates the performance trade-offs of a WebAssembly in-place interpreter against state-of-the-art Wasm engines across translation time, translation space overhead, and execution time on a standard benchmark suite. Use when the user wants to benchmark on PolyBenchC-4.2.1 MEDIUM, or asks about evaluating this task. Reports translation_time.
Evaluates a model's ability to recommend relevant Web APIs for a given mashup application based on textual descriptions. It also probes multi-task learning capability through an auxiliary mashup category classification task. Use when the user wants to benchmark on ProgrammableWeb, or asks about evaluating this task. Reports Precision@N.
Evaluates instruction-following capabilities of LLMs in Thai across culture-aware, domain-specific (Medical, Law, Finance, Retail), and multitask settings. Probes factual accuracy, reasoning quality, and fluency in both zero-shot and fine-tuned regimes. Use when the user wants to benchmark on WangchanThaiInstruct, Thai LLM Leaderboard, Thai MT-Bench, or asks about evaluating this task. Reports Accuracy.
Evaluates the fidelity of geometrically grounded 3D reconstruction and novel view synthesis for open-world urban environments, and benchmarks the performance of embodied navigation policies trained and evaluated in these simulated environments. Use when the user wants to benchmark on Wanderland, or asks about evaluating this task. Reports Navigation Error (NE).
Evaluates the model's multimodal and text understanding capabilities across a suite of standard academic benchmarks covering visual question answering, reasoning, hallucination detection, and text-based reasoning. Use when the user wants to benchmark on MMMU, MMStar, MathVista, HalluBench, MMBench, OCRBench, AI2D, AIME, GPQA, HLE, LCBV6, or asks about evaluating this task. Reports average score.
Probes large language models' ability to recall and identify structural and grammatical properties of languages across diverse linguistic domains. It measures whether models have internalized typological facts from the World Atlas of Language Structures (WALS) by answering multiple-choice questions about specific language features. Use when the user wants to benchmark on WALS, WALS-100, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness and accuracy of TinyML person detection models across diverse demographic, environmental, and visual conditions. It benchmarks binary classification performance on large-scale, real-world image datasets tailored for resource-constrained devices. Use when the user wants to benchmark on Wake Vision, or asks about evaluating this task. Reports accuracy.
Evaluates models on detecting online abuse (personal attacks, aggression, toxicity) in conversational contexts reconstructed from Wikipedia talk pages. It probes the ability to classify individual messages as abusive or non-abusive while leveraging or ignoring conversational structure depending on the method. Use when the user wants to benchmark on WAC (Wikipedia Conversations Corpus), or asks about evaluating this task. Reports Macro F-measure.
Probes the ability of surrogate models to accurately predict FPGA resource usage (e.g., LUTs, BRAMs) and inference latency (clock cycles) for neural networks synthesized via hls4ml. It evaluates how well learned models approximate time-consuming hardware synthesis processes without running the full compilation pipeline. Use when the user wants to benchmark on wa-hls4ml benchmark, or asks about evaluating this task. Reports R².
Evaluates a model's ability to reconstruct the 4-momentum and mass of a W-boson from jet constituents, testing regression accuracy and physical consistency under truth-level and detector-simulated (Delphes) conditions. It compares equivariant neural networks against traditional physics-based taggers. Use when the user wants to benchmark on W-boson 4-momentum regression dataset, or asks about evaluating this task. Reports resolution ($\sigma_{p^{T}}$, $\sigma_{m}$, $\sigma_{\Delta R}$).
This benchmark evaluates multilingual sentence alignment and machine translation capabilities across 11 official South African languages using government-themed corpora. Use when the user wants to benchmark on Vuk'uzenzele, ZA-gov-multilingual, or asks about evaluating this task. Reports cosine similarity.
Evaluates Large Language Models' ability to customize and edit TikZ code based on a specified visual intent. The benchmark probes three core capabilities: locating relevant code features, synthesizing correct code variants, and validating that the modified code produces the intended visual output. Use when the user wants to benchmark on vTikZ, or asks about evaluating this task. Reports visual result validation.
Evaluates the ability of visual token compression methods to retain task-relevant visual information across five dimensions: global understanding, spatial and counting, reasoning and common sense, style and emotion, and local details. Use when the user wants to benchmark on VTCBench, or asks about evaluating this task. Reports accuracy.
This evaluation protocol assesses a vision-language model's ability to perform long-context mathematical and scientific reasoning using vision-text compression. It measures both reasoning accuracy across diverse benchmarks and computational efficiency in terms of token usage and inference latency. Use when the user wants to benchmark on GSM8K, MATH500, AIME25, AMC23, GPQA-Diamond, or asks about evaluating this task. Reports Accuracy (ACC).