Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,193–6,216 of 20,796 skills
Evaluates a model's ability to perform few-shot classification by learning to map a small support set of labeled examples to a classifier for unseen classes without fine-tuning. It probes rapid adaptation and generalization across vision and language modalities using an attention-based non-parametric memory mechanism. Use when the user wants to benchmark on Omniglot, ImageNet, miniImageNet, Penn Treebank, or asks about evaluating this task. Reports accuracy.
Evaluates the matching ability, correspondence sufficiency, and computational efficiency of local feature matchers across short- and wide-baseline image pairs. It measures how well matchers recover camera pose and how many correct correspondences they produce, enabling fair comparison for real-time applications like SLAM. Use when the user wants to benchmark on SfM/SLAM datasets (sequences 01-08), or asks about evaluating this task. Reports AUC (SP curve).
Evaluates the synthesis speed, intelligibility, and naturalness of a non-autoregressive TTS model trained with conditional flow matching on English speech. It measures how efficiently the model converts text to audio and how closely the output matches human perception of naturalness. Use when the user wants to benchmark on LJ Speech, or asks about evaluating this task. Reports MOS.
Evaluates machine learning models on predicting materials properties (e.g., elastic moduli, band gaps, formation energies) from crystal structures or compositions. It probes generalization across diverse data sizes, input types, and property domains using a standardized, pre-cleaned suite of 13 supervised tasks. Use when the user wants to benchmark on Matbench test suite v0.1, or asks about evaluating this task. Reports error estimation (RMSE/Accuracy).
Evaluates a model's ability to recommend functionally indispensable (must-cite) papers for a given query paper based only on its title and abstract. It probes scientific retrieval capability by measuring how well systems rank baseline, core-relevant, or frequently mentioned papers from a large candidate pool. Use when the user wants to benchmark on MasterSet-CoreML-v1, or asks about evaluating this task. Reports Recall@K.
Evaluates the ability of a deep reinforcement learning scheduler to allocate wireless resources (users to resource blocks) in massive MIMO networks. It probes the model's capacity to maximize spectral efficiency and user fairness under varying channel conditions (static vs. mobile) and network scales. Use when the user wants to benchmark on QuaDRiGa 3GPP_3D_UMi_LOS, or asks about evaluating this task. Reports normalized spectral efficiency.
Evaluates multilingual natural language understanding capabilities, specifically intent classification and slot filling, across 51 typologically diverse languages. It measures model robustness to different scripts, spacing conventions, and zero-shot cross-lingual transfer scenarios. Use when the user wants to benchmark on MASSIVE, or asks about evaluating this task. Reports exact match accuracy.
Evaluates a reference-less, masked language model-based metric's ability to predict human judgments on text summarization and simplification quality. It probes the model's capacity to capture multiple quality dimensions such as fluency, consistency, coherence, relevance, simplicity, and meaning preservation without relying on reference texts. Use when the user has predictions and gold and needs to compute pearson_correlation.
Evaluates large language models' domain-specific reasoning and numerical problem-solving capabilities in materials science and metallurgical engineering. It probes their ability to accurately answer multiple-choice, matching, and numerical questions, highlighting gaps in scientific reasoning and computational precision. Use when the user wants to benchmark on MaScQA, or asks about evaluating this task. Reports accuracy.
Evaluates named entity recognition (NER) capabilities across 20 typologically and geographically diverse African languages. It probes zero-shot cross-lingual transfer performance and measures how well models generalize to unseen entities and languages when fine-tuned on limited African language data. Use when the user wants to benchmark on MasakhaNER 2.0, or asks about evaluating this task. Reports F1.
Evaluates named entity recognition (NER) capabilities across ten African languages, probing models' ability to identify PER, ORG, and LOC entities in low-resource, morphologically complex, and culturally specific news text. Use when the user wants to benchmark on MasakhaNER, or asks about evaluating this task. Reports F1 score.
Evaluates multimodal large language models on multidimensional abstract visual reasoning and perceptual grounding. It probes the model's ability to recognize complex geometric and abstract patterns, track temporal/spatial changes, and perform multi-step visual reasoning across diverse puzzle configurations. Use when the user wants to benchmark on MARVEL, or asks about evaluating this task. Reports accuracy.
Evaluates an LLM's ability to refuse harmful or unsafe requests while maintaining helpfulness on benign prompts. It probes safety alignment through automatic reward-model scoring and human flagging of violations across in-distribution and out-of-domain benchmarks. Use when the user wants to benchmark on SafeEval, HelpEval, AlpacaEval, Anthropic Harmless, or asks about evaluating this task. Reports violation_rate.
Evaluates embodied agents and multimodal LLMs on long-horizon manipulation tasks in procedurally generated supermarket environments. Specifically, it probes spatial reasoning, occlusion handling, and collision avoidance during checkout unloading and in-aisle item collection. Use when the user wants to benchmark on MarketGen Benchmark, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates a model's ability to perform video question answering with varying levels of temporal reasoning complexity. It probes whether models can correctly link visual events in gameplay videos to answer questions that require single-frame, event-level, or multi-step causal/temporal understanding. Use when the user wants to benchmark on MarioQA, or asks about evaluating this task. Reports accuracy.
MarineEval probes the marine domain expertise and visual understanding capabilities of vision-language models. It evaluates tasks including species identification, spatial reasoning, ecological knowledge integration, and precise object localization under real-world marine conditions. Use when the user wants to benchmark on MarineEval, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of Large Vision-Language Models (LVLMs) to mitigate object hallucinations during text generation. It probes visual-text alignment by measuring hallucination rates, recall of existing objects, and accuracy on binary probing questions, alongside GPT-4V-aided assessments of response accuracy and detailness. Use when the user wants to benchmark on MSCOCO val2014, or asks about evaluating this task. Reports CHAIRS.
Evaluates a unified neural TTS framework's ability to disentangle speaker identity and emotional style, measuring speaker fidelity, emotional expressiveness, and overall speech quality in both English and Mandarin. Use when the user wants to benchmark on LibriTTS, AISHELL-3, CSEMOTIONS, or asks about evaluating this task. Reports Emotional expressiveness.
Evaluates molecular property prediction using explicit conformer ensembles versus single-conformer or 1D/2D baselines, probing how 3D structural flexibility and ensemble encoding strategies impact regression accuracy. Use when the user wants to benchmark on MARCEL, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
Evaluates pre-trained music audio representation models across a unified taxonomy of 18 downstream tasks spanning acoustic, performance, score, and high-level description levels. It assesses model generalization and representation quality under constrained training settings, including sequence labeling tasks like beat tracking and source separation. Use when the user wants to benchmark on MelodyDB, Muljam, Jamendo, GuitarSet, MUSDB18, NSynth, or asks about evaluating this task. Reports Accuracy.
Evaluates fine-grained spatial reasoning and pixel-accurate route tracing on commercial map images. Models must generate precise path coordinates or masks corresponding to text-based navigation queries. Use when the user wants to benchmark on MapTrace, Map-Bench, or asks about evaluating this task. Reports NDTW.
Evaluates the performance of explainable multi-robot motion planning algorithms (MAPS-X and Lazy MAPS-X) combined with sampling-based planners like RRT* across custom environments. It measures the trade-off between planning efficiency (runtime, success rate) and explainability (number of trajectory segments), highlighting how segmentation constraints impact computational cost and plan optimality. Use when the user wants to benchmark on MAPS-X custom environments, or asks about evaluating this...
Evaluates the performance and security robustness of agentic AI systems when operating in multilingual settings. It measures how task completion accuracy and vulnerability to adversarial prompts degrade or shift when instructions are translated from English into 11 typologically diverse languages. Use when the user wants to benchmark on GAIA, SWE-bench, MATH, ASB, or asks about evaluating this task. Reports accuracy.
Evaluates a robot's ability to navigate to a goal in unknown environments using only local sensor data without a pre-built map. It probes the policy's generalization across varying obstacle densities, passage widths, and real-world conditions, as well as its energy efficiency on neuromorphic hardware. Use when the user wants to benchmark on Gazebo Training Environments, Gazebo Test Environment, Real-world Office Environment, or asks about evaluating this task. Reports success rate.