Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,921–7,944 of 20,853 skills
Evaluates aerial vehicle detection capability using aligned RGB and infrared image pairs. It specifically probes a model's ability to fuse cross-modal features and handle uncertainty in low-light or complex urban backgrounds. Use when the user wants to benchmark on DroneVehicle, or asks about evaluating this task. Reports mAP.
Evaluates the sustainability and robustness of a dynamic behavioral profiling approach (DroidSpan) for Android malware detection over time and against code obfuscation, compared to a static baseline (MamaDroid). Use when the user wants to benchmark on all-data, oldBen+oldMal, MalObf, or asks about evaluating this task. Reports F1-measure.
Evaluates the robustness and generalization of robot manipulation policies when co-trained with diverse, in-the-wild data. It probes the model's ability to handle distractors, novel objects, and scene variations across short, medium, and long-horizon tasks. Use when the user wants to benchmark on DROID Evaluation Tasks, or asks about evaluating this task. Reports success rate.
Evaluates a vision-language model's ability to perform multi-label multiple-choice question answering on real-world driving scenarios, requiring precise visual grounding and spatial reasoning to select all correct answers from a set of options. Use when the user wants to benchmark on DrivingVQA, or asks about evaluating this task. Reports exam score.
Evaluates generative video world models for autonomous driving by jointly assessing visual realism, trajectory plausibility, temporal and agent-level consistency, and ego-conditioned motion controllability over a 100-frame prediction horizon. It benchmarks both general-purpose and driving-specific models to reveal trade-offs between photorealism and physical motion fidelity. Use when the user wants to benchmark on DrivingGen, or asks about evaluating this task. Reports Avg. Rank.
Evaluates the generalization capability of reinforcement learning policies for autonomous driving across procedurally generated traffic scenarios. It probes how well agents trained on a fixed set of road layouts and traffic dynamics perform when transferred to unseen environments with varying vehicle interactions and partial observability. Use when the user wants to benchmark on Driver Dojo, or asks about evaluating this task. Reports Interquartile Mean (IQM) reward.
Evaluates a model's ability to judge autonomous driving trajectory pairs based on context-aware reasoning, safety, and human preferences, rather than relying on rigid rule-based thresholds. It probes whether the model can integrate visual and symbolic context to reason about nuanced traffic situations like lateral buffer maintenance or stop sign compliance. Use when the user wants to benchmark on DriveCritic, or asks about evaluating this task. Reports accuracy.
Evaluates the reliability, visual grounding, and corruption resilience of vision-language models in autonomous driving. It probes whether models genuinely interpret degraded visual inputs or rely on textual priors and hallucinated reasoning when visual cues are missing or corrupted. Use when the user wants to benchmark on DriveBench, or asks about evaluating this task. Reports GPT score.
Evaluates fine-grained driver action recognition in constrained in-cabin environments using multimodal video inputs (RGB, IR, Depth). It probes the model's ability to classify 34 specific driver activities under variable illumination and occlusion by measuring both overall and per-class recognition accuracy. Use when the user wants to benchmark on Drive&Act, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates a session-based recommendation model's ability to predict the next item a user will purchase based on their recent browsing history. It specifically probes how well the model handles cold-start scenarios and varying data availability by measuring ranking quality and hit rates on short retail sessions. Use when the user wants to benchmark on Dressipi, or asks about evaluating this task. Reports Recall@20.
Evaluates a model's ability to perform multimodal instruction-based image editing and generation, specifically testing adherence to text instructions while manipulating concrete objects and abstract attributes (e.g., texture, style) using multiple reference images. Use when the user wants to benchmark on DreamOmni2 benchmark, or asks about evaluating this task. Reports success editing ratio.
Evaluates a diffusion model's ability to generate images that preserve subject identity from reference images while adhering to text prompts, covering both single-subject and multi-subject scenarios. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports DINO.
Evaluates subject-driven image generation by measuring how well the model follows text instructions and preserves the reference subject from the source image. It tests the model's ability to extract and reuse specific objects without fine-tuning. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports CLIP-T.
Evaluates a text-to-image model's ability to customize generated images with specific subjects from reference images, both individually and in combination, while maintaining alignment with text prompts. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports CLIP-I.
Evaluates large language models' fluid intelligence and abstract rule generalization across four hierarchical cognitive levels (Attribute, Spatial, Sequential, Conceptual). It probes the model's ability to dynamically adapt to varying task complexity and apply learned rules to novel grid-based reasoning problems. Use when the user wants to benchmark on DRE-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform span-based machine reading comprehension in traditional Chinese. It probes factual retrieval and exact answer extraction from provided context paragraphs without requiring complex inference or multiple-choice reasoning. Use when the user wants to benchmark on DRCD, or asks about evaluating this task. Reports F1 score.
Evaluates generative models' ability to perform clinical diagnostic reasoning, including medical knowledge representation, evidence synthesis, and diagnosis generation. The benchmark spans sentence-level to full-note tasks to probe abstractive reasoning and clinical knowledge inference. Use when the user wants to benchmark on DR.BENCH, or asks about evaluating this task. Reports accuracy.
Evaluates evidence-grounded visual reasoning in diagrams by requiring models to localize bounding boxes of supporting visual elements (e.g., labels, axes, connectors) rather than just predicting the correct answer. It probes a model's ability to visually ground reasoning steps within complex diagrammatic contexts and assesses how well localization quality correlates with reasoning performance. Use when the user wants to benchmark on DRAGON, or asks about evaluating this task. Reports Max Pair...
Evaluates the robustness of text-to-SQL models against semantic-preserving and semantic-changing perturbations across database schemas, natural language questions, and SQL queries. It measures how well models maintain execution accuracy when inputs are altered, revealing architectural vulnerabilities in entity linking, decoder design, and value prediction. Use when the user wants to benchmark on Dr.Spider, or asks about evaluating this task. Reports execution accuracy (EX).
Evaluates an automated framework's ability to encode natural-language data-governance policies into a formal model, extract data-flow graphs from provenance traces, and correctly trigger compliance obligations across decentralized scientific workflows. Use when the user wants to benchmark on Cyclone tracking workflow, MT3D (Moment Tensor in 3D) workflow, Real-world data-governance policies, or asks about evaluating this task. Reports actioning rules.
Evaluates a model's ability to retrieve relevant passages from a large unstructured corpus for open-domain question answering. It probes semantic matching and dense retrieval capabilities by measuring how often the correct answer span appears in the top-k retrieved passages. Use when the user wants to benchmark on Natural Questions, TriviaQA, WebQuestions, CuratedTREC, SQuAD v1.1, or asks about evaluating this task. Reports top-k retrieval accuracy.
This evaluation protocol assesses a language model's ability to align with human preferences across open-ended text generation tasks. It measures how well the model optimizes a reward objective while staying close to a reference policy, and evaluates practical performance via pairwise win rates against baselines. Use when the user wants to benchmark on IMDb, Reddit TL;DR, Anthropic HH, or asks about evaluating this task. Reports win rate.
This protocol evaluates language models across factual knowledge, mathematical reasoning, instruction following, code generation, truthfulness, and safety/refusal capabilities. It uses standardized benchmarks to measure how preference optimization methods and data quality impact model performance. Use when the user wants to benchmark on MMLU, GSM8k, Big Bench Hard, TruthfulQA, AlpacaEval, IFEval, HumanEval+, MBPP+, ToxiGen, XSTest, or asks about evaluating this task. Reports average accuracy.
Evaluates language models on natural language understanding, commonsense reasoning, and reading comprehension to measure alignment quality and reasoning preservation. It compares standard log-probability predictions against scores derived from a learned Direct Preference Head (DPH) reward model to assess self-evaluation capabilities. Use when the user wants to benchmark on GLUE, RACE, ARC, OpenBookQA, HellaSwag, WinoGrande, BoolQ, PIQA, or asks about evaluating this task. Reports accuracy.