Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,137–8,160 of 20,861 skills
Evaluates the ability of computer vision models to perform multi-label semantic segmentation for identifying and localizing various reinforced concrete defects on bridge infrastructure. It probes pixel-level classification accuracy across diverse, real-world damage types and structural components. Use when the user wants to benchmark on dacl10k, or asks about evaluating this task. Reports mean IoU.
This benchmark evaluates AI-based data assimilation models for weather forecasting. It tests their ability to integrate sparse or noisy observations with background atmospheric fields to produce accurate analysis fields, which are then used to initialize skillful medium-range weather predictions. Use when the user wants to benchmark on ERA5, GDAS, or asks about evaluating this task. Reports RMSE.
Evaluates a model's ability to predict driver maneuvers and generate hierarchical, human-understandable explanations from ego-centric driving videos. It probes spatio-temporal feature disentanglement and the impact of gaze modality on interpretability. Use when the user wants to benchmark on DAAD-X, or asks about evaluating this task. Reports Acc.
Evaluates a diffusion-based reinforcement learning agent's capability to dynamically assign AI-generated content tasks to edge service providers under stochastic workloads, optimizing for user utility while preventing system crashes and minimizing training time. Use when the user wants to benchmark on Custom AIGC Edge Simulation Environment, Gym Benchmark Tasks, or asks about evaluating this task. Reports Cumulative Reward.
Evaluates enterprise network vulnerability to lateral attacks by simulating adversarial movement across authentication graphs. It also measures the effectiveness of defense strategies in predicting attacker movement based on graph topology and credential hygiene levels. Use when the user wants to benchmark on G_s, G_l, G_lanl, or asks about evaluating this task. Reports Network Vulnerability.
Evaluates LLMs' vulnerability to deceptive reasoning and jailbreak attacks by measuring how well models align their internal chain-of-thought with malicious instructions while producing benign final outputs. It probes detection evasion, output camouflage, and internal malicious reasoning under adversarial system prompt injections. Use when the user wants to benchmark on D-REX, or asks about evaluating this task. Reports Target-Specific Success (%).
Evaluates a model's ability to detect and quantify the degree of replication between an original image and a diffusion-generated replica. It probes continuous replication level prediction rather than binary copy detection, measuring how well predicted scores align with manually annotated replication levels. Use when the user wants to benchmark on D-Rep, or asks about evaluating this task. Reports PCC.
Evaluates NLP models on four challenging text classification tasks using a large-scale Czech news dataset: identifying the news source, article category, inferred author gender, and publication day of the week. The benchmark tests a model's ability to capture deep textual dependencies and contextual cues beyond simple keyword matching. Use when the user wants to benchmark on CZE-NEC, or asks about evaluating this task. Reports F1 Macro.
Evaluates LLMs' ability to generate precise Cypher queries from natural language questions over large-scale property graphs. It probes complex graph retrieval capabilities including multi-hop reasoning, temporal constraints, aggregations, and strict schema adherence. Use when the user wants to benchmark on CypherBench, or asks about evaluating this task. Reports EX.
This benchmark evaluates the computational performance and scaling behavior of the cygrid gridding module. It measures processing time and parallelization efficiency across varying input sample sizes, field dimensions, and core counts. Use when the user has predictions and gold and needs to compute processing_time.
Evaluates a model's ability to generate scene graphs from aerial video sequences by predicting object relationships and interactions. It specifically probes long-range temporal dependency modeling and periodic interaction recognition in drone-captured footage across predicate classification, scene graph classification, and scene graph detection tasks. Use when the user wants to benchmark on AeroEye, PVSG, ASPIRe, or asks about evaluating this task. Reports mean Recall@K (mR@K).
Evaluates the ability of generative models to perform unpaired image-to-image translation while preserving structural integrity and achieving perceptual realism. Probes domain mapping capabilities without requiring paired training data. Use when the user wants to benchmark on Cityscapes, Google Maps aerial photos & maps, or asks about evaluating this task. Reports FCN score.
This evaluation probes the impact of LLM assistance on human cybersecurity practitioners' ability to execute novel cyberattack challenges. It measures objective performance metrics (phase completion rates and time) and subjective perception (sentiment/mental effort) across inexperienced and highly skilled cohorts, comparing LLM-assisted versus unassisted conditions. Use when the user wants to benchmark on CYBERSECEVAL 3 Challenge Set, or asks about evaluating this task. Reports phase completi...
Evaluates large language models' knowledge of cybersecurity certifications across a spectrum from general IT security to specialized operational technology (OT) and vendor-specific procedural knowledge. It probes whether models can meet professional certification passing standards and identifies gaps in formal, safety-critical industrial protocols. Use when the user wants to benchmark on CyberCertBench, or asks about evaluating this task. Reports accuracy.
Evaluates language models' cybersecurity capabilities by testing their ability to solve real-world Capture the Flag (CTF) challenges in an agent-based environment. It probes iterative problem-solving, command execution in a Linux container, and vulnerability exploitation under constrained iteration and token limits. Use when the user wants to benchmark on Cybench, or asks about evaluating this task. Reports Unguided Performance.
Evaluates robust multi-label classification under long-tailed class distributions and open-world zero-shot generalization to unseen rare diseases in chest X-rays. Use when the user wants to benchmark on PadChest + NIH, or asks about evaluating this task. Reports mAP.
Evaluates multi-turn diagnostic reasoning agents on chest X-rays, measuring their ability to identify tasks, extract evidence, maintain coverage, avoid hallucinations, and sustain coherent dialogue success. Use when the user wants to benchmark on CXReasonDial, or asks about evaluating this task. Reports Faithfulness (Faith).
Evaluates multi-stage structured diagnostic reasoning in chest X-rays, probing a model’s ability to perform visual grounding, anatomical segmentation, quantitative measurement derivation, and clinical threshold application. It tests whether models can consistently link abstract diagnostic criteria with accurate visual interpretation across direct and guided reasoning paths. Use when the user wants to benchmark on CXReasonBench, or asks about evaluating this task. Reports Completion.
Evaluates the capability of object detection models to localize thoracic abnormalities in chest X-rays under a weakly semi-supervised setting. It specifically probes how well models can leverage sparse point-level annotations alongside a small fraction of fully bounding-box-labeled images to achieve accurate region detection. Use when the user wants to benchmark on RSNA, VinDr-CXR, or asks about evaluating this task. Reports mAP.
This benchmark evaluates the capability of vision-language models to generate accurate and clinically relevant free-text radiology reports from chest X-ray images. It probes both linguistic quality through standard NLG metrics and diagnostic accuracy by extracting and comparing clinical abnormality labels against ground truth reports. Use when the user wants to benchmark on IU X-ray, MIMIC-CXR, CheXpert Plus, or asks about evaluating this task. Reports CIDEr.
This benchmark evaluates a model's ability to perform binary classification on lexical complexity, determining whether a given word is perceived as complex or non-complex by human readers. Use when the user wants to benchmark on SemEval CWI, or asks about evaluating this task. Reports F1 score.
Evaluates the quality of novel view synthesis from sparse input views using 3D radiance fields. It measures how well the model reconstructs unseen images and maintains 3D consistency across different sparsity levels (3, 6, or 9 input views). Use when the user wants to benchmark on DTU dataset, Synthetic dataset, or asks about evaluating this task. Reports PSNR.
Evaluates the ranking performance of models for click-through rate (CTR) and post-click conversion rate (CVR) estimation in recommendation systems. It probes the model's ability to correctly rank items by their predicted probability of conversion, while mitigating sample selection bias and false independence assumptions between clicks and conversions. Use when the user wants to benchmark on Industrial Benchmark, Ali-CCP, or asks about evaluating this task. Reports AUC.
This benchmark evaluates the cultural and linguistic understanding of multimodal vision-language models by testing their ability to answer multiple-choice questions about images in diverse languages and cultural contexts. It probes zero-shot generalization across location-aware and location-agnostic prompts, highlighting performance gaps in low-resource languages. Use when the user wants to benchmark on CVQA, or asks about evaluating this task. Reports accuracy.