Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,145–9,168 of 21,247 skills
Evaluates a generative model's ability to simulate physical camera effects (bokeh, focal length, shutter speed, color temperature) while preserving scene consistency and adhering to text prompts. Use when the user wants to benchmark on Custom Camera Control Dataset, or asks about evaluating this task. Reports CorrCoef.
Evaluates speech emotion recognition (SER) models across multiple languages and emotional states. It probes a model's ability to map raw audio inputs to discrete emotional categories without relying on speaker or language metadata. Use when the user wants to benchmark on CAMEO, or asks about evaluating this task. Reports macro-averaged F1 score.
Evaluates domain generalization capability for metastatic breast cancer detection in histopathology images. The protocol trains models on data from known medical centers and tests them on completely unseen centers with different staining protocols and scanners to measure adaptability. Use when the user wants to benchmark on Camelyon17 WILDS, or asks about evaluating this task. Reports accuracy (%).
Evaluates a multimodal framework's ability to classify thoracic diseases at the study level by jointly modeling multi-view chest X-rays, clinical indications, and vital signs. It probes the model's capacity to integrate heterogeneous clinical data for accurate multi-label diagnosis across head, body, and tail disease categories. Use when the user wants to benchmark on MIMIC-CXR, CXR-LT 2023, CXR-LT 2024, or asks about evaluating this task. Reports mAP.
Evaluates camera-only 4D occupancy forecasting for autonomous driving by predicting current and future voxel states in a 3D grid. It specifically probes the model's ability to distinguish between general movable objects (GMO) and general static objects (GSO) across multiple future time steps. Use when the user wants to benchmark on nuScenes, nuScenes-Occupancy, Lyft-Level5, or asks about evaluating this task. Reports ~IoU_f.
Evaluates a robot policy's ability to chain multiple sub-goals sequentially using only visual observations and goal images. It probes long-horizon planning, goal-conditioned control, and the capacity to generalize to unseen goal configurations without explicit reward signals. Use when the user wants to benchmark on CALVIN, or asks about evaluating this task. Reports Success rate.
Evaluates deep learning models for high-energy physics calorimetry tasks, specifically particle shower generation and particle reconstruction (identification and energy regression) using simulated detector data. Use when the user wants to benchmark on LCD Calorimeter Dataset (GEN & REC), or asks about evaluating this task. Reports accuracy.
Evaluates the scalability and logical correctness of OWL2 EL reasoners when processing large ontologies with complex owl:hasValue restrictions. It measures whether systems can successfully materialize subclass hierarchies and infer individual and literal assertions, and how long they take under memory and time constraints. Use when the user wants to benchmark on CaLiGraph, or asks about evaluating this task. Reports Inferrable Assertions.
Evaluates multi-objective ad ranking models on click-through rate (CTR) and conversion rate (CVR) prediction tasks, focusing on ranking quality, score calibration, and counterfactual utility estimation under position and selection bias. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
Evaluates the sample complexity and verification cost required to estimate calibration error in AI models under rare-error regimes. It probes whether passive querying or active querying can reliably detect miscalibration and how estimation error scales with sample size and model smoothness. Use when the user has predictions and gold and needs to compute estimation error.
Evaluates the calibration quality of 3D scene understanding models by measuring the discrepancy between predicted confidence and actual accuracy. It probes reliability under both in-domain conditions and out-of-domain stressors like adverse weather, sensor failures, and domain shifts. Use when the user wants to benchmark on nuScenes, SemanticKITTI, Waymo Open, SemanticPOSS, SemanticSTF, ScribbleKITTI, Synth4D, S3DIS, or asks about evaluating this task. Reports ECE.
Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect conflicts accurately. Use when the user wants to benchmark on Calendar Scheduling, or asks about evaluating this task. Reports Exact Match (EM).
Probes large language models' vulnerability to culturally-adapted adversarial prompts across Korean and Khmer contexts. It measures how effectively models resist harmful intent when framed within local socio-technical norms, safety policies, and cultural specifics, rather than generic or translated attacks. Use when the user wants to benchmark on KorSET, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates whether a post-processing framework can reduce group-level disparities in predictions while maintaining predictive accuracy. It probes a model's ability to balance fairness constraints (demographic parity and equalized odds) against standard classification performance across multiple benchmark datasets. Use when the user wants to benchmark on Adult Income (UCI), COMPAS Recidivism, German Credit, or asks about evaluating this task. Reports Accuracy.
Evaluates whether context-aided forecasting models can effectively leverage supplementary descriptive context to improve probabilistic time series forecasts, and tests generalization to out-of-domain real-world datasets. Use when the user wants to benchmark on CAF-7M, CGTSF, GIFT-Eval, or asks about evaluating this task. Reports CRPS.
Evaluates AI models on voxel-level segmentation of 167 whole-body anatomical structures from CT scans. It probes generalization across diverse imaging protocols, patient demographics, and pathological conditions, with a focus on clinical utility in radiation oncology. Use when the user wants to benchmark on CADS-dataset, 18 Public Benchmark Datasets, or asks about evaluating this task. Reports Dice coefficient.
Evaluates the efficiency and generation quality of a KV cache compression and streaming system for LLM serving. It measures how effectively the system reduces network bandwidth and time-to-first-token across varying context lengths, network conditions, and concurrent requests while maintaining task-specific accuracy, F1, or perplexity. Use when the user wants to benchmark on LongChat, TriviaQA, NarrativeQA, Wikitext, or asks about evaluating this task. Reports TTFT.
This benchmark evaluates a model's ability to accurately delineate wildfire-affected regions from satellite imagery. It probes pixel-level binary segmentation and change detection capabilities using pre- and post-fire Sentinel-2 multispectral data. Use when the user wants to benchmark on CaBuAr, or asks about evaluating this task. Reports pixel-level accuracy.
Evaluates whether LLMs exhibit biased responses to automatically generated, realistic open-ended questions. It probes for hidden biases across sensitive attributes (e.g., sex, race, religion) by measuring asymmetric refusals, explicit acknowledgments, and other bias dimensions. Use when the user wants to benchmark on CAB, or asks about evaluating this task. Reports fitness score.
Evaluates a model's ability to detect contextual anomalies where normality depends on subject-context alignment rather than intrinsic appearance. It probes cross-context generalization by testing on unseen subject-context combinations and zero-shot transfer to real-world out-of-context benchmarks. Use when the user wants to benchmark on CAAD-3K, MVTec-AD, VisA, MIT-OOC, COCO-OOC, or asks about evaluating this task. Reports I-AUROC.
Evaluates the robustness of Large Audio-Language Models (LALMs) against adversarial audio attacks in conversational settings. It probes response consistency, semantic preservation, and linguistic quality when audio inputs are perturbed with content, emotional, explicit, or implicit noise. Use when the user wants to benchmark on CAA, or asks about evaluating this task. Reports WER.
Evaluates federated learning models for human activity recognition under non-IID data distributions. It measures classification accuracy, fairness across heterogeneous clients, and communication efficiency during model pruning and clustering. Use when the user wants to benchmark on WISDM, UCI-HAR, or asks about evaluating this task. Reports Accuracy.
Evaluates how optimal batch size and learning rate scale with model size and target loss during pre-training. It probes the stability of hyperparameters across different model scales and quantifies the relationship between batch size and validation loss on a standard text corpus. Use when the user wants to benchmark on C4, or asks about evaluating this task. Reports C4 Loss.
Evaluates the robustness and multi-tasking capabilities of LLM-based agents by probing their ability to handle complex tool dependencies, propagate hidden information across tasks, and maintain stable decision policies under dynamic, multi-round interactions. Use when the user wants to benchmark on C^3-Bench, or asks about evaluating this task. Reports accuracy.