Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,247
skills in category
886
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 9,145–9,168 of 21,247 skills

Camera Control EvalA

Evaluates a generative model's ability to simulate physical camera effects (bokeh, focal length, shutter speed, color temperature) while preserving scene consistency and adhering to text prompts. Use when the user wants to benchmark on Custom Camera Control Dataset, or asks about evaluating this task. Reports CorrCoef.

researchpythonaws
0
3
Cameo EvalA

Evaluates speech emotion recognition (SER) models across multiple languages and emotional states. It probes a model's ability to map raw audio inputs to discrete emotional categories without relying on speaker or language metadata. Use when the user wants to benchmark on CAMEO, or asks about evaluating this task. Reports macro-averaged F1 score.

researchpythongo
0
3
Camelyon17 Wilds EvalA

Evaluates domain generalization capability for metastatic breast cancer detection in histopathology images. The protocol trains models on data from known medical centers and tests them on completely unseen centers with different staining protocols and scanners to measure adaptability. Use when the user wants to benchmark on Camelyon17 WILDS, or asks about evaluating this task. Reports accuracy (%).

researchpythonnode
0
3
Camchex EvalA

Evaluates a multimodal framework's ability to classify thoracic diseases at the study level by jointly modeling multi-view chest X-rays, clinical indications, and vital signs. It probes the model's capacity to integrate heterogeneous clinical data for accurate multi-label diagnosis across head, body, and tail disease categories. Use when the user wants to benchmark on MIMIC-CXR, CXR-LT 2023, CXR-LT 2024, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Cam4docc EvalA

Evaluates camera-only 4D occupancy forecasting for autonomous driving by predicting current and future voxel states in a 3D grid. It specifically probes the model's ability to distinguish between general movable objects (GMO) and general static objects (GSO) across multiple future time steps. Use when the user wants to benchmark on nuScenes, nuScenes-Occupancy, Lyft-Level5, or asks about evaluating this task. Reports ~IoU_f.

researchpythongit
0
3
Calvin Long Horizon EvalA

Evaluates a robot policy's ability to chain multiple sub-goals sequentially using only visual observations and goal images. It probes long-horizon planning, goal-conditioned control, and the capacity to generalize to unseen goal configurations without explicit reward signals. Use when the user wants to benchmark on CALVIN, or asks about evaluating this task. Reports Success rate.

researchpythongo
0
3
Calorimetry Reconstruction EvalA

Evaluates deep learning models for high-energy physics calorimetry tasks, specifically particle shower generation and particle reconstruction (identification and energy regression) using simulated detector data. Use when the user wants to benchmark on LCD Calorimeter Dataset (GEN & REC), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Caligraph Owl Reasoning EvalA

Evaluates the scalability and logical correctness of OWL2 EL reasoners when processing large ontologies with complex owl:hasValue restrictions. It measures whether systems can successfully materialize subclass hierarchies and infer individual and literal assertions, and how long they take under memory and time constraints. Use when the user wants to benchmark on CaLiGraph, or asks about evaluating this task. Reports Inferrable Assertions.

researchpythongo
0
3
Calicalcausalrank EvalA

Evaluates multi-objective ad ranking models on click-through rate (CTR) and conversion rate (CVR) prediction tasks, focusing on ranking quality, score calibration, and counterfactual utility estimation under position and selection bias. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Calibration Error EstimationA

Evaluates the sample complexity and verification cost required to estimate calibration error in AI models under rare-error regimes. It probes whether passive querying or active querying can reliably detect miscalibration and how estimation error scales with sample size and model smoothness. Use when the user has predictions and gold and needs to compute estimation error.

researchpythongo
0
3
Calib3d EvalA

Evaluates the calibration quality of 3D scene understanding models by measuring the discrepancy between predicted confidence and actual accuracy. It probes reliability under both in-domain conditions and out-of-domain stressors like adverse weather, sensor failures, and domain shifts. Use when the user wants to benchmark on nuScenes, SemanticKITTI, Waymo Open, SemanticPOSS, SemanticSTF, ScribbleKITTI, Synth4D, S3DIS, or asks about evaluating this task. Reports ECE.

researchpythongo
0
3
Calendar Scheduling EvalA

Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect conflicts accurately. Use when the user wants to benchmark on Calendar Scheduling, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Cage Korset EvalA

Probes large language models' vulnerability to culturally-adapted adversarial prompts across Korean and Khmer contexts. It measures how effectively models resist harmful intent when framed within local socio-technical norms, safety policies, and cultural specifics, rather than generic or translated attacks. Use when the user wants to benchmark on KorSET, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythongit
0
3
Cafp Fairness EvalA

Evaluates whether a post-processing framework can reduce group-level disparities in predictions while maintaining predictive accuracy. It probes a model's ability to balance fairness constraints (demographic parity and equalized odds) against standard classification performance across multiple benchmark datasets. Use when the user wants to benchmark on Adult Income (UCI), COMPAS Recidivism, German Credit, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
Caf 7m EvalA

Evaluates whether context-aided forecasting models can effectively leverage supplementary descriptive context to improve probabilistic time series forecasts, and tests generalization to out-of-domain real-world datasets. Use when the user wants to benchmark on CAF-7M, CGTSF, GIFT-Eval, or asks about evaluating this task. Reports CRPS.

researchpythongit
0
3
Cads Whole Body Ct EvalA

Evaluates AI models on voxel-level segmentation of 167 whole-body anatomical structures from CT scans. It probes generalization across diverse imaging protocols, patient demographics, and pathological conditions, with a focus on clinical utility in radiation oncology. Use when the user wants to benchmark on CADS-dataset, 18 Public Benchmark Datasets, or asks about evaluating this task. Reports Dice coefficient.

researchpythongo
0
3
Cache Gen EvalA

Evaluates the efficiency and generation quality of a KV cache compression and streaming system for LLM serving. It measures how effectively the system reduces network bandwidth and time-to-first-token across varying context lengths, network conditions, and concurrent requests while maintaining task-specific accuracy, F1, or perplexity. Use when the user wants to benchmark on LongChat, TriviaQA, NarrativeQA, Wikitext, or asks about evaluating this task. Reports TTFT.

researchpythongo
0
3
Cabuar Burned Area Delineation EvalA

This benchmark evaluates a model's ability to accurately delineate wildfire-affected regions from satellite imagery. It probes pixel-level binary segmentation and change detection capabilities using pre- and post-fire Sentinel-2 multispectral data. Use when the user wants to benchmark on CaBuAr, or asks about evaluating this task. Reports pixel-level accuracy.

researchpythongit
0
3
Cab EvalA

Evaluates whether LLMs exhibit biased responses to automatically generated, realistic open-ended questions. It probes for hidden biases across sensitive attributes (e.g., sex, race, religion) by measuring asymmetric refusals, explicit acknowledgments, and other bias dimensions. Use when the user wants to benchmark on CAB, or asks about evaluating this task. Reports fitness score.

researchpythongo
0
3
Caad 3k EvalA

Evaluates a model's ability to detect contextual anomalies where normality depends on subject-context alignment rather than intrinsic appearance. It probes cross-context generalization by testing on unseen subject-context combinations and zero-shot transfer to real-world out-of-context benchmarks. Use when the user wants to benchmark on CAAD-3K, MVTec-AD, VisA, MIT-OOC, COCO-OOC, or asks about evaluating this task. Reports I-AUROC.

researchpythontesting
0
3
Caa EvalA

Evaluates the robustness of Large Audio-Language Models (LALMs) against adversarial audio attacks in conversational settings. It probes response consistency, semantic preservation, and linguistic quality when audio inputs are perturbed with content, emotional, explicit, or implicit noise. Use when the user wants to benchmark on CAA, or asks about evaluating this task. Reports WER.

researchpythongit
0
3
Ca Afp EvalA

Evaluates federated learning models for human activity recognition under non-IID data distributions. It measures classification accuracy, fairness across heterogeneous clients, and communication efficiency during model pruning and clustering. Use when the user wants to benchmark on WISDM, UCI-HAR, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
C4 Loss Scaling EvalA

Evaluates how optimal batch size and learning rate scale with model size and target loss during pre-training. It probes the stability of hyperparameters across different model scales and quantifies the relationship between batch size and validation loss on a standard text corpus. Use when the user wants to benchmark on C4, or asks about evaluating this task. Reports C4 Loss.

researchpythongit
0
3
C3 Bench EvalA

Evaluates the robustness and multi-tasking capabilities of LLM-based agents by probing their ability to handle complex tool dependencies, propagate hidden information across tasks, and maintain stable decision policies under dynamic, multi-round interactions. Use when the user wants to benchmark on C^3-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3