Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,247
skills in category
886
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 9,121–9,144 of 21,247 skills

Carebench EvalA

Evaluates multimodal fusion of Electronic Health Records (EHR) and Chest X-Rays (CXR) for clinical decision support, specifically testing robustness to missing modalities, temporal imbalance, and subgroup fairness across phenotyping, mortality, and length-of-stay prediction. Use when the user wants to benchmark on CareBench, or asks about evaluating this task. Reports AUROC, AUPRC, F1, Accuracy, Cohen's Kappa.

researchpythongo
0
3
Care EvalA

Evaluates information extraction systems on clinical literature for fine-grained extraction of experimental findings, including entities, attributes, and complex n-ary relations with discontinuous spans and variable arity. Use when the user wants to benchmark on CARE, or asks about evaluating this task. Reports relaxed overlap F1.

researchpythongo
0
3
Cardiac Cmr EvalA

Evaluates automated segmentation and diagnostic classification capabilities on cardiac magnetic resonance imaging (CMR) sequences. Probes the model's ability to accurately delineate cardiac structures across multiple anatomical views and classify cardiovascular diseases with clinical-grade metrics. Use when the user wants to benchmark on BAAI Cardiac CMR Cohort, or asks about evaluating this task. Reports DSC, AUC.

researchpythongo
0
3
Cardiac Arrhythmia Detection EvalA

Evaluates binary classification of ECG heartbeat signals into normal versus abnormal categories. Probes the effectiveness of feature extraction and optimization pipelines for medical signal processing. Use when the user wants to benchmark on MIT-BIH Arrhythmia database (subset), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Card Long Term Forecasting EvalA

Evaluates multivariate time series forecasting models on capturing temporal and cross-channel dependencies across multiple real-world benchmarks. It probes the model's ability to predict future values over varying horizons using fixed historical lookback windows. Use when the user wants to benchmark on ETTm1, ETTm2, ETTh1, ETTh2, Weather, Electricity, Traffic, or asks about evaluating this task. Reports MSE.

researchpythonperformance
0
3
Card Anomaly Detection EvalA

Evaluates the robustness and energy efficiency of a neuromorphic spiking neural network for real-time anomaly detection on lunar rover sensor telemetry. It specifically probes the model's ability to maintain classification accuracy under gradient-based and temporal adversarial attacks while measuring hardware-level power consumption. Use when the user wants to benchmark on Cislunar Anomaly and Risk Dataset (CARD), or asks about evaluating this task. Reports Adversarial Success Rate (ASR).

researchpythongo
0
3
Car EvalA

Evaluates continual semi-supervised learning on activity recognition by measuring how well a model adapts to time-varying unlabeled data streams across sequential sessions without predefined class boundaries. Use when the user wants to benchmark on Continual Activity Recognition (CAR), or asks about evaluating this task. Reports F1-score (class average).

researchpythongo
0
3
Car Bench EvalA

Evaluates LLM agents' ability to resolve uncertainty and adhere to safety policies in automotive environments. It probes limit-awareness, consistency across multiple attempts, and robust multi-turn tool use under incomplete or ambiguous user requests. Use when the user wants to benchmark on CAR-bench, or asks about evaluating this task. Reports Passˆ3.

researchpythongo
0
3
Caqa Attribution EvalA

Evaluates the quality and validity of citations in generated answers for complex question answering. It probes whether models can correctly classify attributions as supportive, insufficient, contradictory, or irrelevant, and assesses their ability to handle varying reasoning complexities. Use when the user wants to benchmark on CAQA, ACLE-Manual, or asks about evaluating this task. Reports micro-F1.

researchpythongo
0
3
Capx Adversarial EvalA

Evaluates the robustness of classifiers against universal and individual adversarial perturbations under domain-specific linear constraints. Probes the trade-off between attack success rate and computational efficiency across finance, network security, medical IoT, and cyber-physical systems. Use when the user wants to benchmark on LCLD, IDS, IoMT, SWaT, WADI, or asks about evaluating this task. Reports ASR.

researchpythonsecurity
0
3
Captrack EvalA

This evaluation framework probes systematic capability drift and forgetting in large language models after post-training. It measures degradation across latent competence (knowledge, reasoning), default behavioral preferences (refusal, verbosity, formatting), and protocol compliance (instruction following, tool use, citation) in legal and medical domains. Use when the user wants to benchmark on CapTrack Evaluation Suite, or asks about evaluating this task. Reports average forgetting.

researchpythongo
0
3
Captionqa EvalA

This benchmark evaluates whether model-generated image captions retain sufficient visual information to answer domain-specific multiple-choice questions without access to the original image. It measures caption utility for downstream reasoning by testing if a text-only QA model can reliably select correct answers or explicitly acknowledge missing information when prompted only with the caption. Use when the user wants to benchmark on CaptionQA, or asks about evaluating this task. Reports Capt...

researchpythongo
0
3
Capsul EvalA

Evaluates the ability of protein sequence and structure models to predict the subcellular localization compartments of human proteins. It probes multi-label classification performance under severe class imbalance, testing whether models can leverage 3D structural motifs or sequence embeddings to identify fine-grained organelle targeting patterns. Use when the user wants to benchmark on CAPSUL, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Capspeech EvalA

This benchmark evaluates text-to-speech models on generating high-fidelity, intelligible speech conditioned on free-form natural language style captions. It probes the model's ability to control intrinsic speaker traits, expressive styles, accents, emotions, and integrate non-verbal sound events across diverse real-world scenarios. Use when the user wants to benchmark on CapTTS, EmoCapTTS, AccCapTTS, CapTTS-SE, AgentTTS, or asks about evaluating this task. Reports binary_correctness.

researchpythongo
0
3
Caprl EvalA

Evaluates the quality of dense image captions by measuring how well they enable downstream multimodal models to answer visual questions accurately. It probes fine-grained visual perception, structured description capability, and the utility of captions for non-visual reasoning. Use when the user wants to benchmark on InfoVQA, DocVQA, ChartQA, Real World QA, Math Vista, SEED2 Plus, MME, MMB, MMStar, MMVet, AI2D, GQA, MMMU, WeMath, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cape Sr EvalA

This evaluation protocol assesses the effectiveness of context-aware position encoding methods in sequential recommendation systems. It measures how well models rank a target item given a user's historical interaction sequence, testing the model's ability to capture temporal and semantic dependencies in user behavior. Use when the user wants to benchmark on AmazonElectronics, KuaiVideo, AmazonBooks, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Capbencher EvalA

Evaluates LLMs on standard benchmarks modified with randomized answers to measure performance tracking and detect data contamination via a Bayes accuracy ceiling. The protocol compares model accuracy against a predefined theoretical maximum (Bayes accuracy) to identify overfitting or memorization. It also assesses robustness to reverse-engineering attacks and cross-lingual contamination. Use when the user wants to benchmark on GSM8K, ARC-Challenge, GPQA, MathQA, MMLU, HLE-MC, MMLU-ProX, BoolQ...

researchpythongo
0
3
Canvas EvalA

Evaluates robotic navigation policies under diverse simulated environments, testing robustness to both precise and misleading natural language instructions. Use when the user wants to benchmark on CANVAS, or asks about evaluating this task. Reports success rate (%).

researchpythontesting
0
3
Canmt EvalA

Evaluates large language models and specialized MT systems on culture-aware machine translation across 12 language pairs. It probes the models' ability to preserve cultural nuances and adapt to explicit semantic versus communicative translation constraints. Use when the user wants to benchmark on CanMT, or asks about evaluating this task. Reports translation performance.

researchpythongit
0
3
Canmt Contrastive EvalA

Evaluates context-aware neural machine translation models on their ability to correctly translate context-dependent discourse phenomena (e.g., anaphoric pronouns, deixis, ellipsis) across sentence boundaries in concatenated input windows. Use when the user wants to benchmark on En→Ru movie subtitles (Voita et al., 2019), En→De TED talk subtitles (IWSLT17), Voita contrastive set, ContraPro, or asks about evaluating this task. Reports Contrastive accuracy.

researchpythongo
0
3
Canadafiresat EvalA

Evaluates deep learning models for high-resolution (100 m) wildfire forecasting using multi-modal satellite and environmental data. It probes the model's ability to predict fire occurrence at the patch level across different temporal splits and land cover types, particularly under severe class imbalance and varying fire danger conditions. Use when the user wants to benchmark on CanadaFireSat, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Cameramotion Vqa EvalA

Evaluates fine-grained camera motion recognition in VideoLLMs using multiple-choice questions and multi-label classification. Probes whether models can distinguish geometric camera movements from object motion and static shots. Use when the user wants to benchmark on CameraMotionVQA, CameraMotionDataset, or asks about evaluating this task. Reports answer accuracy.

researchpythongo
0
3
Camera Trap Ai EvalA

Evaluates AI-powered platforms for processing camera trap images, measuring their ability to detect animals and classify species against ground truth labels. Use when the user wants to benchmark on Colombian rainforest camera trap images, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Camera Pose Nvs EvalA

Evaluates the efficiency-effectiveness trade-off of Structure-from-Motion (SfM) strategies for novel view synthesis. It probes how different feature extractors, matchers, and mappers impact rendering quality and computational runtime across diverse indoor and outdoor scenes. Use when the user wants to benchmark on Mip-NeRF 360, Tanks and Temples, Zip-NeRF, or asks about evaluating this task. Reports PSNR.

researchpythonperformance
0
3