Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,231
skills in category
885
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,737–8,760 of 21,231 skills

Contamination RateA

Measures the extent to which multimodal evaluation benchmarks are contaminated by pre-training data, assessing both visual similarity and textual inference leakage to quantify data contamination risks. Use when the user has predictions and gold and needs to compute image-only contamination rate.

researchpythongo
0
3
Contact Rich Manipulation EvalA

Evaluates the success rate of imitation learning policies in contact-rich manipulation tasks requiring precise force control, slip detection, and in-hand pose estimation. It compares vision-only baselines against visuo-tactile policies with and without temporal-aware contrastive pretraining. Use when the user wants to benchmark on Contact-Rich Manipulation Tasks, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Construction Site 10k EvalA

Evaluates vision-language models on construction site safety inspection tasks, including image captioning, safety rule violation detection, reasoning, and visual grounding of specific objects. Use when the user wants to benchmark on ConstructionSite 10k, or asks about evaluating this task. Reports IoU.

researchpythongo
0
3
Construct EvalA

Evaluates the trustworthiness and accuracy of LLM-generated structured outputs (JSON) against a ground truth or expected schema. It probes the model's ability to detect per-field and per-document errors in data extraction tasks without requiring labeled data. Use when the user wants to benchmark on Four real-world datasets (unspecified in excerpt), or asks about evaluating this task. Reports rating.

researchpythonrust
0
3
Constrained Adaptive Attack EvalA

Evaluates the adversarial robustness of tabular deep learning models under realistic, domain-aware constraints. It measures how easily an attacker can flip model predictions while respecting feature mutability, boundaries, types, and relational constraints across ten progressively restricted threat models. Use when the user wants to benchmark on phishing, credit scoring, botnet detection, or asks about evaluating this task. Reports adversarial_label_flip.

researchpythongo
0
3
Constitution Ai Feedback EvalA

This evaluation probes how different instructional guidelines (constitutions) shape AI-generated medical dialogues across specific socio-communicative dimensions like empathy, information gathering, and decision-making. It measures human preference for dialogue quality under varying constitutional constraints. Use when the user wants to benchmark on Custom AI-generated medical dialogues, or asks about evaluating this task. Reports Bradley-Terry preference rate.

researchpythongo
0
3
Consistencychecker EvalA

Evaluates LLM generalization and functional consistency by measuring how well models preserve core functionality after iterative, reversible transformations. It probes cumulative error and path-specific divergence across multi-step transformation sequences without relying on static benchmarks. Use when the user wants to benchmark on ConsistencyChecker (Dynamic), or asks about evaluating this task. Reports forest-level consistency score (C3(F)).

researchpythonnode
0
3
Consensus Layer Pruning EvalA

Evaluates a multi-metric layer pruning method (Consensus) on image classification models, measuring trade-offs between computational efficiency (FLOPs reduction) and predictive performance (accuracy drop), while also assessing robustness against adversarial and out-of-distribution attacks. Use when the user wants to benchmark on CIFAR-10, ImageNet, CIFAR-10.2, CIFAR-C, ImageNet-C, or asks about evaluating this task. Reports Δ Acc. (difference in accuracy).

researchpythonperformance
0
3
Connect The Dots EvalA

Evaluates a model's ability to locate and connect dots in sequential order across various visual patterns. It probes precise spatial reasoning and the capacity to generate non-destructive SVG overlays that explain the reasoning process. Use when the user wants to benchmark on Connect-the-Dots, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Conll2012 Coref EvalA

Evaluates a model's ability to jointly detect mentions and resolve coreferential chains in English text without relying on external syntactic parsers or hand-crafted features. It measures how well the model groups word spans into entity clusters based on contextual and structural cues. Use when the user wants to benchmark on CoNLL-2012 (English), or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Conll Ner EvalA

Evaluates a model's ability to perform Named Entity Recognition (NER) across multiple languages, specifically testing its robustness to out-of-domain text, orthographic variations, and cross-lingual transfer when trained on noisy Wikipedia-derived data. Use when the user wants to benchmark on CoNLL 2002/2003 NER, or asks about evaluating this task. Reports Exact F1.

researchpythongo
0
3
Confready Checklist EvalA

This benchmark evaluates a model's ability to accurately answer conference submission checklist questions based on manuscript content. It specifically probes long-form document understanding, retrieval-augmented generation (RAG) effectiveness, and the model's capacity to reflect on ethical considerations, reproducibility, and societal impacts. Use when the user wants to benchmark on ConfReady Evaluation Set, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Conformal Prediction EvalA

Evaluates the ability of conformal prediction frameworks to produce statistically valid prediction sets with instance-level uncertainty quantification for encoder-only transformers, measuring both classification accuracy and calibration efficiency across standard NLP benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, or asks about evaluating this task. Reports Test Accuracy.

researchpythongo
0
3
Conformal Lesion Segmentation EvalA

This evaluation protocol assesses the ability of 3D medical image segmentation models to control false negative rates under user-specified risk constraints while maintaining spatial precision. It benchmarks a model-agnostic conformal prediction calibration method against fixed heuristic thresholds across multiple anatomical datasets. Use when the user wants to benchmark on KiTS21, LiTS, NIH-LN ABD, LIDC-IDRI, MDSC-Colon, MDSC-Pancreas, or asks about evaluating this task. Reports ECR.

researchpythongo
0
3
Conformal Anomaly Detection EvalA

Evaluates the statistical validity (False Discovery Rate control) and detection sensitivity (statistical power) of cross-conformal anomaly detection methods against split-conformal baselines across datasets of varying sizes and dimensionalities. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports False Discovery Rate (FDR).

researchpythongit
0
3
Confer EvalA

Evaluates continual learning methods for facial expression recognition under incremental, non-i.i.d. data settings. It probes a model's ability to learn new expressions sequentially while preserving prior knowledge, measuring both forward adaptation and backward forgetting. Use when the user wants to benchmark on CK+ (Extended Cohn-Kanade), or asks about evaluating this task. Reports Average Accuracy Score.

researchpythongo
0
3
Condmedqa EvalA

Evaluates a model's ability to perform conditional multi-hop reasoning in biomedical question answering, specifically how well it modulates clinical answers based on patient-specific constraints like comorbidities, contraindications, and special population factors. Use when the user wants to benchmark on CondMedQA, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Conditional Unigram Tokenization EvalA

Evaluates a conditional unigram tokenizer's cross-lingual alignment quality and its impact on downstream machine translation and language modeling tasks. It measures intrinsic tokenization properties, alignment accuracy, and task-specific performance metrics. Use when the user wants to benchmark on NLLB, MultiParaCrawl, WMT2020, Flores, WMT2020 test set, or asks about evaluating this task. Reports chrF++.

researchpythonrust
0
3
Conda EvalA

Evaluates in-game toxicity detection using a dual-level NLU framework that jointly predicts utterance-level toxicity intent and token-level semantic slots. It probes a model's ability to understand contextual, game-specific language and distinguish between explicit, implicit, and action-based toxicity. Use when the user wants to benchmark on CONDA, or asks about evaluating this task. Reports UCA.

researchpythongo
0
3
Cond P Diff EvalA

Evaluates a conditional latent diffusion framework's ability to synthesize task-specific LoRA parameters for NLP and image style-transfer tasks. It probes whether generated parameters can match or exceed standard fine-tuning and model-averaging baselines across diverse domains. Use when the user wants to benchmark on GLUE benchmark, SemArt, WikiArt, or asks about evaluating this task. Reports Average accuracy.

researchpythonperformance
0
3
Concurrence EvalA

Measures how consistently a modeling approach's performance ranking holds across different question answering benchmarks. It probes whether improvements in QA models generalize across datasets with varying data collection procedures, passage/question distributions, and targeted linguistic phenomena. Use when the user wants to benchmark on SQuAD, NewsQA, NaturalQuestions, DROP, HotpotQA, QAMR, or asks about evaluating this task. Reports concurrence (Spearman's τ).

researchpythonperformance
0
3
Concode EvalA

Probes a model's ability to generate syntactically valid Java member functions from natural language documentation, conditioned on a full class environment including variable types, method signatures, and their interdependencies. It evaluates context-aware code generation, identifier disambiguation, and code reusability. Use when the user wants to benchmark on CONCODE, or asks about evaluating this task. Reports Exact match accuracy.

researchpythonjava
0
3
Conceptmix EvalA

Evaluates the compositional generalization capability of text-to-image models by testing their ability to generate images that satisfy multiple, simultaneously specified visual concepts (objects, colors, shapes, spatial relationships, etc.) within a single prompt. The benchmark probes model robustness to increasing compositional complexity (k) and reveals limitations in handling less frequent concept combinations. Use when the user wants to benchmark on ConceptMix, or asks about evaluating th...

researchpythongo
0
3
Concap Nids EvalA

Evaluates the ability of machine learning and flow-based intrusion detection systems to accurately classify network traffic flows as benign or malicious. It probes the model's capacity to generalize across real-world benchmarks and synthetically generated, automatically labeled traffic for multi-step attack scenarios. Use when the user wants to benchmark on CICIDS17, ConCap ssh-patator, or asks about evaluating this task. Reports tpr.

researchpythongo
0
3