Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,231
skills in category
885
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,785–8,808 of 21,231 skills

Commonforms EvalA

Evaluates an object detection model's ability to locate and classify form field widgets (text inputs, checkboxes/radio buttons, and signatures) on scanned or digital form pages. It probes sensitivity to input resolution and robustness across different languages and document domains. Use when the user wants to benchmark on CommonForms, or asks about evaluating this task. Reports mAP50-95.

researchpythongit
0
3
Commoncanvas EvalA

Evaluates the image quality and text-image alignment of a text-to-image diffusion model trained on Creative-Commons licensed data, benchmarking it against Stable Diffusion 2 using both automated distribution metrics and human pairwise preference. Use when the user wants to benchmark on MS COCO, PartiPrompts, or asks about evaluating this task. Reports User preference rate.

researchpythongo
0
3
Common Voice Asr EvalA

Evaluates multilingual automatic speech recognition (ASR) capabilities, specifically testing speaker generalization and low-resource language adaptation via transfer learning from an English model. It measures how well a model can transcribe audio from diverse, crowdsourced speakers across multiple languages with varying data sizes. Use when the user wants to benchmark on Common Voice, or asks about evaluating this task. Reports character error rate.

researchpythongo
0
3
Commitment Audit EvalA

Evaluates the extent to which authors fulfill promises made during peer review rebuttals in their final camera-ready papers, and classifies unfulfilled commitments by severity and difficulty. Use when the user wants to benchmark on ICLR 2025, EMNLP 2024, or asks about evaluating this task. Reports fulfillment rate.

researchpythongo
0
3
Commit Message Completion EvalA

Evaluates how well models generate or complete commit messages given code diffs and optional historical context. It probes the model's ability to follow coding conventions, match ground truth exactly, and maintain semantic similarity under varying context lengths. Use when the user wants to benchmark on CMG_test, or asks about evaluating this task. Reports ExactMatch@1.

researchpythongo
0
3
Comet Thermal Sim EvalA

Evaluates the accuracy and overhead of an integrated thermal simulation toolchain (CoMeT) for modeling processor-memory thermal dynamics across 2D, 2.5D, and 3D architectures. It probes the tool's ability to capture thermal coupling, leakage power effects, and DVFS/DTM interactions under diverse compute and memory-intensive workloads. Use when the user wants to benchmark on PARSEC 2.1, SPLASH-2, SPEC CPU2017, or asks about evaluating this task. Reports Temperature.

researchpythonperformance
0
3
Comet Mt EvalA

Probes the semantic fidelity and grammatical correctness of automated translation pipelines when converting English benchmarks into low-resource languages. It measures how well translation methods preserve task structure and downstream model performance consistency. Use when the user wants to benchmark on FLORES, WMT24++, MMLU, or asks about evaluating this task. Reports COMET.

researchpythongo
0
3
Combigraph Vis EvalA

Evaluates multimodal discrete mathematical reasoning, specifically the ability to parse and solve combinatorial problems involving graphs, grids, and geometric diagrams. It also probes susceptibility to deliberately crafted distractors in multiple-choice formats versus genuine solution construction. Use when the user wants to benchmark on CombiGraph-Vis, or asks about evaluating this task. Reports avg@8.

researchpythongo
0
3
Combibench EvalA

This benchmark evaluates large language models on formal combinatorial mathematics reasoning within the Lean 4 proof assistant. It probes the model's ability to generate correct, compilable proof scripts and accurately solve fill-in-the-blank combinatorial problems under rigorous automated verification. Use when the user wants to benchmark on CombiBench, or asks about evaluating this task. Reports pass@N.

researchpythongo
0
3
Comat EvalA

Evaluates large language models on mathematical reasoning across diverse difficulty levels and languages. It probes the model's ability to convert natural language word problems into structured symbolic representations and execute step-by-step logical derivations without external solvers. Use when the user wants to benchmark on AQUA, MultiArith, GSM8K, MMLU-Redux, Olympiad Bench (English), GaoKao, Olympiad Bench (Chinese), or asks about evaluating this task. Reports exact match.

researchpythongo
0
3
Colour Mnist Bias EvalA

Evaluates how well a classifier maintains performance on a biased dataset when trained on different coreset selection strategies. It probes the model's robustness to dataset bias and measures the effectiveness of data frugality techniques in mitigating bias while preserving accuracy across varying data budgets. Use when the user wants to benchmark on Colour-MNIST, or asks about evaluating this task. Reports classifier performance.

researchpythongo
0
3
Colosseum EvalA

Evaluates reinforcement learning agents on tabular Markov Decision Processes (MDPs) to measure their performance under varying theoretical hardness criteria, specifically state-action coverage (diameter) and reward structure (environmental value norm). Use when the user wants to benchmark on Colosseum, or asks about evaluating this task. Reports per-step normalized cumulative regret.

researchpythongo
0
3
Colo Dataset EvalA

Evaluates object detection and localization capabilities for indoor cows under varying camera viewpoints (top, side, external) and lighting conditions (day, night). It probes model generalization across domain shifts in perspective and illumination, testing whether pre-trained weights and model complexity transfer effectively to agricultural environments. Use when the user wants to benchmark on COLO, or asks about evaluating this task. Reports mAP@0.5:0.95.

researchpythontesting
0
3
Collective Constitutional Ai EvalA

This protocol evaluates how fine-tuning language models on publicly derived constitutional principles impacts their core reasoning capabilities, social bias propensity, political representativeness, and perceived helpfulness versus harmlessness. It probes whether aligning models with democratic deliberation outputs reduces bias without degrading performance or increasing refusal rates. Use when the user wants to benchmark on MMLU, GSM8K, BBQ, OpinionQA, or asks about evaluating this task. Rep...

researchpythongo
0
3
Coliee Task4 Legal Qa EvalA

Evaluates large language models' ability to perform legal textual entailment and question answering in monolingual and cross-lingual settings. It probes how well models handle linguistic and structural disparities between English and Japanese legal contexts and questions. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Coliee Task4 EvalA

Evaluates the ability of large language models to perform legal textual entailment, specifically measuring how model accuracy changes over time based on the year of the Japanese statute law data used. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cole EvalA

Evaluates French language understanding across 23 diverse tasks, including sentiment analysis, paraphrase detection, grammatical judgment, reasoning, and extractive QA. It specifically probes capabilities like morphological richness, grammatical gender, syntactic nuance, and regional language variation in a zero-shot setting. Use when the user wants to benchmark on COLE, or asks about evaluating this task. Reports task-specific metrics.

researchpythongo
0
3
Cold Start Al 3d Medical Seg EvalA

Evaluates cold-start active learning sample selection strategies for 3D medical image segmentation by comparing how well diversity-based, uncertainty-based, and random methods perform when annotation budgets are extremely limited. Use when the user wants to benchmark on Medical Segmentation Decathlon (MSD), or asks about evaluating this task. Reports Dice score.

researchpythongit
0
3
Cold Offensive Rate EvalA

This benchmark probes the safety and bias of Chinese generative language models by measuring how frequently they produce offensive content when prompted with various inputs, including offensive, non-offensive, and anti-bias contexts. Use when the user wants to benchmark on COLDataset, or asks about evaluating this task. Reports offensive rate.

researchpythongo
0
3
Cold Dti EvalA

Evaluates a model's ability to predict drug-target binding interactions under cold-start conditions where either drugs, proteins, or both are completely unseen during training. It probes the model's capacity to generalize across different protein structural granularities (primary to quaternary) and handle severe class imbalance inherent in biological interaction datasets. Use when the user wants to benchmark on DrugBank, BindingDB, BioSNAP, Human, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Cola EvalA

This benchmark evaluates a model's ability to classify English sentences as grammatically acceptable or unacceptable. It probes syntactic competence by measuring performance on both in-domain and out-of-domain linguistic data. Use when the user wants to benchmark on CoLA, or asks about evaluating this task. Reports MCC.

researchpythongo
0
3
Coir EvalA

Evaluates code information retrieval models across diverse tasks including text-to-code, code-to-code, code-to-text, and hybrid code retrieval. It probes a model's ability to handle semi-structured, syntactically complex code snippets and natural language queries across multiple programming languages and domains. Use when the user wants to benchmark on APPS, CosQA, Synthetic Text2SQL, CodeSearchNet, CodeSearchNet-CCR, CodeTransOcean-DL, CodeTransOcean-Contest, StackOverflow QA, CodeFeedQA, Co...

researchpythonsql
0
3
Coinfra Sync EvalA

Evaluates the temporal synchronization accuracy and robustness of a multi-node cooperative perception system under adverse weather and network conditions. It measures how well a delay-aware protocol aligns sensor data across nodes compared to naive asynchronous methods, focusing on timing errors, fusion completeness, and reaction latency. Use when the user wants to benchmark on CoInfra, or asks about evaluating this task. Reports full_match_rate.

researchpythonreact
0
3
Coin Inbreast EvalA

Evaluates a deep learning model's ability to classify breast masses as benign or malignant in mammography images. It probes the effectiveness of adversarial data augmentation and contrastive manifold learning in improving discriminative feature extraction under data scarcity. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3