Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,900
skills in category
871
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,329–8,352 of 20,900 skills

Cot EvalA

Evaluates the zero-shot and few-shot reasoning capabilities of language models, specifically probing their ability to generate step-by-step chain-of-thought rationales and produce correct answers across classification and generation tasks. Use when the user wants to benchmark on BigBench Hard (BBH), P3, MGSM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cosyvoice Tts EvalA

Evaluates zero-shot text-to-speech synthesis quality, focusing on content consistency (how well generated speech matches input text) and speaker similarity (how well the cloned voice matches the reference speaker) across English and Chinese. It also probes emotion controllability and the utility of synthesized speech for augmenting ASR training data. Use when the user wants to benchmark on LibriTTS, AISHELL-3, or asks about evaluating this task. Reports WER (%), CER (%).

researchpythonshell
0
3
Costnav EvalA

Economic viability and cost-aware performance of embodied agents in urban sidewalk delivery navigation. It evaluates how technical metrics like collision rate and arrival success translate into real-world financial outcomes, including maintenance costs, energy usage, revenue, and break-even points. Use when the user wants to benchmark on CostNav Urban Sidewalk Navigation Simulation, or asks about evaluating this task. Reports Profit/run.

researchpythongit
0
3
Cosql EvalA

Evaluates conversational text-to-SQL systems on cross-domain database querying. It probes dialogue state tracking via SQL grounding, response generation from query results, and user intent/dialogue act prediction under real-world ambiguity and clarification dynamics. Use when the user wants to benchmark on CoSQL, or asks about evaluating this task. Reports Question Match.

researchpythongo
0
3
Cosql Cg EvalA

This benchmark probes a model's ability to perform compositional generalization in context-dependent Text-to-SQL. It evaluates whether models can correctly combine previously seen SQL query structures with novel modification patterns (e.g., new WHERE or ORDER BY clauses) in multi-turn dialogues. Use when the user wants to benchmark on CoSQL-CG, or asks about evaluating this task. Reports question match (QM).

researchpythongo
0
3
Cosmos Qa EvalA

This benchmark evaluates a model's ability to perform contextual commonsense reasoning in machine reading comprehension. It probes whether systems can make non-literal, implicit inferences about causes, effects, and counterfactuals based on personal narratives, rather than relying on explicit textual evidence or simple semantic matching. Use when the user wants to benchmark on Cosmos QA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Cosmos Drive Dreams EvalA

Evaluates the effectiveness of a synthetic driving data generation pipeline by measuring performance gains in downstream autonomous driving perception tasks, including 3D lane detection, 3D object detection, and LiDAR-based detection, particularly under challenging conditions like extreme weather and nighttime. Use when the user wants to benchmark on Waymo Open Dataset, RDS-HQ, RDS-HQ-HL, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Cosmoflow Hpc ScalingA

Evaluates the compute efficiency and horizontal scalability of a 3D convolutional neural network framework on supercomputers, measuring sustained floating-point throughput and parallel scaling efficiency across thousands of nodes. Use when the user has predictions and gold and needs to compute Pflop/s.

researchpythongo
0
3
Cosmic Symmetry Benchmark EvalA

Evaluates the ability of graph neural networks to extract cosmological parameters and local velocity fields from large-scale 3D point clouds of dark matter halos. It probes both global long-range correlation capture (via cosmological parameter regression) and local geometric dependency modeling (via per-node velocity prediction). Use when the user wants to benchmark on Quijote BSQ Point Cloud, or asks about evaluating this task. Reports MSE.

researchpythonnode
0
3
Corrected P ValueA

Evaluates whether turn-level conversational metrics in LLM interactions suffer from temporal autocorrelation that inflates statistical significance. It compares naive pooled hypothesis testing against cluster-robust corrections to measure false positive rates and classify metric robustness. Use when the user has predictions and gold and needs to compute corrected_p_value.

researchpythongo
0
3
Corpus Technical Validation EvalA

Evaluates the structural integrity, metadata completeness, and text quality of a legally screened chemistry corpus derived from S2ORC. It verifies schema compliance, metadata field availability, subfield label validity, chunking consistency, and embedding reproducibility against predefined thresholds. Use when the user wants to benchmark on Lit2Vec Chemistry Corpus, or asks about evaluating this task. Reports schema_pass_rate.

researchpython
0
3
Corporate Fraud Detection EvalA

Evaluates a model's ability to detect corporate fraud using financial graphs. It specifically probes robustness to information overload from noisy support nodes (e.g., directors) and label noise caused by delayed fraud detection. Use when the user wants to benchmark on MBM, SME, GEM, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Corebt EvalA

Evaluates multimodal fusion models for robust brain tumor typing by integrating MRI, histopathology, and diagnostic text under variable modality availability conditions. The benchmark probes a model's ability to perform fine-grained hierarchical classification across six glioma subtypes when some modalities are missing or degraded. Use when the user wants to benchmark on CoRe-BT, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Core Fewshot Rc EvalA

Evaluates few-shot relation classification models on company and business entity relations, testing their ability to resolve entity ambiguity and adapt across domains using limited labeled examples. Use when the user wants to benchmark on CORE, or asks about evaluating this task. Reports Micro F1.

researchpythontesting
0
3
Corda Peft EvalA

Evaluates parameter-efficient fine-tuning methods across mathematical reasoning, code generation, instruction following, and general language understanding tasks, while measuring their ability to retain pre-trained world knowledge. Use when the user wants to benchmark on MetaMathQA, GSM8k, Math, CodeFeedback, HumanEval, MBPP, WizardLM-Evol-Instruct, MTBench, TriviaQA, NQ open, WebQS, GLUE, Wikitext-2, Penn TreeBank (PTB), or asks about evaluating this task. Reports exact match scores.

researchpythongo
0
3
Coralscapes EvalA

Probes semantic segmentation models on complex underwater scenes characterized by high morphological variability, degradation states, and visual distortions. It evaluates the model's ability to generalize across geographically distinct reef sites and handle severe class imbalance and fine-grained benthic classification. Use when the user wants to benchmark on Coralscapes, or asks about evaluating this task. Reports mean Intersection over Union (mIoU).

researchpythongo
0
3
Cora Crl EvalA

Evaluates continual reinforcement learning agents across sequential task environments, probing their ability to learn new tasks while retaining old ones (plasticity vs stability) and generalizing to unseen contexts. Use when the user wants to benchmark on Procgen, MiniHack, CHORES, Atari, or asks about evaluating this task. Reports Continual Evaluation ($\mathcal{C}$).

researchpythongo
0
3
Coqstoq EvalA

Evaluates a language model's ability to synthesize complete formal proofs in Coq by dynamically retrieving relevant project-specific lemmas and proofs. It measures how effectively retrieval-augmented proving and search strategies improve theorem synthesis success rates over time. Use when the user wants to benchmark on CoqStoq, or asks about evaluating this task. Reports Theorems Proven.

researchpythongo
0
3
Coqa EvalA

Evaluates a model's ability to answer free-form questions in a multi-turn conversational setting. It probes coreference resolution, pragmatic reasoning, and the capacity to maintain and leverage dialogue history over a given context passage. Use when the user wants to benchmark on CoQA, or asks about evaluating this task. Reports macro-average F1 score of word overlap.

researchpythongo
0
3
Copyright Tracking Tmr EvalA

Evaluates the robustness of copyright tracking methods in fine-tuned Large Vision-Language Models (LVLMs) by measuring whether adversarial image triggers can consistently elicit a predefined target response after the model has been adapted on various downstream datasets. Use when the user wants to benchmark on ImageNet 2012 (validation subset), V7W, ST-VQA, TextVQA, PaintingForm, MathV360k, ChEBI-20, or asks about evaluating this task. Reports target match rate (TMR).

researchpythonperformance
0
3
Copo Hallucination EvalA

Evaluates the ability of Multimodal Large Language Models to generate factually grounded captions and answers while suppressing object-level hallucinations. It probes visual grounding, reasoning consistency, and alignment with human or GPT-4 preferences across multiple reasoning and perception benchmarks. Use when the user wants to benchmark on CHAIR, POPE, MMBench, MME, or asks about evaluating this task. Reports POPE F1 Score.

researchpythongo
0
3
Copep EvalA

Evaluates the ability of protein language models to adapt to evolving biological databases through continual pretraining. It probes how well models maintain performance on high-quality sequence validation, predict mutation fitness effects, and generalize across diverse protein understanding tasks over time. Use when the user wants to benchmark on UniProt Validation Set, ProteinGym, PEER, DGEB, or asks about evaluating this task. Reports Spearman correlation.

researchpythongo
0
3
Convsearch R1 EvalA

Evaluates conversational query reformulation (CQR) by measuring how effectively a model rewrites multi-turn queries into standalone search queries that retrieve relevant passages. It probes the model's ability to optimize rewrites using only retrieval signals, without human annotations or LLM distillation. Use when the user wants to benchmark on TopiOCQA, QReCC, or asks about evaluating this task. Reports MRR@3.

researchpython
0
3
Convomem EvalA

Evaluates conversational memory capabilities across six dimensions: recalling user facts, tracking assistant statements, abstaining when information is missing, inferring preferences, handling changing facts, and making implicit connections. It specifically tests the ability to synthesize evidence distributed across multiple conversation turns. Use when the user wants to benchmark on ConvoMem, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3