Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,869
skills in category
870
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,209–8,232 of 20,869 skills

Csaw M EvalA

Evaluates models on ordinal classification of mammographic masking potential (levels 1–8) and their clinical utility in predicting interval and large invasive cancers. It probes the model's ability to respect ordinal relationships in breast tissue obscuration and correlate these estimates with cancer outcomes. Use when the user wants to benchmark on CSAW-M, or asks about evaluating this task. Reports average mean absolute error (AMAE).

researchpythongo
0
3
Cs Kws EvalA

Evaluates a model's ability to detect and localize multiple spoken keywords within continuous, untrimmed audio streams, distinguishing target keywords from background speech and silence. Use when the user wants to benchmark on LibriTop-20, CMAK-7, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Cs Dialogue Asr EvalA

This benchmark evaluates automatic speech recognition (ASR) systems on their ability to accurately transcribe spontaneous, full-length dialogues that alternate between Mandarin and English. It probes a model's robustness to language alternation, phonetic mismatches, and contextual dependencies in naturalistic code-switching scenarios. Use when the user wants to benchmark on CS-Dialogue, or asks about evaluating this task. Reports MER.

researchpythonperformance
0
3
Cs 4k EvalA

Evaluates LLMs on end-to-end computer science research workflows by testing their ability to answer scientific questions grounded in academic papers. It probes domain-specific reasoning, factual recall, and methodological understanding across eight research workflow categories. Use when the user wants to benchmark on CS-4k, or asks about evaluating this task. Reports model response score.

researchpythongo
0
3
Crystal Structure Discovery EvalA

Evaluates an LLM's capacity to discover stable crystal structures by iteratively mutating and crossing over parent structures to minimize deformation energy. Use when the user wants to benchmark on MatBenchbandgap, or asks about evaluating this task. Reports deformation energy.

researchpythongo
0
3
Cruxeval EvalA

Evaluates a model's ability to reason about and execute short Python functions by predicting outputs given inputs (CRUXEval-I) and predicting inputs given outputs (CRUXEval-O). It probes fundamental code execution and understanding capabilities beyond simple code generation. Use when the user wants to benchmark on CRUXEval, or asks about evaluating this task. Reports pass@1.

researchpythongo
0
3
Crumqs EvalA

Evaluates RAG systems on synthetic unanswerable and multi-hop queries to measure their refusal behavior, hallucination rates, and susceptibility to reasoning shortcuts when contexts are disjointed or out-of-distribution. Use when the user wants to benchmark on CRUMQs, UAEval4RAG, MultiHop-RAG, or asks about evaluating this task. Reports cheatability ratio.

researchpythongo
0
3
Crumb EvalA

Evaluates information retrieval models on complex, multi-aspect, and logically structured queries across eight diverse domains. It probes the model's ability to handle nuanced document alignments, set-based operations, and context-rich instructions beyond simple keyword matching. Use when the user wants to benchmark on CRUMB, or asks about evaluating this task. Reports nDCG@10.

researchpythongit
0
3
Crows Pairs EvalA

Measures social biases in masked language models by comparing the likelihood assigned to stereotypical versus anti-stereotypical sentence pairs. It quantifies how strongly models favor historically disadvantaged groups' stereotypes across nine demographic categories. Use when the user wants to benchmark on CrowS-Pairs, or asks about evaluating this task. Reports bias metric.

researchpythongo
0
3
Crown EvalA

Evaluates conversational passage ranking by measuring how effectively a model ranks relevant documents across multi-turn search queries, balancing term similarity with contextual coherence. Use when the user wants to benchmark on TREC CAsT 2019, or asks about evaluating this task. Reports nDCG.

researchpythonnode
0
3
Crowdspeech EvalA

Evaluates algorithms for aggregating multiple noisy, crowdsourced transcriptions of the same audio recording into a single high-quality reference. It probes how well methods handle sequential textual noise, estimate worker reliability, and adapt across different audio quality domains. Use when the user wants to benchmark on CROWDSPEECH, VOXDIY, CROWDWSA2019, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Crowdsensing Id Dfl EvalA

Evaluates the capability of decentralized federated learning (DFL) models to detect malware and classify benign states in IoT crowdsensing environments. It probes robustness under varying node counts, peer-to-peer network topologies, and data heterogeneity (IID vs. non-IID Dirichlet splits). Use when the user wants to benchmark on Crowdsensing Intrusion Detection Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Crowdhuman EvalA

Evaluates object detectors' ability to identify humans in highly crowded and heavily occluded scenes. It covers three annotation levels (full body, visible body, head) and assesses cross-dataset generalization for pedestrian and head detection tasks. Use when the user wants to benchmark on CrowdHuman, or asks about evaluating this task. Reports mMR.

researchpythongo
0
3
Crowdflow EvalA

Evaluates dense optical flow estimation accuracy and long-term temporal consistency in complex crowd surveillance scenarios, specifically testing robustness to non-rigid, self-occluding motion and small object tracking. Use when the user wants to benchmark on CrowdFlow, or asks about evaluating this task. Reports EPE.

researchpythontesting
0
3
Crowd Pose Estimation EvalA

Evaluates the ability of pose estimation models to accurately predict 2D keypoints for humans and animals in crowded, occluded, and multi-instance scenarios. It probes robustness to detection ambiguity, overlapping instances, and the transferability of conditional pose inputs from bottom-up detectors to top-down refiners. Use when the user wants to benchmark on CrowdPose, OCHuman, COCO, Multi-Animal (SchoolingFish, Marmosets, Tri-Mouse), or asks about evaluating this task. Reports AP.

researchpythonperformance
0
3
Crosswoz Dst EvalA

Evaluates the ability of generative dialogue state tracking models to accurately predict and maintain the complete set of user intents (domain-slot-value triples) across dialogue turns. It specifically probes cross-lingual and cross-ontology transfer capabilities by measuring how well models trained on one language or ontology generalize to another. Use when the user wants to benchmark on CrossWOZ-en, or asks about evaluating this task. Reports Joint Goal Accuracy.

researchpythongo
0
3
Crossvoice S2st EvalA

Evaluates cross-lingual speech-to-speech translation (S2ST) systems on translation accuracy and prosody preservation. It measures how well a cascade-based S2ST pipeline preserves speaker identity and naturalness while translating speech across different language pairs. Use when the user wants to benchmark on CVSS-T, Indic-TTS, Fisher, MuST-C, VoxPopuli, or asks about evaluating this task. Reports BLEU.

researchpython
0
3
Crosssum Alignment EvalA

Evaluates the quality of automatically induced cross-lingual summary alignments in the CrossSum dataset by measuring human agreement on whether two summaries correspond to the same source article. Use when the user wants to benchmark on CrossSum, or asks about evaluating this task. Reports alignment_accuracy.

researchpythongit
0
3
Crosspoint Bench EvalA

Evaluates Vision-Language Models' ability to perform precise point-level geometric correspondence across multiple viewpoints. It probes fine-grained spatial grounding, visibility reasoning, cross-view correspondence judgment, and continuous 2D coordinate pointing. Use when the user wants to benchmark on CrossPoint-Bench, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3
Crossnews Ua EvalA

Evaluates cross-lingual semantic similarity between news article pairs across four dimensions (Who, What, Where, When) in Ukrainian, Polish, Russian, and English. It probes a model's ability to align event-level information across languages while ignoring publication dates. Use when the user wants to benchmark on CrossNews-UA, or asks about evaluating this task. Reports macro-averaged F1-score.

researchpythongo
0
3
Crossner EvalA

Evaluates cross-domain named entity recognition by measuring how well models adapt from a source domain (CoNLL2003) to five specialized target domains. These domains feature unique, domain-specific entity types that test the model's ability to generalize beyond standard categories. Use when the user wants to benchmark on CrossNER, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Crossmodal 3600 EvalA

Evaluates multilingual image captioning models across 36 languages, probing their ability to generate stylistically coherent and culturally representative descriptions without relying on direct translation artifacts. It measures how well models generalize to low-resource and geographically diverse languages. Use when the user wants to benchmark on Crossmodal-3600, or asks about evaluating this task. Reports CIDEr.

researchpythonperformance
0
3
Crossmoda EvalA

Evaluates unsupervised cross-modality domain adaptation for medical image segmentation (Vestibular Schwannoma and Cochlea) and tumour grading (Koos classification) from ceT1 to T2 MRI. Use when the user wants to benchmark on crossMoDA, or asks about evaluating this task. Reports DSC.

researchpython
0
3
Crossmed EvalA

Evaluates compositional generalization in medical vision-language models across a structured Modality–Anatomy–Task (MAT) schema. It probes zero-shot cross-task transfer, generalization to novel MAT combinations, and robustness under low-data regimes using a unified visual question answering interface. Use when the user wants to benchmark on CrossMed, or asks about evaluating this task. Reports top-1 classification accuracy.

researchpythongo
0
3