Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,865
skills in category
870
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,185–8,208 of 20,865 skills

Culemo EvalA

Probes LLMs' cross-cultural emotion understanding by testing their ability to predict emotions and sentiments across six languages, specifically examining how prompt language and explicit country context influence model performance. Use when the user wants to benchmark on CULEMO, or asks about evaluating this task. Reports emotion prediction.

researchpythongo
0
3
Culane EvalA

Evaluates lane detection performance in diverse urban and highway scenarios using an F1-measure based on IoU between predicted and ground truth lane lines. Use when the user wants to benchmark on CULane, or asks about evaluating this task. Reports F1-measure.

researchpythonperformance
0
3
Cuge EvalA

Evaluates Chinese language understanding and generation capabilities across a hierarchical framework. It probes discourse comprehension, conversational interaction, mathematical reasoning, and multilingual tasks using a multi-level scoring strategy that normalizes model performance against a fixed baseline. Use when the user wants to benchmark on CUGE (lite version), or asks about evaluating this task. Reports normalized capability performance.

researchpythongo
0
3
Cue R EvalA

This evaluation probes the per-evidence-item utility and trace sensitivity in single-shot retrieval-augmented generation. It measures how removing, replacing, or duplicating retrieved context chunks affects answer correctness, grounding faithfulness, confidence calibration, and reasoning trace stability. Use when the user wants to benchmark on HotpotQA (distractor setting), 2WikiMultihopQA, or asks about evaluating this task. Reports Soft Correctness.

researchpythongo
0
3
Cubert Fine Tuning EvalA

Evaluates contextual code embeddings on six Python source-code understanding tasks, including classification of variable misuse, incorrect binary operators, swapped operands, function-docstring mismatches, and exception types, plus a joint localization and repair task. Use when the user wants to benchmark on ETH Py150 Open Benchmarks, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Cuad EvalA

Evaluates a model's ability to identify and extract relevant text spans from legal contracts corresponding to specific clause categories. It probes domain-specific information extraction and needle-in-a-haystack detection under severe class imbalance. Use when the user wants to benchmark on CUAD, or asks about evaluating this task. Reports Precision@80% Recall.

researchpythongo
0
3
Ctta Text Understanding EvalA

Evaluates continual test-time adaptation (CTTA) for text understanding across sequential, unobserved domains. It probes a model's ability to adapt to shifting domains using only unlabeled test data while mitigating error accumulation and maintaining cross-domain generalization. Use when the user wants to benchmark on CTTA-Text-Understanding-Benchmark, or asks about evaluating this task. Reports exact match (EM), F1 score.

researchpythongo
0
3
Ctr Welfare EvalA

Evaluates click-through rate (CTR) prediction models for their ability to maximize economic welfare in simulated and real-world ad auction settings, while also measuring standard classification performance. Use when the user wants to benchmark on Synthetic Dataset, Criteo Display Advertising Challenge, or asks about evaluating this task. Reports test-time welfare.

researchpythonperformance
0
3
Ctr Prediction EvalA

Evaluates the ability of deep learning models to predict click-through rates (CTR) from sparse, high-dimensional categorical features in advertising and recommendation scenarios. It probes how well models capture multi-scale semantic interactions and handle large-scale, imbalanced binary classification tasks typical of real-world ad systems. Use when the user wants to benchmark on Avazu, MovieLens, Weibo, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Ctr Prediction Dhan EvalA

Evaluates a model's ability to predict the next item a user will click or review based on their historical interaction sequence. It probes hierarchical user interest modeling across multiple product dimensions and abstraction levels in recommendation systems. Use when the user wants to benchmark on Amazon Review (Six-Category, Kindle Shop, Electronics), or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Cti Plausibility EvalA

Evaluates whether neural machine translation models correctly rely on contextual cues when generating target tokens. It compares model-extracted cue-target pairs against human-annotated discourse-level expectations to measure the plausibility of context reliance. Use when the user wants to benchmark on SCAT+, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Ct Rate Zero Shot EvalA

This benchmark evaluates the zero-shot multi-abnormality detection capability of a visual-language foundation model on 3D chest CT volumes. It probes the model's ability to generalize to unseen data distributions and classify multiple pathologies simultaneously without task-specific supervised training. Use when the user wants to benchmark on CT-RATE, RAD-ChestCT, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Ct Brain Segmentation EvalA

Evaluates the ability of segmentation models to accurately delineate brain tissue, cerebrospinal fluid (CSF), and subdural hematomas in post-operative CT scans of hydrocephalic infants. It probes robustness to intensity overlap, anatomical distortion, and limited training data in a real-world clinical setting. Use when the user wants to benchmark on CURE Children's Hospital of Uganda CT Brain Dataset, or asks about evaluating this task. Reports dice-overlap coefficient.

researchpython
0
3
Csymr EvalA

Evaluates compositional symbolic music reasoning by requiring models to chain atomic analyses across multiple musical dimensions (e.g., rhythm, harmony, key, structure) to answer multiple-choice questions derived from expert forums and professional exams. Use when the user wants to benchmark on CSyMR-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Csvqa EvalA

This benchmark evaluates the scientific reasoning and domain-grounded visual question answering capabilities of Vision-Language Models (VLMs) in Chinese. It probes the ability to integrate multimodal STEM evidence across physics, chemistry, biology, and mathematics with domain knowledge to solve both multiple-choice and open-ended questions. Use when the user wants to benchmark on CSVQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Css10 Tts EvalA

Evaluates the quality of synthesized speech from single-speaker TTS models trained on the CSS10 datasets across 10 languages. It probes how well models can reproduce natural-sounding audio and accurate pronunciation for held-out test sentences. Use when the user wants to benchmark on CSS10, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Csrec Sequential Rec EvalA

This evaluation protocol assesses the ranking performance and robustness of sequential recommendation models trained with confident soft labels. It measures whether predicted item sequences align with actual user interactions and verifies if recommendations correspond to genuinely positive user preferences using explicit rating thresholds. Use when the user wants to benchmark on Last.FM, Yelp, Amazon Electronics, Amazon Movies and TV, or asks about evaluating this task. Reports Recall@n, NDCG@n.

researchpythongo
0
3
Csr L EvalA

Evaluates the robustness of information retrieval models when processing code-switched queries (English mixed with Mandarin Chinese or Japanese). It probes whether multilingual retrievers and rerankers suffer embedding divergence or performance degradation compared to monolingual English queries across argument, code, biomedical, and instruction-following retrieval tasks. Use when the user wants to benchmark on Touché 2020, HumanEval, TRECCOVID, FollowIR, or asks about evaluating this task. R...

researchpythonperformance
0
3
Csi Bert2 EvalA

Evaluates a transformer-based framework for Channel State Information (CSI) time-series prediction and wireless sensing classification. It probes the model's ability to recover missing data, predict future CSI sequences, and classify human actions or environmental states from Wi-Fi signals. Use when the user wants to benchmark on WiGesture, WiFall, WiCount, CommPre, or asks about evaluating this task. Reports Accuracy.

researchpython
0
3
Csegg EvalA

Evaluates continual learning capabilities in scene graph generation by measuring how models retain prior object-relationship knowledge while learning new tasks, handle long-tailed data distributions, and generalize to unseen objects and relationships across incremental learning scenarios. Use when the user wants to benchmark on CSEGG, or asks about evaluating this task. Reports Avg. R@20.

researchpythongo
0
3
Csaw M EvalA

Evaluates models on ordinal classification of mammographic masking potential (levels 1–8) and their clinical utility in predicting interval and large invasive cancers. It probes the model's ability to respect ordinal relationships in breast tissue obscuration and correlate these estimates with cancer outcomes. Use when the user wants to benchmark on CSAW-M, or asks about evaluating this task. Reports average mean absolute error (AMAE).

researchpythongo
0
3
Cs Kws EvalA

Evaluates a model's ability to detect and localize multiple spoken keywords within continuous, untrimmed audio streams, distinguishing target keywords from background speech and silence. Use when the user wants to benchmark on LibriTop-20, CMAK-7, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Cs Dialogue Asr EvalA

This benchmark evaluates automatic speech recognition (ASR) systems on their ability to accurately transcribe spontaneous, full-length dialogues that alternate between Mandarin and English. It probes a model's robustness to language alternation, phonetic mismatches, and contextual dependencies in naturalistic code-switching scenarios. Use when the user wants to benchmark on CS-Dialogue, or asks about evaluating this task. Reports MER.

researchpythonperformance
0
3
Cs 4k EvalA

Evaluates LLMs on end-to-end computer science research workflows by testing their ability to answer scientific questions grounded in academic papers. It probes domain-specific reasoning, factual recall, and methodological understanding across eight research workflow categories. Use when the user wants to benchmark on CS-4k, or asks about evaluating this task. Reports model response score.

researchpythongo
0
3