Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,861
skills in category
870
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,161–8,184 of 20,861 skills

Cvc Value Alignment EvalA

Evaluates how well large language models align with culturally grounded Chinese value rules compared to Western benchmarks. It probes moral reasoning, preference alignment, and boundary separation across six sensitive themes like drugs, firearms, politics, and suicide. Use when the user wants to benchmark on CVC, or asks about evaluating this task. Reports preference.

researchpythongo
0
3
Cv Inference EvalA

Evaluates end-to-end inference latency and hardware efficiency of computer vision models on edge AI hardware. It probes how well a hardware-software co-design optimizes data movement and compute utilization under strict memory and bandwidth constraints. Use when the user wants to benchmark on ImageNet, COCO 2017, or asks about evaluating this task. Reports Latency [ms].

researchpythonperformance
0
3
Cv 18 Ner EvalA

Evaluates end-to-end and cascaded named entity recognition from Arabic speech. It probes a model's ability to jointly transcribe spoken Arabic and predict fine-grained entity types (21 categories) using BIO-style tagging, as well as extract entity spans and values. Use when the user wants to benchmark on CV-18 NER, or asks about evaluating this task. Reports CoER.

researchpythongo
0
3
Custom 101 EvalA

Evaluates a kernel-level safety gateway's ability to correctly classify MCP tool-call prompts as dangerous or benign across 18 attack and benign domains. It probes the system's semantic understanding of tool intent versus surface-form rule matching, measuring how well the logit-based safety primitive prevents privilege escalation and adversarial bypasses. Use when the user wants to benchmark on Custom-101, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Curriculum Word2vec EvalA

Evaluates how the ordering of training data (curriculum) affects the quality of task-specific word embeddings. It probes whether optimized data sequencing improves downstream performance in sentiment analysis, NER, POS tagging, and parsing compared to random or heuristic orderings. Use when the user wants to benchmark on Wikipedia Paragraph Corpus, or asks about evaluating this task. Reports test results.

researchpythongo
0
3
Curriculum Dpo++ EvalA

Evaluates text-to-image generation models on their ability to align generated images with text prompts, produce visually appealing outputs, and match human preferences. It tests the effectiveness of curriculum-based fine-tuning strategies on standard generative benchmarks. Use when the user wants to benchmark on D1 (Black-ICLR-2024), D2 (DrawBench), D3 (Pick-a-Pic), or asks about evaluating this task. Reports Text Alignment.

researchpython
0
3
Curr Reft EvalA

Evaluates the out-of-domain generalization and reasoning capabilities of vision-language models across visual detection, classification, and multimodal mathematical reasoning tasks, alongside standard multimodal benchmarks. Use when the user wants to benchmark on RefCOCO, RefGTA, Pascal-VOC, Math360K, CLEVER-70k-Counting, MathVista, MATH, AI2D, MMBench, MMVet, OCRBench, LLaVABench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Curlora Continual EvalA

Tests a model's ability to learn sequentially across multiple NLP tasks while retaining prior knowledge. It specifically probes catastrophic forgetting mitigation during continual fine-tuning by measuring performance drops on earlier tasks after learning new ones. Use when the user wants to benchmark on GLUE-MRPC, GLUE-SST-2, Sentiment140, WikiText-2, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Curll EvalA

Evaluates continual learning capabilities in language models by measuring skill retention, forward/backward transfer, and catastrophic forgetting across a developmental skill graph spanning ages 5–10. It probes how sequential, joint, and independent training affect performance on instruction, context-question-answer, and context-sentence-question-answer tasks. Use when the user wants to benchmark on CurLL, or asks about evaluating this task. Reports LLM rating score (1-5).

researchpythonperformance
0
3
Curiosity Redteam EvalA

Evaluates automated red-teaming methods on their ability to generate diverse and effective prompts that elicit toxic responses from target LLMs. It probes both the effectiveness (toxicity elicitation rate) and diversity (textual and semantic variation) of generated test cases across text continuation and instruction-following tasks. Use when the user wants to benchmark on IMDb review dataset, Alpaca dataset, Databricks dataset, or asks about evaluating this task. Reports toxic response rate.

researchpythonperformance
0
3
Cure EvalA

This benchmark evaluates multimodal large language models' ability to perform clinical differential diagnosis using patient history and medical images. It explicitly disentangles intrinsic diagnostic reasoning from external evidence retrieval by testing models under four paradigms: no context, physician-curated references, standard RAG, and agentic web search. The protocol probes how well models leverage retrieved literature versus relying on internal knowledge, and how retrieval noise impact...

researchpythongo
0
3
Cura Mimic Iv EvalA

Probes a clinical language model's ability to predict binary adverse outcomes from free-text EHR notes while simultaneously calibrating its prediction uncertainty. It evaluates both discriminative accuracy and probabilistic calibration across multiple clinical risk stratification tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Cuni Wmt22 Csuk EvalA

Evaluates machine translation quality for Czech-Ukrainian and Ukrainian-Czech translation using constrained back-translation systems and a proprietary unconstrained system. Probes the impact of data preprocessing techniques like romanization and ensemble methods on translation performance. Use when the user wants to benchmark on Flores 101 development set, WMT22 Czech-Ukrainian test set, or asks about evaluating this task. Reports BLEU.

researchpythonperformance
0
3
Cumulative RegretA

Evaluates piecewise-stationary multi-armed bandit algorithms by measuring the expected cumulative regret over a sequence of time steps. It probes how well an algorithm adapts to changing arm reward distributions (change-points) while balancing exploration and exploitation. Use when the user has predictions and gold and needs to compute cumulative regret.

researchpythongo
0
3
Cultureguard Multilingual Safety EvalA

Evaluates multilingual content safety guard models on their ability to detect harmful or unsafe prompts and responses across diverse languages and cultural contexts, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on CultureGuard, PolyGuardPrompts, RTP-LX, MultiJail, XSafety, Aya Red-teaming, or asks about evaluating this task. Reports harmful-F1.

researchpythongo
0
3
Cultural Positioning EvalA

Evaluates whether an LLM's value profile aligns with specific cultural norms using World Values Survey data, and tests the model's steerability when provided with diverse cultural contexts. It probes the extent to which constitutional AI codifies dominant cultural biases and resists prompt-based cultural adaptation. Use when the user wants to benchmark on World Values Survey (WVS) Wave 7, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Cultural Nuance Mt EvalA

This benchmark evaluates how well multilingual LLMs preserve cultural nuance, idioms, puns, and culturally embedded concepts during machine translation. It probes the persistent gap between grammatical accuracy and cultural resonance by measuring translation quality across different figurative and non-figurative segment categories. Use when the user wants to benchmark on Cultural Nuance MT Benchmark, or asks about evaluating this task. Reports overall quality.

researchpythongo
0
3
Cultural Awareness EvalA

Assesses the ability of large multimodal models to identify the geographical origin (country, subregion, or continent) of an image based on visual cultural cues. It probes implicit stereotypical associations and geographic bias in vision-language models. Use when the user wants to benchmark on Dalle Street, Dollar Street, MaRVL, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Cultural Aware Mt EvalA

Evaluates machine translation systems on culturally specific items (CSIs) to measure how well they preserve cultural nuances, entities, and meanings compared to reference translations. It probes both automated lexical/semantic alignment and human judgment on translation accuracy for culturally grounded content. Use when the user wants to benchmark on Wikipedia Cultural Parallel Corpus, or asks about evaluating this task. Reports CSI-Match.

researchpythongo
0
3
Cult Eval EvalA

This benchmark probes a machine translation model's ability to accurately translate culture-loaded expressions (idioms, proverbs, culture-specific items) while preserving their figurative, contextual, and cultural meanings. It evaluates whether models can avoid literal or superficial translations that strip away culturally grounded nuances. Use when the user wants to benchmark on CulT-Eval, or asks about evaluating this task. Reports ACRE.

researchpythonexpress
0
3
Culemo EvalA

Probes LLMs' cross-cultural emotion understanding by testing their ability to predict emotions and sentiments across six languages, specifically examining how prompt language and explicit country context influence model performance. Use when the user wants to benchmark on CULEMO, or asks about evaluating this task. Reports emotion prediction.

researchpythongo
0
3
Culane EvalA

Evaluates lane detection performance in diverse urban and highway scenarios using an F1-measure based on IoU between predicted and ground truth lane lines. Use when the user wants to benchmark on CULane, or asks about evaluating this task. Reports F1-measure.

researchpythonperformance
0
3
Cuge EvalA

Evaluates Chinese language understanding and generation capabilities across a hierarchical framework. It probes discourse comprehension, conversational interaction, mathematical reasoning, and multilingual tasks using a multi-level scoring strategy that normalizes model performance against a fixed baseline. Use when the user wants to benchmark on CUGE (lite version), or asks about evaluating this task. Reports normalized capability performance.

researchpythongo
0
3
Cue R EvalA

This evaluation probes the per-evidence-item utility and trace sensitivity in single-shot retrieval-augmented generation. It measures how removing, replacing, or duplicating retrieved context chunks affects answer correctness, grounding faithfulness, confidence calibration, and reasoning trace stability. Use when the user wants to benchmark on HotpotQA (distractor setting), 2WikiMultihopQA, or asks about evaluating this task. Reports Soft Correctness.

researchpythongo
0
3