Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,161–8,184 of 20,861 skills
Evaluates how well large language models align with culturally grounded Chinese value rules compared to Western benchmarks. It probes moral reasoning, preference alignment, and boundary separation across six sensitive themes like drugs, firearms, politics, and suicide. Use when the user wants to benchmark on CVC, or asks about evaluating this task. Reports preference.
Evaluates end-to-end inference latency and hardware efficiency of computer vision models on edge AI hardware. It probes how well a hardware-software co-design optimizes data movement and compute utilization under strict memory and bandwidth constraints. Use when the user wants to benchmark on ImageNet, COCO 2017, or asks about evaluating this task. Reports Latency [ms].
Evaluates end-to-end and cascaded named entity recognition from Arabic speech. It probes a model's ability to jointly transcribe spoken Arabic and predict fine-grained entity types (21 categories) using BIO-style tagging, as well as extract entity spans and values. Use when the user wants to benchmark on CV-18 NER, or asks about evaluating this task. Reports CoER.
Evaluates a kernel-level safety gateway's ability to correctly classify MCP tool-call prompts as dangerous or benign across 18 attack and benign domains. It probes the system's semantic understanding of tool intent versus surface-form rule matching, measuring how well the logit-based safety primitive prevents privilege escalation and adversarial bypasses. Use when the user wants to benchmark on Custom-101, or asks about evaluating this task. Reports F1.
Evaluates how the ordering of training data (curriculum) affects the quality of task-specific word embeddings. It probes whether optimized data sequencing improves downstream performance in sentiment analysis, NER, POS tagging, and parsing compared to random or heuristic orderings. Use when the user wants to benchmark on Wikipedia Paragraph Corpus, or asks about evaluating this task. Reports test results.
Evaluates text-to-image generation models on their ability to align generated images with text prompts, produce visually appealing outputs, and match human preferences. It tests the effectiveness of curriculum-based fine-tuning strategies on standard generative benchmarks. Use when the user wants to benchmark on D1 (Black-ICLR-2024), D2 (DrawBench), D3 (Pick-a-Pic), or asks about evaluating this task. Reports Text Alignment.
Evaluates the out-of-domain generalization and reasoning capabilities of vision-language models across visual detection, classification, and multimodal mathematical reasoning tasks, alongside standard multimodal benchmarks. Use when the user wants to benchmark on RefCOCO, RefGTA, Pascal-VOC, Math360K, CLEVER-70k-Counting, MathVista, MATH, AI2D, MMBench, MMVet, OCRBench, LLaVABench, or asks about evaluating this task. Reports accuracy.
Tests a model's ability to learn sequentially across multiple NLP tasks while retaining prior knowledge. It specifically probes catastrophic forgetting mitigation during continual fine-tuning by measuring performance drops on earlier tasks after learning new ones. Use when the user wants to benchmark on GLUE-MRPC, GLUE-SST-2, Sentiment140, WikiText-2, or asks about evaluating this task. Reports Accuracy.
Evaluates continual learning capabilities in language models by measuring skill retention, forward/backward transfer, and catastrophic forgetting across a developmental skill graph spanning ages 5–10. It probes how sequential, joint, and independent training affect performance on instruction, context-question-answer, and context-sentence-question-answer tasks. Use when the user wants to benchmark on CurLL, or asks about evaluating this task. Reports LLM rating score (1-5).
Evaluates automated red-teaming methods on their ability to generate diverse and effective prompts that elicit toxic responses from target LLMs. It probes both the effectiveness (toxicity elicitation rate) and diversity (textual and semantic variation) of generated test cases across text continuation and instruction-following tasks. Use when the user wants to benchmark on IMDb review dataset, Alpaca dataset, Databricks dataset, or asks about evaluating this task. Reports toxic response rate.
This benchmark evaluates multimodal large language models' ability to perform clinical differential diagnosis using patient history and medical images. It explicitly disentangles intrinsic diagnostic reasoning from external evidence retrieval by testing models under four paradigms: no context, physician-curated references, standard RAG, and agentic web search. The protocol probes how well models leverage retrieved literature versus relying on internal knowledge, and how retrieval noise impact...
Probes a clinical language model's ability to predict binary adverse outcomes from free-text EHR notes while simultaneously calibrating its prediction uncertainty. It evaluates both discriminative accuracy and probabilistic calibration across multiple clinical risk stratification tasks. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUROC.
Evaluates machine translation quality for Czech-Ukrainian and Ukrainian-Czech translation using constrained back-translation systems and a proprietary unconstrained system. Probes the impact of data preprocessing techniques like romanization and ensemble methods on translation performance. Use when the user wants to benchmark on Flores 101 development set, WMT22 Czech-Ukrainian test set, or asks about evaluating this task. Reports BLEU.
Evaluates piecewise-stationary multi-armed bandit algorithms by measuring the expected cumulative regret over a sequence of time steps. It probes how well an algorithm adapts to changing arm reward distributions (change-points) while balancing exploration and exploitation. Use when the user has predictions and gold and needs to compute cumulative regret.
Evaluates multilingual content safety guard models on their ability to detect harmful or unsafe prompts and responses across diverse languages and cultural contexts, including zero-shot generalization to unseen languages. Use when the user wants to benchmark on CultureGuard, PolyGuardPrompts, RTP-LX, MultiJail, XSafety, Aya Red-teaming, or asks about evaluating this task. Reports harmful-F1.
Evaluates whether an LLM's value profile aligns with specific cultural norms using World Values Survey data, and tests the model's steerability when provided with diverse cultural contexts. It probes the extent to which constitutional AI codifies dominant cultural biases and resists prompt-based cultural adaptation. Use when the user wants to benchmark on World Values Survey (WVS) Wave 7, or asks about evaluating this task. Reports Pearson correlation.
This benchmark evaluates how well multilingual LLMs preserve cultural nuance, idioms, puns, and culturally embedded concepts during machine translation. It probes the persistent gap between grammatical accuracy and cultural resonance by measuring translation quality across different figurative and non-figurative segment categories. Use when the user wants to benchmark on Cultural Nuance MT Benchmark, or asks about evaluating this task. Reports overall quality.
Assesses the ability of large multimodal models to identify the geographical origin (country, subregion, or continent) of an image based on visual cultural cues. It probes implicit stereotypical associations and geographic bias in vision-language models. Use when the user wants to benchmark on Dalle Street, Dollar Street, MaRVL, or asks about evaluating this task. Reports classification accuracy.
Evaluates machine translation systems on culturally specific items (CSIs) to measure how well they preserve cultural nuances, entities, and meanings compared to reference translations. It probes both automated lexical/semantic alignment and human judgment on translation accuracy for culturally grounded content. Use when the user wants to benchmark on Wikipedia Cultural Parallel Corpus, or asks about evaluating this task. Reports CSI-Match.
This benchmark probes a machine translation model's ability to accurately translate culture-loaded expressions (idioms, proverbs, culture-specific items) while preserving their figurative, contextual, and cultural meanings. It evaluates whether models can avoid literal or superficial translations that strip away culturally grounded nuances. Use when the user wants to benchmark on CulT-Eval, or asks about evaluating this task. Reports ACRE.
Probes LLMs' cross-cultural emotion understanding by testing their ability to predict emotions and sentiments across six languages, specifically examining how prompt language and explicit country context influence model performance. Use when the user wants to benchmark on CULEMO, or asks about evaluating this task. Reports emotion prediction.
Evaluates lane detection performance in diverse urban and highway scenarios using an F1-measure based on IoU between predicted and ground truth lane lines. Use when the user wants to benchmark on CULane, or asks about evaluating this task. Reports F1-measure.
Evaluates Chinese language understanding and generation capabilities across a hierarchical framework. It probes discourse comprehension, conversational interaction, mathematical reasoning, and multilingual tasks using a multi-level scoring strategy that normalizes model performance against a fixed baseline. Use when the user wants to benchmark on CUGE (lite version), or asks about evaluating this task. Reports normalized capability performance.
This evaluation probes the per-evidence-item utility and trace sensitivity in single-shot retrieval-augmented generation. It measures how removing, replacing, or duplicating retrieved context chunks affects answer correctness, grounding faithfulness, confidence calibration, and reasoning trace stability. Use when the user wants to benchmark on HotpotQA (distractor setting), 2WikiMultihopQA, or asks about evaluating this task. Reports Soft Correctness.