Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,017–8,040 of 20,861 skills
Evaluates machine translation systems on four discourse phenomena: anaphora resolution, lexical consistency, coherence/readability, and discourse connectives. It probes whether context-aware models can maintain discourse-level quality and consistency across different language pairs beyond standard n-gram overlap. Use when the user wants to benchmark on DiP Benchmark, or asks about evaluating this task. Reports BLEU.
Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
Evaluates the cross-domain generalization and scaling behavior of a natural-image pre-trained vision transformer (DINOv3) across diverse medical imaging modalities, including 2D/3D classification and segmentation tasks. Use when the user wants to benchmark on NIH-14, RSNA-Pneumonia, Camelyon16, Camelyon17, BCNB, Kvasir-Capsule, AutoLaparo, EndoVis18, EDD 2020, CT-RATE, Medical Segmentation Decathlon (MSD), CREMI, AC3/4, AutoPET-II, HECKTOR 2022, or asks about evaluating this task. Reports AUC...
Evaluates the cross-task generalizability of the DINOv2 vision foundation model on medical image analysis tasks, specifically disease classification and organ segmentation across X-ray, CT, and MRI modalities. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, SARS-CoV-2, Brain Tumor, Montgomery County (MC), AMOS, MSD Heart, MSD Hipp, MSD Spleen, or asks about evaluating this task. Reports AUROC.
This benchmark evaluates Vision-Language Models on fine-grained visual discrimination, nutritional quantification from images, and complex food-related visual question answering. It probes the models' ability to fuse multi-view imagery, perform volumetric reasoning, and avoid parametric knowledge biases when identifying dishes and estimating macronutrients. Use when the user wants to benchmark on DiningBench, or asks about evaluating this task. Reports Accuracy, MAPE.
This benchmark evaluates the robustness of network intrusion detection models against distribution shifts between different network environments. It specifically probes cross-domain generalization by training on one NetFlow dataset and testing on another, measuring how well domain-invariant feature extraction mitigates performance degradation when facing unseen attack distributions. Use when the user wants to benchmark on NFv2-UNSW-NB15, NFv2-CIC-2018, or asks about evaluating this task. Repo...
Evaluates a model's ability to generate syntactically and semantically correct SQL queries from natural language questions across diverse database schemas. It probes schema linking, handling of complex query structures (joins, nested subqueries, aggregations), and the capacity for iterative self-correction when initial generations fail. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports Execution Accuracy (EX).
Evaluates multilingual and multidomain aspect-based sentiment analysis by predicting continuous valence-arousal (VA) scores alongside aspect, opinion, and category extraction. It probes a model's ability to perform fine-grained dimensional sentiment regression and structured information extraction across diverse languages and domains. Use when the user wants to benchmark on DimABSA, or asks about evaluating this task. Reports RMSE_VA, cF1.
Evaluates visuomotor policy learning in robotics by measuring how well a model generates sequential actions to complete manipulation tasks under both state and image observations. It probes the policy's ability to handle multimodal action distributions, long-horizon dependencies, and latency robustness across rigid and fluid object manipulation. Use when the user wants to benchmark on Robomimic, Push-T, Block Push, Franka Kitchen, or asks about evaluating this task. Reports success_rate.
This protocol evaluates the vision-language alignment and zero-shot generalization capabilities of fine-tuned VLMs across diverse multimodal tasks. It measures how effectively aligning VLM cross-attention with diffusion model attention maps improves performance on document understanding, reasoning, real-world visual comprehension, and hallucination detection benchmarks. Use when the user wants to benchmark on AI2D, ChartQA, OCRBench, DocVQA, InfoVQA, MME, MMBench, ScienceQA, MMStar, MMMU, Rea...
Evaluates a model's ability to predict the next item in a user's interaction sequence by capturing dynamic preferences and multi-aspect item representations. It probes sequential recommendation performance under varying sequence lengths and item popularities. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Movielens-1M, Steam, or asks about evaluating this task. Reports HR@K.
Pixel-level localization of diffusion-based AI edits in images, shifting from whole-image classification to semantic segmentation to identify precisely which regions have been altered by generative models. Use when the user wants to benchmark on DiffSeg30k, or asks about evaluating this task. Reports localization accuracy.
This evaluation probes an adversarial auditing framework where a blue team must identify a compromised model among a pair of nearly identical models. It tests the ability to detect hidden backdoors, misaligned behaviors, or injected instructions using various probing strategies under varying levels of prior knowledge. Use when the user wants to benchmark on CIFAR-10, Truthful QA, HHH, or asks about evaluating this task. Reports accuracy.
Evaluates whether LLMs recognize meaningful demographic group differences (Difference Awareness) and understand when differential treatment is contextually appropriate (Contextual Awareness), challenging the standard 'color-blind' fairness paradigm. Use when the user wants to benchmark on DiffAware and CtxtAware Benchmark Suite, or asks about evaluating this task. Reports win rate.
This benchmark evaluates a language model's ability to disentangle factual gender knowledge from gender bias in masked language modeling. It measures whether a model can correctly predict gendered tokens in gender-specific contexts while remaining gender-neutral in gender-neutral contexts, revealing the trade-off between fairness and factual performance. Use when the user wants to benchmark on DIFAIR, or asks about evaluating this task. Reports GIS.
Evaluates a model's ability to predict click-through rates (CTR) by modeling sequential user behavior and dynamically evolving latent interests relative to a target item. Use when the user wants to benchmark on Amazon Books, Amazon Electronics, Industrial (Taobao), or asks about evaluating this task. Reports AUC.
Evaluates a disease-weighted attention refinement framework for medical image analysis. It probes the model's ability to align cross-attention maps with radiologist annotations (grounding) while maintaining diagnostic accuracy on chest X-ray classification tasks. Use when the user has predictions and gold and needs to compute Dice Coefficient.
Evaluates large language models' financial reasoning capabilities and general problem-solving skills across multiple benchmarks. It measures how well models can answer domain-specific financial questions and general math/science reasoning tasks, while also assessing compliance rule adherence in Chinese financial contexts. Use when the user wants to benchmark on CFLUE, FinQA, CCC, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.
Evaluates Theory of Mind and participant-centric reasoning in multi-party dialogues by testing a model's ability to track dynamic numerical variables, filter distractors, and reason from specific character perspectives (including false beliefs) rather than using omniscient context. Use when the user wants to benchmark on DIAMONDs, or asks about evaluating this task. Reports accuracy.
Evaluates dialogue segmentation on a benchmark constructed by joining disparate task-oriented dialogues. It probes the model's ability to detect abrupt, artificial context shifts and identify segment boundaries in synthetic multi-intent conversations. Use when the user wants to benchmark on DialSeg711, or asks about evaluating this task. Reports Pk.
Evaluates the quality of unsupervised abstractive dialogue summarization across multiple domains by comparing generated summaries against human references using standard n-gram and LCS overlap metrics. It tests the model's ability to compress and rephrase conversational transcripts into coherent summaries without training data. Use when the user wants to benchmark on AMI, ICSI, DialogSum, SAMSum, MediaSum, SummScreen, ADS, or asks about evaluating this task. Reports ROUGE-1.
Evaluates the robustness of offensive language detection models against adversarial human attacks in single-turn and multi-turn dialogue contexts. It measures classifier resilience when exposed to iterative, context-aware attacks designed to evade safety filters. Use when the user wants to benchmark on Wikipedia Toxic Comments, or asks about evaluating this task. Reports Weighted-F1.
Evaluates a model's ability to generate task-oriented and knowledge-grounded dialogue responses in zero-shot and few-shot settings. It measures lexical overlap and unigram F1 against ground-truth responses to assess generalization across open-domain and multi-domain conversational tasks. Use when the user wants to benchmark on CoQA, MultiWOZ 2.2, or asks about evaluating this task. Reports ROUGE-L.
Evaluates large language models' ability to understand and reason across multiple Arabic dialects and standard Arabic across diverse academic and professional domains. It measures dialectal generalization and sensitivity to linguistic context by comparing performance under default, dialect-conditioned, and dialect-identification prompts. Use when the user wants to benchmark on DialectalArabicMMLU, or asks about evaluating this task. Reports accuracy.