Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,409–6,432 of 20,820 skills
This benchmark probes an evaluator model's ability to accurately judge preference between two model responses while resisting bias towards superficial qualities like verbosity, fluency, and formality. It measures instruction-following accuracy by comparing judgments on natural preference data against adversarially crafted instances designed to confound less capable judges. Use when the user wants to benchmark on LLMBar, or asks about evaluating this task. Reports accuracy.
Evaluates LLM harmlessness and trustworthiness across four domains: Bias, Hate, Illegal content, and Sensitiveness. It measures the model's ability to correctly identify harmful or biased prompts and respond appropriately using a multiple-choice format where safe or neutral responses are designated as correct. Use when the user wants to benchmark on LLM Trustworthiness Benchmark, or asks about evaluating this task. Reports accuracy.
Evaluates the throughput, latency, and scalability of LLM inference serving systems under varying request rates and context lengths. It probes how efficiently a system manages KV cache, batching, and resource allocation for both short and long-context instruction-following workloads. Use when the user wants to benchmark on Alpaca, LongBench, or asks about evaluating this task. Reports Throughput.
Evaluates the intrinsic self-correction capability of LLMs across safety, reasoning, and vision-language tasks. It measures how iterative self-refinement reduces model uncertainty and improves calibration, toxicity mitigation, and bias reduction. Use when the user wants to benchmark on AdvBench, CommonGen-Hard, BBQ, MMVP, MS-COCO (Visual Grounding Subset), Real Toxicity Prompts, or asks about evaluating this task. Reports toxicity_score.
Evaluates the end-to-end latency and answer accuracy of LLM-based search agents operating in a ReAct workflow with external Wikipedia API calls. It probes the system's ability to balance speculative action execution with verification to reduce inference time while maintaining multi-hop reasoning quality. Use when the user wants to benchmark on HotPotQA, 2WikiMultihopQA, TriviaQA, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to answer open-ended protein questions accurately and predict Enzyme Commission (EC) numbers hierarchically. It probes semantic understanding of biological knowledge and fine-grained functional classification across multiple taxonomic levels. Use when the user has predictions and gold and needs to compute LLM-Score, Hierarchical Micro-F1.
Evaluates LLM safety and robustness against adversarial attacks by measuring overall safety scores, attack success rates, and toxicity levels. It also assesses whether safety alignment preserves general capabilities across standard reasoning, instruction-following, and knowledge benchmarks. Use when the user wants to benchmark on ALERT, LLM Leaderboard, or asks about evaluating this task. Reports Safety Score S.
Evaluates the performance and behavioral characteristics of Large Language Models when deployed as recommender systems. It probes traditional recommendation accuracy and novelty, alongside LLM-specific traits like history length sensitivity, candidate position bias, and hallucination rates. Use when the user wants to benchmark on Unspecified recommendation datasets (four datasets referenced in paper), or asks about evaluating this task. Reports HR.
This evaluation protocol assesses the mathematical reasoning and multi-step problem-solving capabilities of LLMs after reinforcement learning fine-tuning. It measures how well models generalize from a specialized arithmetic training task to standard academic and competitive math benchmarks. Use when the user wants to benchmark on GSM8K, BBH, MATH, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
Evaluates memory-efficient gradient compression optimizers against full-rank baselines during LLM pre-training and fine-tuning, measuring final model quality, convergence speed, memory footprint, and training throughput. Use when the user wants to benchmark on C4, MMLU, GLUE, or asks about evaluating this task. Reports Validation PPL, Accuracy.
Evaluates whether an LLM-based pairwise comparison framework can effectively identify high-impact academic papers compared to human peer review and traditional rating-based LLM methods. It probes the system's predictive accuracy for future scholarly influence, decision consistency with human committees, and susceptibility to biases in topic novelty and institutional representation. Use when the user wants to benchmark on OpenReview Conference Papers (ICLR, NeurIPS, CoRL, EMNLP), or asks about...
This evaluation protocol assesses the performance of an on-device LLM inference framework (llm.npu) on mobile NPUs. It probes the system's ability to accelerate prefill and decoding stages, manage quantization without accuracy loss, and optimize energy efficiency across various LLM sizes and real-world task datasets. Use when the user wants to benchmark on LAMBADA, HellaSwag, WinoGrande, MMLU, LongBench, DroidTask, Persona-Chat, or asks about evaluating this task. Reports end-to-end latency.
Evaluates the utility of LLM text generation across code, math, and summarization tasks under inference budget constraints. It probes whether jointly tuning generation hyperparameters (e.g., temperature, top-p, number of responses) improves task performance compared to default or benchmark configurations. Use when the user wants to benchmark on APPS, HumanEval, MATH, XSum, or asks about evaluating this task. Reports pass_rate (code).
Evaluates how reducing the size of the language model component in multimodal models impacts task performance, specifically isolating and measuring the bottlenecks in visual perception versus logical reasoning across multiple benchmarks. Use when the user wants to benchmark on Grounding, NIGHTS, PieAPP, OCR-VQA, Fine-grained Perception, Logical Reasoning, Math, Science & Technology, or asks about evaluating this task. Reports performance.
This evaluation probes whether explicitly unbiased LLMs exhibit automatic, stereotype-driven preferences in decision-making scenarios. It measures implicit bias by comparing model agreement rates across stereotypical versus counter-stereotypical social categories (e.g., gender-career, race-health) using a single-choice decision prompt rather than a relative comparison. Use when the user wants to benchmark on LLM Decision Bias (Absolute Variant), or asks about evaluating this task. Reports nor...
This evaluation probes the propensity of large language models to generate stereotypical or biased predictions across multiple demographic and social categories. It measures how often models align with human-annotated stereotypes versus anti-stereotypes or neutral alternatives when completing masked sentences or answering repurposed benchmark questions. Use when the user wants to benchmark on StereoSet, WinoBias, UnQover, CrowS-Pairs, Real Toxicity Prompts (RTP), Equity Evaluation Corpus (EEC...
Evaluates a language model's general capabilities, including in-context learning, instruction following, mathematical reasoning, code generation, and bidirectional reversal reasoning across a suite of standard and custom benchmarks. Use when the user wants to benchmark on MMLU, GSM8K, HumanEval, Chinese Poem Sentence Pairs, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of LLMs to accurately classify academic papers as discussing LLM limitations and to extract supporting evidence from abstracts. It measures alignment with human expert annotations using ordinal rating agreement and span-level extraction metrics. Use when the user wants to benchmark on ACL Anthology & arXiv (crawled 2022-2025), or asks about evaluating this task. Reports weighted-cohens-kappa.
This evaluation probes the multimodal reasoning, visual question answering, OCR, and chart understanding capabilities of large multimodal models. It tests the model's ability to process high-resolution images, extract fine-grained text, and perform complex reasoning across diverse visual domains. Use when the user wants to benchmark on MMStar, MMEBench, MME-RealWorld, SeedBench, CV-Bench, RealWorldQA, MathVista, WeMath, MathVision, MMMU, MMMU-Pro, ChartQA, CharXiv, DocVQA, OCRBench, AI2D, Inf...
Evaluates multimodal instruction-following and complex geological reasoning on lunar surface imagery. It probes the model's ability to interpret crater morphology, degradation states, and inferred geological processes beyond simple visual description. Use when the user wants to benchmark on LUCID (held-out eval set), or asks about evaluating this task. Reports average_overall_score.
Evaluates the effectiveness of structured chain-of-thought prompting and test-time scaling algorithms on multimodal reasoning tasks. It probes whether enforcing a specific reasoning order (summary, caption, reasoning, conclusion) and selecting among multiple generated candidates improves answer accuracy over baseline prompting or dense supervision. Use when the user wants to benchmark on Unspecified multimodal reasoning benchmarks, or asks about evaluating this task. Reports accuracy.
Assesses multimodal chatbot capabilities, including conversation, detailed description, and complex visual reasoning. It measures how well a model follows instructions and understands novel or challenging visual inputs compared to a strong text-only baseline. Use when the user wants to benchmark on LLaVA-Bench, or asks about evaluating this task. Reports relative_score.
Evaluates a LLaMA-based generative speech enhancement model's ability to perform multiple audio restoration tasks (noise suppression, packet loss concealment, target speaker extraction, acoustic echo cancellation, and speech separation) in a task-agnostic manner. It probes the model's capacity to preserve acoustic fidelity and semantic content while generalizing across different acoustic conditions and device types. Use when the user wants to benchmark on DNS Challenge blind test set (Intersp...
Evaluates the language understanding, reasoning, coding, multilingual, and multimodal capabilities of Llama 4 models across a standardized suite of academic and industry benchmarks. It measures performance on text-only, code, and vision-language tasks using few-shot or zero-shot prompting protocols. Use when the user wants to benchmark on MMLU, MMLU-Pro, MATH, MBPP, LiveCodeBench, GPQA Diamond, ChartQA, DocVQA, MMMU, MTOB, or asks about evaluating this task. Reports macro_avg/acc.