Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,785–8,808 of 21,231 skills
Evaluates an object detection model's ability to locate and classify form field widgets (text inputs, checkboxes/radio buttons, and signatures) on scanned or digital form pages. It probes sensitivity to input resolution and robustness across different languages and document domains. Use when the user wants to benchmark on CommonForms, or asks about evaluating this task. Reports mAP50-95.
Evaluates the image quality and text-image alignment of a text-to-image diffusion model trained on Creative-Commons licensed data, benchmarking it against Stable Diffusion 2 using both automated distribution metrics and human pairwise preference. Use when the user wants to benchmark on MS COCO, PartiPrompts, or asks about evaluating this task. Reports User preference rate.
Evaluates multilingual automatic speech recognition (ASR) capabilities, specifically testing speaker generalization and low-resource language adaptation via transfer learning from an English model. It measures how well a model can transcribe audio from diverse, crowdsourced speakers across multiple languages with varying data sizes. Use when the user wants to benchmark on Common Voice, or asks about evaluating this task. Reports character error rate.
Evaluates the extent to which authors fulfill promises made during peer review rebuttals in their final camera-ready papers, and classifies unfulfilled commitments by severity and difficulty. Use when the user wants to benchmark on ICLR 2025, EMNLP 2024, or asks about evaluating this task. Reports fulfillment rate.
Evaluates how well models generate or complete commit messages given code diffs and optional historical context. It probes the model's ability to follow coding conventions, match ground truth exactly, and maintain semantic similarity under varying context lengths. Use when the user wants to benchmark on CMG_test, or asks about evaluating this task. Reports ExactMatch@1.
Evaluates the accuracy and overhead of an integrated thermal simulation toolchain (CoMeT) for modeling processor-memory thermal dynamics across 2D, 2.5D, and 3D architectures. It probes the tool's ability to capture thermal coupling, leakage power effects, and DVFS/DTM interactions under diverse compute and memory-intensive workloads. Use when the user wants to benchmark on PARSEC 2.1, SPLASH-2, SPEC CPU2017, or asks about evaluating this task. Reports Temperature.
Probes the semantic fidelity and grammatical correctness of automated translation pipelines when converting English benchmarks into low-resource languages. It measures how well translation methods preserve task structure and downstream model performance consistency. Use when the user wants to benchmark on FLORES, WMT24++, MMLU, or asks about evaluating this task. Reports COMET.
Evaluates multimodal discrete mathematical reasoning, specifically the ability to parse and solve combinatorial problems involving graphs, grids, and geometric diagrams. It also probes susceptibility to deliberately crafted distractors in multiple-choice formats versus genuine solution construction. Use when the user wants to benchmark on CombiGraph-Vis, or asks about evaluating this task. Reports avg@8.
This benchmark evaluates large language models on formal combinatorial mathematics reasoning within the Lean 4 proof assistant. It probes the model's ability to generate correct, compilable proof scripts and accurately solve fill-in-the-blank combinatorial problems under rigorous automated verification. Use when the user wants to benchmark on CombiBench, or asks about evaluating this task. Reports pass@N.
Evaluates large language models on mathematical reasoning across diverse difficulty levels and languages. It probes the model's ability to convert natural language word problems into structured symbolic representations and execute step-by-step logical derivations without external solvers. Use when the user wants to benchmark on AQUA, MultiArith, GSM8K, MMLU-Redux, Olympiad Bench (English), GaoKao, Olympiad Bench (Chinese), or asks about evaluating this task. Reports exact match.
Evaluates how well a classifier maintains performance on a biased dataset when trained on different coreset selection strategies. It probes the model's robustness to dataset bias and measures the effectiveness of data frugality techniques in mitigating bias while preserving accuracy across varying data budgets. Use when the user wants to benchmark on Colour-MNIST, or asks about evaluating this task. Reports classifier performance.
Evaluates reinforcement learning agents on tabular Markov Decision Processes (MDPs) to measure their performance under varying theoretical hardness criteria, specifically state-action coverage (diameter) and reward structure (environmental value norm). Use when the user wants to benchmark on Colosseum, or asks about evaluating this task. Reports per-step normalized cumulative regret.
Evaluates object detection and localization capabilities for indoor cows under varying camera viewpoints (top, side, external) and lighting conditions (day, night). It probes model generalization across domain shifts in perspective and illumination, testing whether pre-trained weights and model complexity transfer effectively to agricultural environments. Use when the user wants to benchmark on COLO, or asks about evaluating this task. Reports mAP@0.5:0.95.
This protocol evaluates how fine-tuning language models on publicly derived constitutional principles impacts their core reasoning capabilities, social bias propensity, political representativeness, and perceived helpfulness versus harmlessness. It probes whether aligning models with democratic deliberation outputs reduces bias without degrading performance or increasing refusal rates. Use when the user wants to benchmark on MMLU, GSM8K, BBQ, OpinionQA, or asks about evaluating this task. Rep...
Evaluates large language models' ability to perform legal textual entailment and question answering in monolingual and cross-lingual settings. It probes how well models handle linguistic and structural disparities between English and Japanese legal contexts and questions. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of large language models to perform legal textual entailment, specifically measuring how model accuracy changes over time based on the year of the Japanese statute law data used. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.
Evaluates French language understanding across 23 diverse tasks, including sentiment analysis, paraphrase detection, grammatical judgment, reasoning, and extractive QA. It specifically probes capabilities like morphological richness, grammatical gender, syntactic nuance, and regional language variation in a zero-shot setting. Use when the user wants to benchmark on COLE, or asks about evaluating this task. Reports task-specific metrics.
Evaluates cold-start active learning sample selection strategies for 3D medical image segmentation by comparing how well diversity-based, uncertainty-based, and random methods perform when annotation budgets are extremely limited. Use when the user wants to benchmark on Medical Segmentation Decathlon (MSD), or asks about evaluating this task. Reports Dice score.
This benchmark probes the safety and bias of Chinese generative language models by measuring how frequently they produce offensive content when prompted with various inputs, including offensive, non-offensive, and anti-bias contexts. Use when the user wants to benchmark on COLDataset, or asks about evaluating this task. Reports offensive rate.
Evaluates a model's ability to predict drug-target binding interactions under cold-start conditions where either drugs, proteins, or both are completely unseen during training. It probes the model's capacity to generalize across different protein structural granularities (primary to quaternary) and handle severe class imbalance inherent in biological interaction datasets. Use when the user wants to benchmark on DrugBank, BindingDB, BioSNAP, Human, or asks about evaluating this task. Reports AUC.
This benchmark evaluates a model's ability to classify English sentences as grammatically acceptable or unacceptable. It probes syntactic competence by measuring performance on both in-domain and out-of-domain linguistic data. Use when the user wants to benchmark on CoLA, or asks about evaluating this task. Reports MCC.
Evaluates code information retrieval models across diverse tasks including text-to-code, code-to-code, code-to-text, and hybrid code retrieval. It probes a model's ability to handle semi-structured, syntactically complex code snippets and natural language queries across multiple programming languages and domains. Use when the user wants to benchmark on APPS, CosQA, Synthetic Text2SQL, CodeSearchNet, CodeSearchNet-CCR, CodeTransOcean-DL, CodeTransOcean-Contest, StackOverflow QA, CodeFeedQA, Co...
Evaluates the temporal synchronization accuracy and robustness of a multi-node cooperative perception system under adverse weather and network conditions. It measures how well a delay-aware protocol aligns sensor data across nodes compared to naive asynchronous methods, focusing on timing errors, fusion completeness, and reaction latency. Use when the user wants to benchmark on CoInfra, or asks about evaluating this task. Reports full_match_rate.
Evaluates a deep learning model's ability to classify breast masses as benign or malignant in mammography images. It probes the effectiveness of adversarial data augmentation and contrastive manifold learning in improving discriminative feature extraction under data scarcity. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports Accuracy.