Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,761–8,784 of 21,231 skills
Evaluates the ability of neural models to predict human-assigned translation quality scores for Indian language pairs, measuring alignment with crowd-sourced DA+SQM ratings. Use when the user has predictions and gold and needs to compute Pearson correlation, Spearman correlation.
Evaluates large vision-language models on chain-of-thought reasoning that requires generating both textual explanations and intermediate or final images. It probes the model's ability to perform four specific visual operations (creation, deletion, update, and selection) and align its multi-modal reasoning steps with ideal visual states. Use when the user wants to benchmark on CoMT, or asks about evaluating this task. Reports F1 score.
Evaluates the robustness of open-set anomaly segmentation models under complex, real-world driving conditions. It probes a model's ability to detect out-of-distribution objects across diverse landforms and adverse weather while correctly ignoring non-driving-area elements and void regions. Use when the user wants to benchmark on ComsAmy, or asks about evaluating this task. Reports AuPRC.
Evaluates computer-use agents on desktop task completion in online and offline settings, and measures the precision of a video-to-action module in detecting GUI events and extracting interaction parameters from screen recordings. Use when the user wants to benchmark on OSWorld-Verified, AgentNetBench, Video2Action Held-out Test Set, or asks about evaluating this task. Reports task success rate, step success rate.
Evaluates the ability of transformer-based architectures to model long-range dependencies efficiently by compressing past hidden states into a fixed-size memory. It probes sequence modeling capabilities across text, audio, and visual domains, measuring how well compressed representations preserve salient information for next-token prediction and task completion. Use when the user wants to benchmark on Enwiki8, WikiText-103, PG-19, DMLab-30 (rooms_select_nonmatching_object), or asks about eval...
Evaluates the trade-offs between model accuracy and resource efficiency when applying various DNN compression techniques on mobile hardware. It measures how different compression methods affect inference speed, energy consumption, and storage footprint across standard vision and audio datasets. Use when the user wants to benchmark on CIFAR-10, MNIST, CIFAR-100, ImageNet, UbiSound, Har, or asks about evaluating this task. Reports accuracy.
Evaluates the comprehensiveness and fine-grained accuracy of detailed image captions generated by vision-language models. It probes object detection, attribute binding, directional relationship modeling, and perception of tiny objects through hierarchical scene graph alignment and dedicated VQA tasks. Use when the user wants to benchmark on CompreCap, or asks about evaluating this task. Reports S_unified.
This benchmark evaluates hardware-software co-design trade-offs for compound AI applications by measuring end-to-end latency, energy consumption, and accuracy across multi-modal workflows like video QA, evolutionary code generation, and RAG. It probes how different hardware configurations and software optimizations impact system performance under varying latency targets and workload patterns. Use when the user wants to benchmark on Google FRAMES benchmark, or asks about evaluating this task. ...
Evaluates the ability of on-device LLMs to perform two distinct tasks simultaneously in a single forward pass (compositional multi-tasking), such as summarization combined with translation or tone adjustment, while maintaining strict efficiency constraints. Use when the user wants to benchmark on Compositional Multi-tasking Benchmark, or asks about evaluating this task. Reports LLM judge (LLM-J).
This benchmark evaluates systematic generalization in abstract spatial reasoning by testing whether models can infer and compose geometric transformations (e.g., translation, rotation, reflection) from limited few-shot examples. It specifically probes out-of-distribution compositionality by training on known transformation primitives and level-1 compositions, then testing on novel level-2 compositions. Use when the user wants to benchmark on Compositional-ARC, or asks about evaluating this ta...
Evaluates cross-lingual safety degradation in LLMs by measuring how well models refuse harmful prompts and avoid generating unsafe content across English and five Indic languages. Use when the user wants to benchmark on CompositeHarm, or asks about evaluating this task. Reports Refusal Rate (RR), Attack Success Rate (ASR).
Probes a model's ability to perform heterogeneous question answering by integrating information from multiple sources (knowledge bases, text, tables, infoboxes) across diverse domains and complex question intents. It specifically tests whether systems can fuse complementary structured and unstructured data to answer self-contained, human-generated questions. Use when the user wants to benchmark on CompMix, or asks about evaluating this task. Reports answer exact match.
Evaluates an agent's ability to play competitive Pokémon Singles under partial observability and long-horizon uncertainty. It measures strategic decision-making, team building, and adaptation against heuristic opponents, search-based engines, LLM agents, and human players on a ranked ladder. Use when the user wants to benchmark on Competitive Pokémon Singles (CPS) on Pokémon Showdown, or asks about evaluating this task. Reports win rate.
Evaluates the ability of large language models to generate correct, executable Python solutions for competitive programming problems. It probes algorithmic reasoning, code synthesis, and adherence to problem constraints under strict time and complexity limits. Use when the user wants to benchmark on LiveCodeBench, CodeContests, or asks about evaluating this task. Reports pass@1.
Evaluates Multimodal Large Language Models' ability to comprehend composite images (charts, collages, tables, code) and natural images, covering text recognition, visual reasoning, and conversational capabilities. Use when the user wants to benchmark on SEEDBench*, TextVQA, MMBench, MME, LLaVABench, ChartQA, DocVQA, InfoVQA, WebSRC, MathVista, OCRBench, or asks about evaluating this task. Reports Average score.
Evaluates a fairness-aware ensemble learning framework on recidivism risk prediction. It probes the model's ability to balance predictive accuracy against multiple group fairness constraints across racial demographics in a counterfactual causal setting. Use when the user wants to benchmark on COMPAS, or asks about evaluating this task. Reports MSE.
Evaluates Open Information Extraction systems on their ability to extract compact, clause-level facts from text. It measures precision, recall, and F1 using token-level matching against gold triples, with a focus on avoiding over-specific extractions and handling overlapping constituents. Use when the user wants to benchmark on CaRB, Wire57, BenchIE, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to perform comprehensive hierarchical document structure analysis, including detecting page objects, predicting reading order across multiple groups, extracting tables of contents, and reconstructing the overall document hierarchy. Use when the user wants to benchmark on Comp-HRDoc, PubLayNet, DocLayNet, HRDoc, or asks about evaluating this task. Reports segmentation-based mAP.
Evaluates a model's ability to align latent representations across different conditions (e.g., batch effects, treatment, demographic attributes) while preserving task-relevant information. It measures local mixing quality using nearest-neighbour and silhouette metrics, and assesses predictive utility via classification accuracy on held-out labels. Use when the user wants to benchmark on Tumour / Cell Line, Stimulated / untreated single-cell PBMCs, Single-cell RNA-seq data integration (PBMCs),...
This evaluation protocol probes the ability of fake image detectors to generalize across a wide variety of generative models and architectures. It measures how well classifiers trained on diverse synthetic data can distinguish real from generated images in both in-distribution and out-of-distribution settings. Use when the user wants to benchmark on Wang et al. [129], Ojha et al. [90], Synthbuster [7], GenImage [137], Community Forensics (Ours), or asks about evaluating this task. Reports mAP.
Evaluates the stability, uncertainty quantification, and accuracy of consensus-based community detection algorithms against ground-truth partitions on synthetic and real-world benchmark networks. Use when the user wants to benchmark on Zachary's Karate Network, LFR Benchmark, Ring of Cliques (RC) Benchmark, or asks about evaluating this task. Reports NMI.
Evaluates machine translation quality across multiple language pairs and specialized tasks (general translation, terminology-constrained, and automatic post-editing). It measures how well encoder-decoder and decoder-only models generate accurate and fluent target sentences. Use when the user wants to benchmark on ComMT, or asks about evaluating this task. Reports SacreBLEU.
This benchmark evaluates a model's ability to answer multiple-choice questions that require real-world commonsense knowledge. It specifically probes whether models can distinguish a correct answer from semantically plausible but factually incorrect distractors based on spatial, causal, or physical reasoning. Use when the user wants to benchmark on CommonsenseQA, or asks about evaluating this task. Reports accuracy.
This benchmark probes the commonsense reasoning capabilities of vision-language models by evaluating their ability to match images to text riddles (or vice versa) where the subject entity is replaced with a demonstrative pronoun. It specifically tests relational knowledge retrieval and generalization to unseen knowledge triples. Use when the user wants to benchmark on DANCE Diagnostic Set, or asks about evaluating this task. Reports Acc@50.