Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,025–9,048 of 21,244 skills
Evaluates multimodal models' ability to answer questions about charts by extracting visual and textual information. It probes robustness to missing labels and geometric perturbations, distinguishing between simple pattern matching and true structural reasoning. Use when the user wants to benchmark on ChartQA, Charixv, or asks about evaluating this task. Reports Relaxed Accuracy (RA).
This benchmark evaluates large language models' ability to comprehend factual data in charts (FACT task) and their capacity to predict how visual manipulations mislead human readers (MIND task). It probes visual theory-of-mind by measuring whether models can distinguish between objective chart data and subjective human perceptual biases introduced by deceptive visualization techniques. Use when the user wants to benchmark on CHARTOM, or asks about evaluating this task. Reports FACT_accuracy.
Evaluates multimodal vision-language models on chart understanding tasks, including reconstructing plotting code from charts, extracting tabular data, summarizing chart content, and answering complex reasoning questions. Use when the user wants to benchmark on ChartNet Evaluation Set, or asks about evaluating this task. Reports ChartNet Evaluation Metrics.
This benchmark evaluates the visual reasoning capabilities of Large Vision-Language Models (LVLMs) on real-world charts. It specifically probes the model's ability to perform visual extraction, object picking, visual comparisons, and trajectory tracking, while distinguishing these from purely textual inference tasks. Use when the user wants to benchmark on CHARTMUSEUM, or asks about evaluating this task. Reports accuracy.
Evaluates large multimodal models' cross-modal reasoning by generating code to reproduce or modify charts based on visual and textual instructions. It tests visual understanding, code generation, and the integration of textual and visual inputs. Use when the user wants to benchmark on ChartMimic, or asks about evaluating this task. Reports Overall.
This benchmark evaluates multimodal large language models' ability to understand and reason about infographic charts. It probes capabilities in text-based data reasoning, visual-element association, and visual style analysis through structured question-answering tasks. Use when the user wants to benchmark on ChartGalaxy, or asks about evaluating this task. Reports relaxed accuracy (5% margin).
Evaluates multimodal image-to-image editing models on chart editing tasks, measuring both low-level visual fidelity and high-level semantic correctness and consistency against editing instructions. Use when the user wants to benchmark on ChartE³, or asks about evaluating this task. Reports Correctness.
This benchmark evaluates a model's ability to perform cross-chart comparative reasoning by generating natural language summaries that identify and explain differences in trends, fluctuations, and anomalies between pairs of charts. It probes vision-language models on their capacity to synthesize visual information from multiple plots into coherent, human-aligned textual descriptions. Use when the user wants to benchmark on ChartDiff, or asks about evaluating this task. Reports GPT Score.
Evaluates the capability of vision-language models to generate dense, structurally accurate captions for charts while mitigating hallucinations. It probes both reference-based text similarity and a novel visual consistency metric that verifies if a generated caption can successfully reconstruct the original chart image. Use when the user wants to benchmark on ChartCap, or asks about evaluating this task. Reports Visual Consistency Score.
Evaluates a multimodal language model's ability to comprehend, summarize, and answer questions about various chart types (base and specialized). It probes chart-to-text generation, open-ended and numerical question answering, referring question answering, and chart-to-table translation. Use when the user wants to benchmark on ChartQA, Chart-to-Text, OpenCQA, MathQA, ReferQA, RealQA, or asks about evaluating this task. Reports relaxed_correctness.
This benchmark evaluates vision-language models on fine-grained chart understanding, specifically focusing on dense grounding of data and visual attributes, identifying precise differences between paired charts, and measuring robustness to stylistic perturbations like color or font changes. Use when the user wants to benchmark on ChartAB, or asks about evaluating this task. Reports SCRM.
Evaluates multimodal language models' ability to comprehend diverse chart types, extract underlying numerical data, and answer questions across varying complexity levels. It distinguishes between OCR-dependent recognition on annotated charts and true data reasoning on unannotated or raw-data-requiring charts. Use when the user wants to benchmark on ChartQA, PlotQA, ChartDQA, MMC, ChartX, Chart-to-Table, Chart-to-Text, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to translate chart images into executable plotting code, measuring both code executability and visual/textual fidelity of the generated charts. Use when the user wants to benchmark on ChartMimic, Plot2Code, ChartX, or asks about evaluating this task. Reports High-Level Score.
Evaluates a model's ability to understand complex chart visualizations and perform fine-grained visual grounding and numerical reasoning. It probes both in-domain chart comprehension across real-world and synthetic datasets, and out-of-domain generalization to visual mathematical reasoning tasks. Use when the user wants to benchmark on CharXiv, ChartQAPro, ChartQA, ChartBench, ChartX, ReachQA, MathVista, WeMath, MathVerse, or asks about evaluating this task. Reports accuracy.
Evaluates vision-language models on chart question answering, probing their ability to extract numerical data, perform visual interpolation, and reason over diverse chart types. It measures both direct data retrieval and complex reasoning capabilities across real-world and synthetic chart distributions. Use when the user wants to benchmark on FigureQA-Sub, DVQA-Sub, PlotQA-Sub, ChartQA, CharXiv, or asks about evaluating this task. Reports exact accuracy, relaxed accuracy.
Evaluates multimodal large language models' ability to perform counterfactual reasoning over chart visualizations. It probes whether models rely on parametric memory or truly understand the visual data when answering questions that contain hypothetical assumptions about the chart. Use when the user wants to benchmark on Chart-HQA, or asks about evaluating this task. Reports relaxed accuracy.
Evaluates the accuracy of neural networks in regressing charged particle hit positions (x, y) and incident angles (alpha, beta) from pixelated silicon sensor charge patterns. It measures how well on-chip inference models reconstruct particle trajectories compared to ground-truth simulation and traditional reconstruction algorithms. Use when the user wants to benchmark on Simulated silicon tracker charge clusters, or asks about evaluating this task. Reports residual_68pct_interval.
This evaluation probes an LLM-based autorater's ability to predict fine-grained machine translation errors (spans, severities, categories) without using human references. It specifically tests how well the model can specialize to a given test set by leveraging in-context examples of human ratings from other systems on the same inputs. Use when the user has predictions and gold and needs to compute character-level F1.
Evaluates the visual fidelity and spatial accuracy of formula recognition models by comparing rendered images of predicted and ground-truth LaTeX code at the character level. It addresses the misalignment of text-based metrics with human perception by treating each character as a detectable object in an image. Use when the user has predictions and gold and needs to compute CDM.
Evaluates the ability of time series forecasting models to predict future values of noisy, chaotic dynamical systems. It probes how well models capture underlying nonlinear dynamics and handle varying levels of observation noise and system complexity. Use when the user wants to benchmark on Gilpin chaotic systems benchmark, or asks about evaluating this task. Reports SMAPE.
Evaluates a model's ability to detect semantic building changes between two temporally separated remote sensing images. It probes robustness to weak temporal supervision, label noise, and out-of-domain generalization in large-scale urban environments. Use when the user wants to benchmark on b-FLAIR-test, b-FLAIR-test-spot, LEVIR-CD, WHUCD, S2Looking, or asks about evaluating this task. Reports F1-score (F1), Intersection over Union (IoU).
Evaluates a model's ability to recognize and verify traditional Telugu Chandassu metrical patterns in padyam poetry. It measures adherence to structural prosodic constraints including syllable counts, line divisions, sequential gana patterns, rhythmic breaks, and recurring syllable markers. Use when the user wants to benchmark on Telugu Chandassu Padyam Dataset, or asks about evaluating this task. Reports Chandassu Score.
Evaluates the transfer learning capability and generalization of deep learning models (CNNs and ViTs) on patch-level histopathology image classification tasks across multiple cancer-related benchmarks. Use when the user wants to benchmark on Various publicly available histopathology datasets, or asks about evaluating this task. Reports AUROC.
Evaluates the accuracy and inference efficiency of training-free multimodal reasoning methods across diverse vision-language benchmarks. It probes how well atomic visual hint injection reduces redundant reasoning steps while maintaining or improving task performance on math, logic, science, and general visual understanding tasks. Use when the user wants to benchmark on MathVista mini, MathVision, WeMath, MMMU Pro vis, LogicVista, OlympiadBench, VStar, CVBench, ConBench, ChartVQA, SEED-Bench, ...