Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,244
skills in category
886
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 9,025–9,048 of 21,244 skills

Chartqa EvalA

Evaluates multimodal models' ability to answer questions about charts by extracting visual and textual information. It probes robustness to missing labels and geometric perturbations, distinguishing between simple pattern matching and true structural reasoning. Use when the user wants to benchmark on ChartQA, Charixv, or asks about evaluating this task. Reports Relaxed Accuracy (RA).

researchpythongo
0
3
Chartom EvalA

This benchmark evaluates large language models' ability to comprehend factual data in charts (FACT task) and their capacity to predict how visual manipulations mislead human readers (MIND task). It probes visual theory-of-mind by measuring whether models can distinguish between objective chart data and subjective human perceptual biases introduced by deceptive visualization techniques. Use when the user wants to benchmark on CHARTOM, or asks about evaluating this task. Reports FACT_accuracy.

researchpythongo
0
3
Chartnet EvalA

Evaluates multimodal vision-language models on chart understanding tasks, including reconstructing plotting code from charts, extracting tabular data, summarizing chart content, and answering complex reasoning questions. Use when the user wants to benchmark on ChartNet Evaluation Set, or asks about evaluating this task. Reports ChartNet Evaluation Metrics.

researchpythongo
0
3
Chartmuseum EvalA

This benchmark evaluates the visual reasoning capabilities of Large Vision-Language Models (LVLMs) on real-world charts. It specifically probes the model's ability to perform visual extraction, object picking, visual comparisons, and trajectory tracking, while distinguishing these from purely textual inference tasks. Use when the user wants to benchmark on CHARTMUSEUM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chartmimic EvalA

Evaluates large multimodal models' cross-modal reasoning by generating code to reproduce or modify charts based on visual and textual instructions. It tests visual understanding, code generation, and the integration of textual and visual inputs. Use when the user wants to benchmark on ChartMimic, or asks about evaluating this task. Reports Overall.

researchpythongo
0
3
Chartgalaxy EvalA

This benchmark evaluates multimodal large language models' ability to understand and reason about infographic charts. It probes capabilities in text-based data reasoning, visual-element association, and visual style analysis through structured question-answering tasks. Use when the user wants to benchmark on ChartGalaxy, or asks about evaluating this task. Reports relaxed accuracy (5% margin).

researchpythongo
0
3
Charte3 EvalA

Evaluates multimodal image-to-image editing models on chart editing tasks, measuring both low-level visual fidelity and high-level semantic correctness and consistency against editing instructions. Use when the user wants to benchmark on ChartE³, or asks about evaluating this task. Reports Correctness.

researchpythongo
0
3
Chartdiff EvalA

This benchmark evaluates a model's ability to perform cross-chart comparative reasoning by generating natural language summaries that identify and explain differences in trends, fluctuations, and anomalies between pairs of charts. It probes vision-language models on their capacity to synthesize visual information from multiple plots into coherent, human-aligned textual descriptions. Use when the user wants to benchmark on ChartDiff, or asks about evaluating this task. Reports GPT Score.

researchpython
0
3
Chartcap EvalA

Evaluates the capability of vision-language models to generate dense, structurally accurate captions for charts while mitigating hallucinations. It probes both reference-based text similarity and a novel visual consistency metric that verifies if a generated caption can successfully reconstruct the original chart image. Use when the user wants to benchmark on ChartCap, or asks about evaluating this task. Reports Visual Consistency Score.

researchpythongo
0
3
Chartassistant EvalA

Evaluates a multimodal language model's ability to comprehend, summarize, and answer questions about various chart types (base and specialized). It probes chart-to-text generation, open-ended and numerical question answering, referring question answering, and chart-to-table translation. Use when the user wants to benchmark on ChartQA, Chart-to-Text, OpenCQA, MathQA, ReferQA, RealQA, or asks about evaluating this task. Reports relaxed_correctness.

researchpythongo
0
3
Chartab EvalA

This benchmark evaluates vision-language models on fine-grained chart understanding, specifically focusing on dense grounding of data and visual attributes, identifying precise differences between paired charts, and measuring robustness to stylistic perturbations like color or font changes. Use when the user wants to benchmark on ChartAB, or asks about evaluating this task. Reports SCRM.

researchpythongo
0
3
Chart Understanding EvalA

Evaluates multimodal language models' ability to comprehend diverse chart types, extract underlying numerical data, and answer questions across varying complexity levels. It distinguishes between OCR-dependent recognition on annotated charts and true data reasoning on unannotated or raw-data-requiring charts. Use when the user wants to benchmark on ChartQA, PlotQA, ChartDQA, MMC, ChartX, Chart-to-Table, Chart-to-Text, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chart To Code EvalA

Evaluates a model's ability to translate chart images into executable plotting code, measuring both code executability and visual/textual fidelity of the generated charts. Use when the user wants to benchmark on ChartMimic, Plot2Code, ChartX, or asks about evaluating this task. Reports High-Level Score.

researchpythongo
0
3
Chart Reasoning EvalA

Evaluates a model's ability to understand complex chart visualizations and perform fine-grained visual grounding and numerical reasoning. It probes both in-domain chart comprehension across real-world and synthetic datasets, and out-of-domain generalization to visual mathematical reasoning tasks. Use when the user wants to benchmark on CharXiv, ChartQAPro, ChartQA, ChartBench, ChartX, ReachQA, MathVista, WeMath, MathVerse, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chart Qa EvalA

Evaluates vision-language models on chart question answering, probing their ability to extract numerical data, perform visual interpolation, and reason over diverse chart types. It measures both direct data retrieval and complex reasoning capabilities across real-world and synthetic chart distributions. Use when the user wants to benchmark on FigureQA-Sub, DVQA-Sub, PlotQA-Sub, ChartQA, CharXiv, or asks about evaluating this task. Reports exact accuracy, relaxed accuracy.

researchpythongo
0
3
Chart Hqa EvalA

Evaluates multimodal large language models' ability to perform counterfactual reasoning over chart visualizations. It probes whether models rely on parametric memory or truly understand the visual data when answering questions that contain hypothetical assumptions about the chart. Use when the user wants to benchmark on Chart-HQA, or asks about evaluating this task. Reports relaxed accuracy.

researchpythongo
0
3
Charged Particle Tracking ResidualsA

Evaluates the accuracy of neural networks in regressing charged particle hit positions (x, y) and incident angles (alpha, beta) from pixelated silicon sensor charge patterns. It measures how well on-chip inference models reconstruct particle trajectories compared to ground-truth simulation and traditional reconstruction algorithms. Use when the user wants to benchmark on Simulated silicon tracker charge clusters, or asks about evaluating this task. Reports residual_68pct_interval.

researchpythongo
0
3
Character Level F1A

This evaluation probes an LLM-based autorater's ability to predict fine-grained machine translation errors (spans, severities, categories) without using human references. It specifically tests how well the model can specialize to a given test set by leveraging in-context examples of human ratings from other systems on the same inputs. Use when the user has predictions and gold and needs to compute character-level F1.

researchpythongo
0
3
Character Detection MatchingA

Evaluates the visual fidelity and spatial accuracy of formula recognition models by comparing rendered images of predicted and ground-truth LaTeX code at the character level. It addresses the misalignment of text-based metrics with human perception by treating each character as a detectable object in an image. Use when the user has predictions and gold and needs to compute CDM.

researchpythongo
0
3
Chaotic Time Series Forecasting EvalA

Evaluates the ability of time series forecasting models to predict future values of noisy, chaotic dynamical systems. It probes how well models capture underlying nonlinear dynamics and handle varying levels of observation noise and system complexity. Use when the user wants to benchmark on Gilpin chaotic systems benchmark, or asks about evaluating this task. Reports SMAPE.

researchpythontesting
0
3
Change Detection EvalA

Evaluates a model's ability to detect semantic building changes between two temporally separated remote sensing images. It probes robustness to weak temporal supervision, label noise, and out-of-domain generalization in large-scale urban environments. Use when the user wants to benchmark on b-FLAIR-test, b-FLAIR-test-spot, LEVIR-CD, WHUCD, S2Looking, or asks about evaluating this task. Reports F1-score (F1), Intersection over Union (IoU).

researchpython
0
3
Chandassu Metrical EvalA

Evaluates a model's ability to recognize and verify traditional Telugu Chandassu metrical patterns in padyam poetry. It measures adherence to structural prosodic constraints including syllable counts, line divisions, sequential gana patterns, rhythmic breaks, and recurring syllable markers. Use when the user wants to benchmark on Telugu Chandassu Padyam Dataset, or asks about evaluating this task. Reports Chandassu Score.

researchpythongo
0
3
Champkit EvalA

Evaluates the transfer learning capability and generalization of deep learning models (CNNs and ViTs) on patch-level histopathology image classification tasks across multiple cancer-related benchmarks. Use when the user wants to benchmark on Various publicly available histopathology datasets, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Chainv EvalA

Evaluates the accuracy and inference efficiency of training-free multimodal reasoning methods across diverse vision-language benchmarks. It probes how well atomic visual hint injection reduces redundant reasoning steps while maintaining or improving task performance on math, logic, science, and general visual understanding tasks. Use when the user wants to benchmark on MathVista mini, MathVision, WeMath, MMMU Pro vis, LogicVista, OlympiadBench, VStar, CVBench, ConBench, ChartVQA, SEED-Bench, ...

researchpythongo
0
3