Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,244
skills in category
886
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 9,001–9,024 of 21,244 skills

Chestxray14 Multi Label EvalA

Evaluates deep learning models for multi-label chest X-ray pathology classification. It probes the capability of CNN architectures to detect 14 specific thoracic diseases, testing the impact of transfer learning, network depth, and fusion of non-image patient demographics (age, gender, view position) on diagnostic accuracy. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AUC.

researchpythontesting
0
3
Chestxray14 Bias EvalA

Evaluates the effectiveness of attribute-neutralization models in removing demographic bias (sex and age) from chest X-ray images while preserving diagnostic utility for 15 disease findings. It probes the trade-off between demographic leakage suppression and clinical performance across varying edit intensities. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AI-Judge AUC.

researchpythonperformance
0
3
Chestxray Pneumonia EvalA

Evaluates deep learning models for binary classification of chest X-rays into normal versus pneumonia categories, while also assessing the spatial interpretability of model predictions using Grad-CAM heatmaps. Use when the user wants to benchmark on Chest X-Rays dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Chestx Det EvalA

Evaluates instance-level detection and segmentation of thoracic diseases on chest X-rays. It probes a model's ability to localize and classify 13 disease categories while handling domain-specific challenges like ambiguous boundaries and disease co-occurrence. Use when the user wants to benchmark on ChestX-Det, DR-private, or asks about evaluating this task. Reports AP_50^bb.

researchpythongo
0
3
Chestagentbench EvalA

Evaluates an AI agent's ability to perform multi-step medical reasoning and tool orchestration for chest X-ray interpretation. It probes capabilities across seven clinically relevant categories: detection, classification, localization, comparison, relationship, diagnosis, and characterization. Use when the user wants to benchmark on ChestAgentBench, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Chest Xray14 EvalA

Evaluates a model's ability to perform multi-label classification of 14 thoracic diseases on chest X-ray images and localize pathological regions using attention maps. It probes whether anatomically grounded feature weighting improves disease detection over global feature fusion or saliency-based methods. Use when the user wants to benchmark on Chest X-ray14, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Chest Xray Report Generation EvalA

This evaluation probes a model's ability to generate clinically coherent and accurate natural language reports from chest X-ray images. It measures both surface-level linguistic similarity to ground-truth radiology reports and the clinical correctness of extracted pathological findings. Use when the user wants to benchmark on Indiana U. Chest X-Ray, MIMIC-CXR, or asks about evaluating this task. Reports BLEU-1.

researchpythongo
0
3
Chest Xray Classification EvalA

Evaluates a CNN's ability to classify chest X-ray images into disease categories (COVID-19, pneumonia, tuberculosis, normal) using various preprocessing techniques. It probes robustness across different dataset sizes and class distributions. Use when the user wants to benchmark on Multiclass Chest X-ray Dataset, Hamad Medical Corporation Tuberculosis Dataset, Pneumonia Dataset, NIH Chest X-ray Dataset, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Chess Puzzle EvalA

Evaluates a model's ability to play chess and solve tactical puzzles without explicit search, testing its capacity for long-horizon planning and generalization to novel board states. Use when the user wants to benchmark on Lichess puzzles, or asks about evaluating this task. Reports Lichess Elo.

researchpythongo
0
3
Chemo EvalA

Evaluates multimodal large language models' ability to solve Olympiad-level theoretical chemistry problems requiring visual perception, chemical reasoning, and structured problem-solving. It specifically probes the visual perception bottleneck in chemistry tasks and tests the effectiveness of multi-agent orchestration and structured visual enhancement. Use when the user wants to benchmark on ChemO, or asks about evaluating this task. Reports normalized rubric-based score.

researchpythonperformance
0
3
Chemllmbench EvalA

Evaluates LLM-based chemistry agents on four core computational chemistry tasks: text-to-molecule generation, molecule-to-text captioning, molecular property prediction, and reaction product prediction. It probes the model's ability to translate between chemical language and structures, predict physicochemical properties, and correct tool errors via hierarchical agent stacking. Use when the user wants to benchmark on ChemLLMBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Chemcotb EvalA

Evaluates large language models' ability to perform stepwise chemical reasoning and molecular structure manipulation. It probes capabilities in molecular understanding, functional group/ring recognition, scaffold extraction, and SMILES-based molecule editing and reaction prediction. Use when the user wants to benchmark on ChemCoTBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chebi 20 Mm EvalA

Evaluates large language models' ability to process, generate, and retrieve molecular information across multiple modalities (SMILES, InChI, SELFIES, graphs, captions, IUPAC names, images). It probes cross-modal compatibility, chemical knowledge acquisition, and molecular property prediction capabilities. Use when the user wants to benchmark on ChEBI-20-MM, or asks about evaluating this task. Reports ROC_AUC.

researchpythongo
0
3
Chd CmmsA

Evaluates the perceptual quality and human preference alignment of generative image models by comparing their discrete visual token distributions and sequence statistics against human ratings. Use when the user has predictions and gold and needs to compute CHD, CMMS.

researchpythongo
0
3
Chc Comp21 Lia Lin EvalA

Evaluates the effectiveness of bottom-up versus top-down transformations for solving linear constrained Horn clauses (CHCs) using software verification workflows. It measures how many verification tasks a solver can successfully resolve within a strict time limit. Use when the user wants to benchmark on CHC-COMP21 LIA-Lin track, or asks about evaluating this task. Reports solved_tasks.

researchpython
0
3
Chattime EvalA

Evaluates a multimodal time series foundation model on zero-shot forecasting, context-guided forecasting, and time series question answering. It probes the model's ability to handle discretized numerical time series alongside textual prompts for prediction and feature recognition without task-specific fine-tuning. Use when the user wants to benchmark on Electric, Exchange, Traffic, Weather, and ETT datasets, Context-guided multimodal datasets, Synthesized TSQA dataset, or asks about evaluatin...

researchpythongo
0
3
Chatterbox Mrg EvalA

Evaluates a model's ability to perform multi-round multimodal referring and grounding, requiring logical consistency across dialogue turns while generating accurate text responses and bounding box coordinates for visual instances. Use when the user wants to benchmark on CB-LC, RefCOCOg, COCO 2017, or asks about evaluating this task. Reports BERT(·).

researchpythongo
0
3
Chatqa Conversational Qa EvalA

Evaluates multi-turn conversational question answering over both long and short documents, including text-only and tabular contexts. It probes a model's ability to retrieve relevant context, handle topic switching, perform arithmetic reasoning, and correctly identify unanswerable queries. Use when the user wants to benchmark on Doc2Dial, QuAC, QReCC, TopiOCQA, INSCIT, CoQA, DoQA, ConvFinQA, SQA, HybridDial, or asks about evaluating this task. Reports Average F1/EM across 10 datasets.

researchpythongo
0
3
Chatgpt Nlp EvalA

Evaluates ChatGPT's few-shot generation and reasoning capabilities across multiple NLP tasks, including question answering, commonsense reasoning, natural language inference, and sentiment analysis. It probes the model's ability to follow task-specific formalizations, leverage retrieved demonstrations, and mitigate hallucination through self-verification. Use when the user wants to benchmark on SQuADv2, TQA, MRQA-OOD, CSQA, StrategyQA, RTE, CommitmentBank, SST-2, IMDB, Yelp, or asks about eva...

researchpythongo
0
3
Chatgpt Multitask EvalA

Evaluates ChatGPT's zero-shot multitask capabilities across summarization, machine translation, sentiment analysis, question answering, dialogue, and misinformation detection. It probes the model's generalization, reasoning, multilingual understanding, and task-specific performance without fine-tuning. Use when the user wants to benchmark on CNN/DM, SAMSum, FLoRes-200, NusaX, bAbI, EntailmentBank, CLUTRR, StepGame, Pep-3k, COVID-Social, COVID-Scientific, MultiWOZ2.2, OpenDialKG, or asks about...

researchpythongo
0
3
Chatbot Arena EvalA

Evaluates large language models by collecting human preference votes on pairwise responses to real-world prompts, then ranks them using Bradley-Terry models to measure alignment and real-world utility. Use when the user wants to benchmark on Chatbot Arena, or asks about evaluating this task. Reports BT coefficients.

researchpythongo
0
3
Charxiv EvalA

Evaluates multimodal large language models' ability to understand real-world charts through descriptive and reasoning tasks. It probes capabilities like information extraction, pattern recognition, counting, compositional reasoning, and robustness to chart complexity (e.g., multiple subplots) and unanswerable queries. Use when the user wants to benchmark on CharXiv, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chartverse EvalA

Evaluates visual language models on complex chart understanding and reasoning tasks, including data extraction, comparison, and trend analysis across diverse chart types and domains. Use when the user wants to benchmark on ChartQA-Pro, CharXiv, ChartMuseum, ChartX, ChartBench, EvoChart, or asks about evaluating this task. Reports average score.

researchpythongo
0
3
Chartsumm EvalA

Evaluates automatic chart-to-text summarization models on their ability to generate accurate, fluent, and informative summaries from chart metadata and data tables. It probes factual correctness, trend capture, and hallucination resistance across short and long summary formats. Use when the user wants to benchmark on ChartSumm, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3