Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,977–9,000 of 21,233 skills
Evaluates fundamental lexical comprehension and generation capabilities of multilingual LLMs across thousands of languages. It probes word-level translation, context-aware translation, translation-conditioned language modeling, and bag-of-words machine translation tasks to measure basic linguistic competence beyond high-resource languages. Use when the user wants to benchmark on ChiKhaPo, or asks about evaluating this task. Reports language score.
Evaluates the ability of recurrent graph neural networks to forecast spatiotemporal epidemiological time series. It probes how well models capture spatial dependencies between adjacent regions and temporal dynamics like seasonality and zero-inflation over multiple forecasting horizons. Use when the user wants to benchmark on Chickenpox Cases in Hungary, or asks about evaluating this task. Reports mean squared error.
This evaluation probes the transferability of ImageNet-pretrained CNN architectures to chest X-ray interpretation, measuring how model size, architecture family, and pretraining affect classification performance and parameter efficiency on the CheXpert dataset. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports CheXpert AUC.
Evaluates the generalization of chest X-ray deep learning models to two clinically relevant distribution shifts: smartphone photographs of digital X-rays (introducing visual artifacts like blur and glare) and external institutional data. It probes robustness to imaging degradation and cross-institutional heterogeneity in multi-label pathology detection. Use when the user wants to benchmark on CheXpert, NIH, or asks about evaluating this task. Reports MCC.
Evaluates a multi-task deep learning framework for thorax disease classification and weakly-supervised localization on chest X-rays. It probes the model's ability to detect multiple pathologies and accurately localize abnormal regions without requiring pre-annotated bounding boxes during training. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, MIMIC-CXR, or asks about evaluating this task. Reports AUC.
Evaluates an automated rule-based pipeline's ability to extract clinical observations from free-text radiology reports. It specifically tests the system's capacity to classify mentions as positive, negative, or uncertain, and to aggregate them into structured labels for 14 predefined observations. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports F1 score.
Evaluates the diagnostic accuracy of predictive algorithms and human radiologists on chest X-ray images for detecting atelectasis. It specifically probes whether human expertise provides actionable, non-redundant information on input subsets where algorithms are algorithmically indistinguishable. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
Evaluates text-to-image generative models for synthetic chest radiograph generation across three dimensions: fidelity (image quality and mode coverage), privacy (memorization and re-identification risks), and utility (downstream clinical performance for classification and report generation). It probes whether synthetic medical images can match real data in clinical tasks while preserving patient privacy. Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Re...
Evaluates a chest X-ray vision-language foundation model across zero-shot classification, cross-modal retrieval, and adapted downstream tasks (classification, segmentation, report generation). It specifically probes data and compute efficiency, as well as the model's ability to represent long-tailed thoracic diseases without aggressive scaling. Use when the user wants to benchmark on SIIM-PTX, Pneumonia2017, TBX11K, CheXpert, MIMIC-CXR, ChestX-ray14, VinDr-CXR, VinDr-PCXR, or asks about evalu...
Evaluates vision-language foundation models on chest X-ray interpretation across two axes: image perception (view classification, disease identification/classification, VQA, reasoning) and textual understanding (findings generation, summarization). It probes clinical reasoning, visual grounding, and medical text generation capabilities. Use when the user wants to benchmark on MIMIC-CXR, CheXpert, SIIM, RSNA, OpenI, SLAKE, Rad-Restruct, or asks about evaluating this task. Reports accuracy.
Evaluates a vision-language model's ability to perform interactive localization, region classification, and text generation on chest X-rays. It probes zero-shot multitask capabilities, including sentence grounding, pathology detection, and customizable report generation. Use when the user wants to benchmark on MS-CXR, VinDrCXR, NIH8, CIG, MIMIC-CXR, or asks about evaluating this task. Reports mAP.
Evaluates weakly-supervised multi-label classification and spatial localization of eight common thoracic diseases on chest X-rays. It probes a model's ability to detect disease presence from image-level labels and localize pathological regions using only bounding box annotations during testing. Use when the user wants to benchmark on ChestX-ray8, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to detect and localize multiple pathologies in high-resolution chest X-ray images. It specifically probes the model's robustness to severe class imbalance and its capacity to leverage explicit spatial location information for pathology classification. Use when the user wants to benchmark on ChestX-Ray14, PLCO, or asks about evaluating this task. Reports AUC.
Evaluates deep learning models for multi-label chest X-ray pathology classification. It probes the capability of CNN architectures to detect 14 specific thoracic diseases, testing the impact of transfer learning, network depth, and fusion of non-image patient demographics (age, gender, view position) on diagnostic accuracy. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AUC.
Evaluates the effectiveness of attribute-neutralization models in removing demographic bias (sex and age) from chest X-ray images while preserving diagnostic utility for 15 disease findings. It probes the trade-off between demographic leakage suppression and clinical performance across varying edit intensities. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AI-Judge AUC.
Evaluates deep learning models for binary classification of chest X-rays into normal versus pneumonia categories, while also assessing the spatial interpretability of model predictions using Grad-CAM heatmaps. Use when the user wants to benchmark on Chest X-Rays dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates instance-level detection and segmentation of thoracic diseases on chest X-rays. It probes a model's ability to localize and classify 13 disease categories while handling domain-specific challenges like ambiguous boundaries and disease co-occurrence. Use when the user wants to benchmark on ChestX-Det, DR-private, or asks about evaluating this task. Reports AP_50^bb.
Evaluates an AI agent's ability to perform multi-step medical reasoning and tool orchestration for chest X-ray interpretation. It probes capabilities across seven clinically relevant categories: detection, classification, localization, comparison, relationship, diagnosis, and characterization. Use when the user wants to benchmark on ChestAgentBench, or asks about evaluating this task. Reports Accuracy (%).
Evaluates a model's ability to perform multi-label classification of 14 thoracic diseases on chest X-ray images and localize pathological regions using attention maps. It probes whether anatomically grounded feature weighting improves disease detection over global feature fusion or saliency-based methods. Use when the user wants to benchmark on Chest X-ray14, or asks about evaluating this task. Reports AUROC.
This evaluation probes a model's ability to generate clinically coherent and accurate natural language reports from chest X-ray images. It measures both surface-level linguistic similarity to ground-truth radiology reports and the clinical correctness of extracted pathological findings. Use when the user wants to benchmark on Indiana U. Chest X-Ray, MIMIC-CXR, or asks about evaluating this task. Reports BLEU-1.
Evaluates a CNN's ability to classify chest X-ray images into disease categories (COVID-19, pneumonia, tuberculosis, normal) using various preprocessing techniques. It probes robustness across different dataset sizes and class distributions. Use when the user wants to benchmark on Multiclass Chest X-ray Dataset, Hamad Medical Corporation Tuberculosis Dataset, Pneumonia Dataset, NIH Chest X-ray Dataset, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to play chess and solve tactical puzzles without explicit search, testing its capacity for long-horizon planning and generalization to novel board states. Use when the user wants to benchmark on Lichess puzzles, or asks about evaluating this task. Reports Lichess Elo.
Evaluates multimodal large language models' ability to solve Olympiad-level theoretical chemistry problems requiring visual perception, chemical reasoning, and structured problem-solving. It specifically probes the visual perception bottleneck in chemistry tasks and tests the effectiveness of multi-agent orchestration and structured visual enhancement. Use when the user wants to benchmark on ChemO, or asks about evaluating this task. Reports normalized rubric-based score.
Evaluates LLM-based chemistry agents on four core computational chemistry tasks: text-to-molecule generation, molecule-to-text captioning, molecular property prediction, and reaction product prediction. It probes the model's ability to translate between chemical language and structures, predict physicochemical properties, and correct tool errors via hierarchical agent stacking. Use when the user wants to benchmark on ChemLLMBench, or asks about evaluating this task. Reports Accuracy.