Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,929–5,952 of 23,907 skills
Evaluates open-source PDF information extraction tools across multiple content elements (metadata, references, tables, paragraphs, sections, etc.) on academic documents. It probes how well different tools handle layout-based segmentation, text extraction, and structural recognition in real-world academic PDFs. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 score.
Evaluates a CNN's ability to classify chest X-ray images into three diagnostic categories: COVID-19, Normal, and Viral Pneumonia. It probes multi-scale feature extraction and robustness to class imbalance in medical imaging. Use when the user wants to benchmark on Custom benchmark dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates how well 3D binding affinity models generalise to unseen proteins and novel ligands in low-data regimes. It uses a strict low-Tanimoto-similarity split of the PDBBind dataset to prevent data leakage and benchmark generalisation capabilities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports performance.
Predicts the binding affinity between a protein pocket and a ligand from 3D structural data. It probes the model's ability to quantify molecular interaction strength and generalize across protein sequence identities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports RMSE.
This benchmark evaluates the ability of docking tools, deep learning models, and meta-modeling ensembles to predict ligand-protein binding affinities. It probes how well different feature representations (physical scores, sequence-based DL outputs, physicochemical properties) generalize to unseen protein-ligand complexes. Use when the user wants to benchmark on PDBbind, or asks about evaluating this task. Reports Pearson correlation coefficient.
Evaluates a model's ability to classify six specific types of PCB manufacturing defects from cropped defect images. It probes the model's feature extraction and categorization capabilities on a specialized industrial computer vision dataset. Use when the user wants to benchmark on PCB Defect Dataset, or asks about evaluating this task. Reports average_precision_rate.
Evaluates a model's ability to perform referring expression segmentation across five hierarchical levels of semantic complexity, from basic object recognition to fine-grained attribute binding, OCR-based disambiguation, spatial layout understanding, and relational interactions. It also stress-tests long-context generation and instance stability in crowded scenes with high object counts. Use when the user wants to benchmark on PBench, or asks about evaluating this task. Reports per-level perfo...
PAWS-X evaluates a model's ability to identify paraphrases across multiple languages, specifically probing sensitivity to word order and syntactic structure under conditions of high lexical overlap. It measures how well models generalize cross-lingually when trained on machine-translated data versus zero-shot settings. Use when the user wants to benchmark on PAWS-X, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to identify paraphrases in sentence pairs that share high lexical overlap but differ in meaning due to word order and syntactic structure. It specifically probes sensitivity to non-local contextual information and adversarial word scrambling, revealing whether models rely on superficial word matching rather than true semantic understanding. Use when the user wants to benchmark on PAWS_QQP, PAWS_Wiki, or asks about evaluating this task. Reports classi...
Evaluates the ability of program-by-example (PBE) systems to synthesize correct SQL queries from example input/output tables. It probes query generation accuracy, synthesis speed, and scalability to larger database schemas. Use when the user wants to benchmark on ase13, so-top, so-dev, so-rec, kaggle, or asks about evaluating this task. Reports solve_rate.
Evaluates the persona fidelity, factual accuracy, and clinical plausibility of an LLM-based patient simulator in doctor-patient dialogues. It measures how well the model adheres to assigned patient profiles, maintains factual consistency, and handles out-of-profile questions plausibly. Use when the user wants to benchmark on PatientSim Profiles, or asks about evaluating this task. Reports Entail (%).
Evaluates a multimodal chatbot's ability to interpret real-world pathology images (H&E and IHC) and integrate clinical context to produce accurate diagnoses, terminology, and multimodal reasoning across four anatomical systems. Use when the user wants to benchmark on Pathology Clinical Q&A Dataset, or asks about evaluating this task. Reports diagnosis accuracy.
Evaluates a model's ability to make predictions while removing the influence of a sensitive attribute along specific causal pathways, balancing predictive accuracy with path-specific counterfactual fairness constraints. Use when the user wants to benchmark on Berkeley Admission Dataset, UCI Adult Dataset, UCI German Credit Dataset, or asks about evaluating this task. Reports fair accuracy.
Evaluates patent text embedding models across 15 diverse tasks including symmetric/asymmetric retrieval, classification, paraphrase detection, and clustering. It specifically probes domain-specific challenges like cross-domain retrieval, fragment-to-document matching, and temporal citation dynamics. Use when the user wants to benchmark on PatenTEB, or asks about evaluating this task. Reports NDCG@10, Macro-F1, Pearson r, V-measure.
This benchmark evaluates the quality of generated patent claims against expert-annotated reference claims across five dimensions: feature completeness, conceptual clarity, terminology consistency, logical linkage, and overall quality. It probes a model's ability to capture patent-specific linguistic precision, legal formality, and structural requirements rather than just surface-level text overlap. Use when the user wants to benchmark on Patent-CE, or asks about evaluating this task. Reports ...
This evaluation probes a model's ability to generate clinically accurate diagnostic captions from histopathological image patches. It specifically tests the model's capacity to capture subtype-specific terminology and overall caption fluency using standard and custom n-gram overlap metrics. Use when the user wants to benchmark on PatchGastricADC22, or asks about evaluating this task. Reports BLEU@4.
Evaluates a model's ability to ignore out-of-context patches (patch selectivity) and maintain classification accuracy under simulated occlusion and spatial permutation attacks. Use when the user wants to benchmark on ImageNet-1K val, SMD, NVD, ROD, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates large language models' ability to answer present-anchored temporal questions that require up-to-date world knowledge and multi-hop reasoning, such as identifying the current holder of a position or the previous president. It specifically probes performance degradation due to knowledge obsolescence and complex temporal relations. Use when the user wants to benchmark on PAT-Questions, or asks about evaluating this task. Reports exact-match accuracy (EM).
Tests agricultural land cover mapping and crop-type classification by evaluating models on high-resolution satellite imagery combined with optical and radar time series. Use when the user wants to benchmark on PASTIS-HD, or asks about evaluating this task. Reports macro-averaged F1-score.
Evaluates an object detection model's ability to localize and classify objects within images. It measures how well the system predicts bounding boxes and assigns correct class labels across multiple object categories. Use when the user wants to benchmark on PASCAL VOC 2007, or asks about evaluating this task. Reports mAP.
Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.
This benchmark probes a robot policy's ability to follow fine-grained, part-level natural language instructions for long-horizon manipulation. It specifically tests zero-shot task decomposition, 3D part grounding, and multi-step planning under varying object, part, and task generalization conditions. Use when the user wants to benchmark on PartInstruct, or asks about evaluating this task. Reports success.
Evaluates Persian language understanding across six distinct NLU tasks, including reading comprehension, textual entailment, sentiment analysis, and machine translation. It measures how well pre-trained monolingual and multilingual models perform on native-speaker annotated Persian data compared to human baselines. Use when the user wants to benchmark on ParsiNLU, or asks about evaluating this task. Reports F1, Accuracy.
Evaluates the multilingual visual-language understanding capabilities of multimodal large language models (MLLMs) across six languages (English, Chinese, Portuguese, Arabic, Turkish, Russian). It probes how well models align visual features with non-English textual instructions and handle cross-lingual multimodal tasks without relying on naive translation. Use when the user wants to benchmark on MMMB, MMBench, or asks about evaluating this task. Reports Accuracy.