Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,907
skills in category
997
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,929–5,952 of 23,907 skills

Pdf Extraction EvalA

Evaluates open-source PDF information extraction tools across multiple content elements (metadata, references, tables, paragraphs, sections, etc.) on academic documents. It probes how well different tools handle layout-based segmentation, text extraction, and structural recognition in real-world academic PDFs. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Pdcovidnet EvalA

Evaluates a CNN's ability to classify chest X-ray images into three diagnostic categories: COVID-19, Normal, and Viral Pneumonia. It probes multi-scale feature extraction and robustness to class imbalance in medical imaging. Use when the user wants to benchmark on Custom benchmark dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Pdbbind Low Similarity EvalA

Evaluates how well 3D binding affinity models generalise to unseen proteins and novel ligands in low-data regimes. It uses a strict low-Tanimoto-similarity split of the PDBBind dataset to prevent data leakage and benchmark generalisation capabilities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Pdbbind Lba EvalA

Predicts the binding affinity between a protein pocket and a ligand from 3D structural data. It probes the model's ability to quantify molecular interaction strength and generalize across protein sequence identities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports RMSE.

researchpython
0
3
Pdbbind Binding Affinity EvalA

This benchmark evaluates the ability of docking tools, deep learning models, and meta-modeling ensembles to predict ligand-protein binding affinities. It probes how well different feature representations (physical scores, sequence-based DL outputs, physicochemical properties) generalize to unseen protein-ligand complexes. Use when the user wants to benchmark on PDBbind, or asks about evaluating this task. Reports Pearson correlation coefficient.

researchpythongo
0
3
Pcb Defect Classification EvalA

Evaluates a model's ability to classify six specific types of PCB manufacturing defects from cropped defect images. It probes the model's feature extraction and categorization capabilities on a specialized industrial computer vision dataset. Use when the user wants to benchmark on PCB Defect Dataset, or asks about evaluating this task. Reports average_precision_rate.

researchpythongo
0
3
Pbench EvalA

Evaluates a model's ability to perform referring expression segmentation across five hierarchical levels of semantic complexity, from basic object recognition to fine-grained attribute binding, OCR-based disambiguation, spatial layout understanding, and relational interactions. It also stress-tests long-context generation and instance stability in crowded scenes with high object counts. Use when the user wants to benchmark on PBench, or asks about evaluating this task. Reports per-level perfo...

researchpythongo
0
3
Paws X EvalA

PAWS-X evaluates a model's ability to identify paraphrases across multiple languages, specifically probing sensitivity to word order and syntactic structure under conditions of high lexical overlap. It measures how well models generalize cross-lingually when trained on machine-translated data versus zero-shot settings. Use when the user wants to benchmark on PAWS-X, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Paws EvalA

This benchmark evaluates a model's ability to identify paraphrases in sentence pairs that share high lexical overlap but differ in meaning due to word order and syntactic structure. It specifically probes sensitivity to non-local contextual information and adversarial word scrambling, revealing whether models rely on superficial word matching rather than true semantic understanding. Use when the user wants to benchmark on PAWS_QQP, PAWS_Wiki, or asks about evaluating this task. Reports classi...

researchpythongo
0
3
Patsql Sql Synthesis EvalA

Evaluates the ability of program-by-example (PBE) systems to synthesize correct SQL queries from example input/output tables. It probes query generation accuracy, synthesis speed, and scalability to larger database schemas. Use when the user wants to benchmark on ase13, so-top, so-dev, so-rec, kaggle, or asks about evaluating this task. Reports solve_rate.

researchpythongo
0
3
Patientsim EvalA

Evaluates the persona fidelity, factual accuracy, and clinical plausibility of an LLM-based patient simulator in doctor-patient dialogues. It measures how well the model adheres to assigned patient profiles, maintains factual consistency, and handles out-of-profile questions plausibly. Use when the user wants to benchmark on PatientSim Profiles, or asks about evaluating this task. Reports Entail (%).

researchpythongo
0
3
Pathology Vqa EvalA

Evaluates a multimodal chatbot's ability to interpret real-world pathology images (H&E and IHC) and integrate clinical context to produce accurate diagnoses, terminology, and multimodal reasoning across four anatomical systems. Use when the user wants to benchmark on Pathology Clinical Q&A Dataset, or asks about evaluating this task. Reports diagnosis accuracy.

researchpython
0
3
Path Specific Fairness EvalA

Evaluates a model's ability to make predictions while removing the influence of a sensitive attribute along specific causal pathways, balancing predictive accuracy with path-specific counterfactual fairness constraints. Use when the user wants to benchmark on Berkeley Admission Dataset, UCI Adult Dataset, UCI German Credit Dataset, or asks about evaluating this task. Reports fair accuracy.

researchpythongo
0
3
Patenteb EvalA

Evaluates patent text embedding models across 15 diverse tasks including symmetric/asymmetric retrieval, classification, paraphrase detection, and clustering. It specifically probes domain-specific challenges like cross-domain retrieval, fragment-to-document matching, and temporal citation dynamics. Use when the user wants to benchmark on PatenTEB, or asks about evaluating this task. Reports NDCG@10, Macro-F1, Pearson r, V-measure.

researchpythongo
0
3
Patent Ce EvalA

This benchmark evaluates the quality of generated patent claims against expert-annotated reference claims across five dimensions: feature completeness, conceptual clarity, terminology consistency, logical linkage, and overall quality. It probes a model's ability to capture patent-specific linguistic precision, legal formality, and structural requirements rather than just surface-level text overlap. Use when the user wants to benchmark on Patent-CE, or asks about evaluating this task. Reports ...

researchpythongo
0
3
Patchgastricadc22 EvalA

This evaluation probes a model's ability to generate clinically accurate diagnostic captions from histopathological image patches. It specifically tests the model's capacity to capture subtype-specific terminology and overall caption fluency using standard and custom n-gram overlap metrics. Use when the user wants to benchmark on PatchGastricADC22, or asks about evaluating this task. Reports BLEU@4.

researchpythongit
0
3
Patch Selectivity EvalA

Evaluates a model's ability to ignore out-of-context patches (patch selectivity) and maintain classification accuracy under simulated occlusion and spatial permutation attacks. Use when the user wants to benchmark on ImageNet-1K val, SMD, NVD, ROD, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Pat Questions EvalA

Evaluates large language models' ability to answer present-anchored temporal questions that require up-to-date world knowledge and multi-hop reasoning, such as identifying the current holder of a position or the previous president. It specifically probes performance degradation due to knowledge obsolescence and complex temporal relations. Use when the user wants to benchmark on PAT-Questions, or asks about evaluating this task. Reports exact-match accuracy (EM).

researchpythongo
0
3
Pastis Hd EvalA

Tests agricultural land cover mapping and crop-type classification by evaluating models on high-resolution satellite imagery combined with optical and radar time series. Use when the user wants to benchmark on PASTIS-HD, or asks about evaluating this task. Reports macro-averaged F1-score.

researchpythongo
0
3
Pascal Voc Detection EvalA

Evaluates an object detection model's ability to localize and classify objects within images. It measures how well the system predicts bounding boxes and assigns correct class labels across multiple object categories. Use when the user wants to benchmark on PASCAL VOC 2007, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Pas Dataset EvalA

Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.

researchpythonperformance
0
3
Partinstruct EvalA

This benchmark probes a robot policy's ability to follow fine-grained, part-level natural language instructions for long-horizon manipulation. It specifically tests zero-shot task decomposition, 3D part grounding, and multi-step planning under varying object, part, and task generalization conditions. Use when the user wants to benchmark on PartInstruct, or asks about evaluating this task. Reports success.

researchpythongo
0
3
Parsinlu EvalA

Evaluates Persian language understanding across six distinct NLU tasks, including reading comprehension, textual entailment, sentiment analysis, and machine translation. It measures how well pre-trained monolingual and multilingual models perform on native-speaker annotated Persian data compared to human baselines. Use when the user wants to benchmark on ParsiNLU, or asks about evaluating this task. Reports F1, Accuracy.

researchpythongo
0
3
Parrot Multilingual EvalA

Evaluates the multilingual visual-language understanding capabilities of multimodal large language models (MLLMs) across six languages (English, Chinese, Portuguese, Arabic, Turkish, Russian). It probes how well models align visual features with non-English textual instructions and handle cross-lingual multimodal tasks without relying on naive translation. Use when the user wants to benchmark on MMMB, MMBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3