Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,891
skills in category
996
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,809–5,832 of 23,891 skills

Pneumonia Xray Zero Shot EvalA

This benchmark evaluates the zero-shot diagnostic capability of vision-language models on chest X-ray images for binary pneumonia detection. It probes whether models can accurately classify radiological findings without task-specific fine-tuning, relying instead on prompt engineering and pre-trained visual reasoning. Use when the user wants to benchmark on Chest radiographic Images (Pneumonia), or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Pn Summary EvalA

Evaluates the ability of transformer-based models to generate abstractive summaries in Persian. It measures how well generated summaries match reference summaries in terms of lexical overlap and longest common subsequence at the sentence level. Use when the user wants to benchmark on pn-summary, or asks about evaluating this task. Reports ROUGE-1 F-1.

researchpythongit
0
3
Pmmeval EvalA

Evaluates multilingual capabilities of LLMs across understanding, reasoning, and generation tasks in 10 languages. It probes prompt sensitivity and cross-lingual performance consistency to reveal benchmark origin bias and language-specific scaling trends. Use when the user wants to benchmark on MMMLU, MLogiQA, MGSM, MHellaSwag, XNLI, Flores-200, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Pmindia Nmt EvalA

This evaluation probes the quality of automatic machine translation between English and 13 Indian languages using a parallel corpus. It measures how well NMT systems can handle diverse linguistic structures, including abugida scripts and agglutinative morphology, across low-resource language pairs. Use when the user wants to benchmark on PMIndia, or asks about evaluating this task. Reports BLEU.

researchpythonshell
0
3
Pmemo Emotion Recognition EvalA

Binary classification of user-independent emotional states (valence and arousal) from electrodermal activity (EDA) signals. It probes the model's ability to generalize across subjects by using subject-specific thresholds and fusing physiological signals with external music benchmarks. Use when the user wants to benchmark on PMEmo, or asks about evaluating this task. Reports accuracy.

researchpython
0
3
Plutus Ben EvalA

This benchmark evaluates large language models on five core Greek financial NLP tasks: numeric and textual named entity recognition, multiple-choice question answering, abstractive summarization, and financial topic classification. It probes models' ability to handle low-resource language morphology, domain-specific financial terminology, and reasoning within Greek financial contexts. Use when the user wants to benchmark on GRFinNUM, GRFinNER, GRFinQA, GRFNS-2023, GRMultiFin, or asks about ev...

researchpythongo
0
3
Plotchain EvalA

This benchmark evaluates multimodal LLMs on engineering plot reading and visual quantitative reasoning. It probes the model's ability to interpret complex axes (including log scales), read curve values, and compute derived engineering quantities like cutoff frequencies or settling times from rendered plot images. Use when the user wants to benchmark on PlotChain, or asks about evaluating this task. Reports field-level accuracy.

researchpythongo
0
3
Plo Abbreviation Detection EvalA

Evaluates sequence labeling models on detecting abbreviations and extracting their corresponding long forms in scientific text. It probes domain-specific NER capabilities under challenges like context dependency, sub-abbreviations, and polysemy. Use when the user wants to benchmark on PLOD, SDU@AAAI-22 Shared Task, or asks about evaluating this task. Reports F.

researchpythongo
0
3
Plm Structural Pruning EvalA

Evaluates structural pruning methods for pre-trained language models (BERT-base, RoBERTa-base) across eight text classification tasks. It measures the trade-off between model size (parameter count) and task performance (validation error) to identify Pareto-optimal sub-networks. Use when the user wants to benchmark on eight text classification tasks, or asks about evaluating this task. Reports Hypervolume.

researchpythonperformance
0
3
Platinum Benchmarks EvalA

Evaluates LLM reliability on curated, low-noise subsets of standard benchmarks (VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench) by removing ambiguous examples and re-labeling to minimize ground-truth errors, revealing true model failures on elementary reasoning tasks. Use when the user wants to benchmark on VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Planviz EvalA

Evaluates multimodal models' ability to generate and edit images that correctly follow planning-oriented instructions (route planning, workflow diagramming, web/UI displaying). It probes procedural reasoning, spatial consistency, and semantic alignment in visual synthesis. Use when the user wants to benchmark on PlanViz, or asks about evaluating this task. Reports Cor.

researchpythongo
0
3
Plantvillagevqa EvalA

This benchmark evaluates vision-language models on plant science tasks, ranging from basic species and health identification to detailed symptom verification and higher-order causal or counterfactual reasoning. It probes a model's ability to ground visual attributes, diagnose diseases, and generate descriptive or diagnostic text based on leaf images. Use when the user wants to benchmark on PlantVillageVQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Plantseg EvalA

Evaluates pixel-level segmentation capabilities for identifying and localizing plant diseases in real-world, uncontrolled agricultural imagery across 115 disease classes. The benchmark tests a model's ability to handle fine-grained lesion boundaries, overlapping disease symptoms, and high visual diversity typical of field-captured crops. Use when the user wants to benchmark on PlantSeg, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Planetarium EvalA

Evaluates an LLM's ability to translate natural language planning task descriptions into valid, semantically equivalent Planning Domain Definition Language (PDDL) code. It specifically probes the model's capacity to accurately capture initial states, goal states, and object relationships while adhering to formal planning semantics. Use when the user wants to benchmark on Planetarium, or asks about evaluating this task. Reports equivalence.

researchpythongo
0
3
Pkugoodsad EvalA

Evaluates unsupervised visual anomaly detection and segmentation models on real-world supermarket goods. It probes robustness to object misalignment, intra-class appearance variation, and the ability to detect subtle or small anomalies without labeled anomalous training data. Use when the user wants to benchmark on PKU-GoodsAD, or asks about evaluating this task. Reports AUROC, AUPR.

researchpythongo
0
3
Pku Saferealf EvalA

Probes an LLM's ability to generate safe and helpful responses by classifying harmful content across 19 distinct categories and 3 severity levels, while aligning with human preference rankings on Q-A-B triplets. Use when the user wants to benchmark on PKU-SafeRLHF, or asks about evaluating this task. Reports harm_category.

researchpythongo
0
3
Pkgt Kg Construction EvalA

Evaluates the PheKnowLator ecosystem's software features and computational performance against other biomedical KG construction tools, and measures construction efficiency on 12 benchmark knowledge graphs. Use when the user has predictions and gold and needs to compute coverage score.

researchpythongo
0
3
Pkad R EvalA

This evaluation probes the ability of hybrid quantum-classical models to accurately predict residue-level pKa values across diverse protein microenvironments. It tests whether entanglement-aware quantum feature mappings generalize beyond the training distribution to experimental datasets and capture subtle electronic and geometric correlations in flexible peptide regions. Use when the user wants to benchmark on PKAD-R, Aβ40, or asks about evaluating this task. Reports RMSE.

researchpythonperformance
0
3
Pixelrec EvalA

Evaluates the ability of recommender systems to rank items using raw pixel images instead of traditional ID embeddings. It probes cold-start item recommendation, cross-domain transfer learning, and end-to-end vision-based recommendation performance. Use when the user wants to benchmark on PixelRec, or asks about evaluating this task. Reports Recall@N.

researchpythongit
0
3
Pixel3dmm 3dface Recon EvalA

This benchmark evaluates single-image 3D face reconstruction pipelines on two tasks: posed reconstruction (measuring geometric fidelity under diverse expressions) and neutral reconstruction (testing the ability to disentangle identity shape from expression). It probes a model's capacity to recover accurate facial geometry and surface normals from a single input image. Use when the user wants to benchmark on Pixel3DMM Benchmark, or asks about evaluating this task. Reports Chamfer Distance (L1/...

researchpythonexpress
0
3
Pixel Reasoner EvalA

Evaluates multimodal models' ability to perform fine-grained visual reasoning, object counting, temporal video understanding, and complex infographic parsing. It specifically probes whether models can effectively leverage pixel-space operations (e.g., zooming, frame selection) rather than defaulting to text-only reasoning pathways. Use when the user wants to benchmark on V* (V-Star), TallyQA, MVBench, InfographicVQA, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Pixel Face EvalA

Evaluates the capability of 3D face reconstruction models to predict accurate 3D meshes and facial landmarks from 2D RGB images. It probes how well models generalize to high-resolution, in-the-wild faces across diverse ages and expressions, highlighting domain gaps from synthetic training data. Use when the user wants to benchmark on Pixel-Face, or asks about evaluating this task. Reports ARMSE.

researchpythonexpress
0
3
Pistachio EvalA

Evaluates video anomaly detection and understanding on synthetic, long-form videos with diverse scenes and balanced anomaly categories. Probes models' temporal consistency, ability to detect subtle behavioral anomalies, and long-context narrative comprehension. Use when the user wants to benchmark on Pistachio, or asks about evaluating this task. Reports frame-level AUC.

researchpythongo
0
3
Pira Bench EvalA

This benchmark evaluates multimodal large language models on proactive intent recommendation from continuous, noisy GUI visual streams. It probes the model's ability to track long-horizon user behavior, distinguish true latent goals from background noise, and exercise operational restraint by remaining silent when no action is required. Use when the user wants to benchmark on PIRA-Bench, or asks about evaluating this task. Reports S_final.

researchpythongo
0
3