Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,891
skills in category
996
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,857–5,880 of 23,891 skills

Phomt EvalA

This benchmark evaluates Vietnamese-English machine translation quality by comparing neural baselines and commercial engines. It probes translation accuracy across multiple domains and sentence lengths using both automatic metrics and human preference judgments. Use when the user wants to benchmark on PhoMT, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Phocal EvalA

Evaluates category-level 6D object pose estimation on photometrically challenging objects (reflective, transparent, occluded). It tests both in-distribution generalization (seen objects) and out-of-distribution generalization (novel objects within the same category), comparing RGB-D and monocular approaches. Use when the user wants to benchmark on PhoCaL, or asks about evaluating this task. Reports 3D IoU.

researchpythongo
0
3
Phishing Detection EvalA

This benchmark evaluates machine learning classifiers and feature selection strategies for detecting phishing websites. It probes the ability of models to distinguish between legitimate and malicious web pages using content-based, external service, and hybrid feature sets, while measuring classification accuracy and macro F1-score. Use when the user wants to benchmark on Collected Phishing Dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongit
0
3
Phibench EvalA

Evaluates diverse reasoning and coding capabilities using an internal benchmark designed to minimize data contamination and LLM-judge bias. It probes a model's ability to debug, extend, and explain code, as well as identify errors in mathematical proofs and generate related problems. Use when the user wants to benchmark on PhiBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Phi3 Safety EvalA

Evaluates the safety and refusal capabilities of language models across multiple risk categories including harmful content generation, jailbreaking, stereotype bias, privacy leaks, and toxicity detection. It measures how well models balance harmlessness (refusing unsafe prompts) and helpfulness (complying with safe prompts) in both single- and multi-turn interactions. Use when the user wants to benchmark on XSTest, DecodingTrust, ToxiGen, XSafety, RTP-LX, Microsoft Internal Automated Measurem...

researchpythonrust
0
3
Phi Preference Hijacking EvalA

Probes the vulnerability of multi-modal large language models to inference-time adversarial image perturbations that hijack response preferences (e.g., personality, opinions, contrastive biases) without model retraining. It measures how effectively optimized images steer model outputs toward attacker-specified targets across text-only, multi-modal, and universal perturbation settings. Use when the user wants to benchmark on Anthropic Model-Written Evaluation Datasets (Advanced AI Risk & Hallu...

researchpython
0
3
Phi 4 Reasoning EvalA

Evaluates large language models on reasoning-specific capabilities including mathematics, scientific QA, coding, algorithmic planning, and spatial reasoning. It probes the model's ability to generate step-by-step solution traces and produce correct final answers under varying decoding temperatures and run counts. Use when the user wants to benchmark on AIME, GPQA Diamond, OmniMATH, LiveCodeBench, Codeforces, or asks about evaluating this task. Reports pass@1 accuracy.

researchpythongo
0
3
Pheno Ca EvalA

Evaluates deep neural networks on phenotypic drug discovery tasks using high-content screening images. It probes the model's ability to deconvolve mechanisms of action, molecular targets, and compound identities from cellular phenotypes, as well as zero-shot compound retrieval for CRISPR perturbations. Use when the user wants to benchmark on Pheno-CA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Phase Recovery Nmf EvalA

Evaluates the effectiveness of different phase recovery and source separation algorithms on audio mixtures. It probes how well models maintain phase consistency and reconstruct audio quality under blind and oracle conditions, particularly when time-frequency bins overlap. Use when the user wants to benchmark on Audio source separation mixtures (synthetic harmonics, piano notes, MIDI excerpt), or asks about evaluating this task. Reports SDR.

researchpythongo
0
3
Phase Picking EvalA

Evaluates a model's ability to detect seismic events and classify phase types (P vs S) from raw waveform windows, particularly under varying amounts of labeled training data. It probes the effectiveness of self-supervised pretraining in learning generalizable seismic features compared to randomly initialized baselines. Use when the user wants to benchmark on ETHZ, GEOFON, STEAD, or asks about evaluating this task. Reports AUC.

researchpython
0
3
Phase No EvalA

Evaluates the ability of neural phase pickers to detect P- and S-wave arrivals in continuous seismic waveforms across multi-station networks. It probes detection accuracy, timing precision, and generalization to out-of-distribution earthquake sequences under varying signal-to-noise conditions. Use when the user wants to benchmark on NCEDC 2020 Test Set, 2019 Ridgecrest Sequence, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Pharos Benchmark EvalA

Evaluates whether traditional tabular reinforcement learning hardness metrics (MDP diameter, suboptimality gaps, effective horizon) can predict the sample efficiency and performance of deep RL agents across different observation modalities and environment scales. Use when the user wants to benchmark on Pharos Benchmark, or asks about evaluating this task. Reports cumulative regret.

researchpythonperformance
0
3
Pharmaship EvalA

Evaluates layout-aware document understanding models on Chinese pharmaceutical shipping documents. It probes semantic entity recognition, entity linking, and reading order prediction, specifically testing robustness to dense tabular layouts and long-range semantic dependencies. Use when the user wants to benchmark on PharmaShip, or asks about evaluating this task. Reports F1.

researchpythontesting
0
3
Pgexplainer Gnn EvalA

Evaluates the predictive accuracy of GNN models and the effectiveness of post-hoc explanation methods on synthetic and real-world graph classification and node classification tasks. It probes whether a parameterized explainer can learn global explanatory motifs end-to-end and generalize inductively without retraining. Use when the user wants to benchmark on BA-Shapes, BA-Community, Tree-Cycles, Tree-Grid, BA-2motifs, MUTAG, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Pg Gnn EvalA

Evaluates the expressivity and predictive performance of permutation-sensitive Graph Neural Networks (PG-GNN) on synthetic substructure counting tasks and real-world graph classification/regression benchmarks. It probes the model's ability to capture pairwise node correlations and higher-order substructures (triangles, 4-cliques) compared to standard permutation-invariant GNNs. Use when the user wants to benchmark on Erdős-Rényi random graphs, Random regular graphs, TUDataset (PROTEINS, NCI1,...

researchpythongo
0
3
Pfm1 Landmine Detection EvalA

Evaluates the ability of statistical and learning-based detectors to identify sparse PFM-1 landmines in UAV-captured hyperspectral imagery, emphasizing performance under severe class imbalance and varying background clutter. Use when the user wants to benchmark on UAV Hyperspectral Imagery (PFM-1 Landmine Scene), or asks about evaluating this task. Reports AP.

researchpythonperformance
0
3
Pfm Gene Reg EvalA

Evaluates the ability of probability flow matching to infer stochastic gene regulatory dynamics and cell differentiation trajectories from time-resolved single-cell omics data. It probes interpolation accuracy, generalization to unseen initial conditions, and recovery of biologically validated gene-gene regulatory interactions. Use when the user wants to benchmark on 2D Ornstein-Uhlenbeck process, Multistable Waddington-like landscape, Ex vivo Hematopoiesis scRNA-seq, or asks about evaluating...

researchpythongo
0
3
Pevlm Longvideo EvalA

Evaluates the accuracy and latency-constrained performance of vision-language models on long-video understanding tasks. It measures how well models preserve temporal reasoning and answer questions about extended video sequences under strict computational and time budgets. Use when the user wants to benchmark on LongVideoBench, VideoMME, EgoSchema, MVBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Peta Protein EvalA

Evaluates protein language models across 33 downstream tasks (fitness, localization, PPI, solubility, and structure prediction) to measure how sub-word tokenization and vocabulary size affect representation quality and task performance. Use when the user wants to benchmark on PETA Benchmark Suite, or asks about evaluating this task. Reports Spearman correlation.

researchpythongo
0
3
Pet EvalA

This benchmark evaluates the capability of NLP models to extract structured business process elements and their relationships from unstructured natural language text. It specifically probes entity recognition (activities, actors, gateways, data) and relation detection (flow, usage, performer/recipient) under varying information availability assumptions. Use when the user wants to benchmark on PET, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Personamem EvalA

Evaluates large language models' ability to track dynamic user profile evolution over time and generate personalized responses to in-situ queries. It probes long-context memory, preference tracking, and contextual alignment across interleaved multi-session conversations. Use when the user wants to benchmark on PersonaMem, or asks about evaluating this task. Reports multiple-choice selection.

researchpythongo
0
3
Personalized Embodied Navigation EvalA

Evaluates an embodied agent's ability to navigate and ground objects based on user-specific ownership semantics provided only in text. It tests long-term memory, spatial reasoning, and the capacity to interpret personalized queries without relying on visual object cues. Use when the user wants to benchmark on PersONAL, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Persona Chat EvalA

Evaluates a dialogue agent's ability to generate or rank contextually appropriate next utterances while maintaining consistency with a given personal profile. It probes the model's capacity for persona-conditioned chit-chat and its ability to infer or reflect speaker interests during conversation. Use when the user wants to benchmark on PersonaChat, or asks about evaluating this task. Reports Hits@1.

researchpythongo
0
3
Person Detection EvalA

Evaluates person and body part detection accuracy using standard object detection metrics, while assessing a self-monitoring framework's ability to reduce false negatives and false positives through part-based plausibility checks. Use when the user wants to benchmark on DensePose, MS-COCO, Pascal VOC, or asks about evaluating this task. Reports AP@0.5.

researchpythonperformance
0
3