Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,891
skills in category
996
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,881–5,904 of 23,891 skills

Persian Ner EvalA

Evaluates the quality and cross-lingual transferability of machine-translated Persian named entity recognition datasets by measuring model performance against original English benchmarks. It probes how well translation-based dataset generation preserves entity boundaries and labels across languages with different scripts and linguistic structures. Use when the user wants to benchmark on CoNLL 2003, OntoNotes 5.0, NCBI Disease, WNUT 2017, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Persian Instruction Following EvalA

Evaluates the instruction-following capability of Persian large language models across multiple NLP tasks, including paraphrasing, sentiment analysis, and textual entailment. Use when the user wants to benchmark on parsinlu queryparaphrasing, Digikala SentimentAnalysis, FarsTail, ParsinluEntailment, or asks about evaluating this task. Reports ROUGE-L F1.

researchpython
0
3
Persia Ctr EvalA

Evaluates the training efficiency, scalability, and convergence of a hybrid deep learning recommender system against baselines on click-through rate (CTR) prediction tasks. It measures end-to-end training time to reach target AUC, final test AUC for statistical efficiency, and training throughput across varying model scales up to 100 trillion parameters. Use when the user wants to benchmark on Taobao-Ad, Avazu-Ad, Criteo-Ad, Kwai-Video, Criteo-Syn, or asks about evaluating this task. Reports ...

researchpythongo
0
3
Persense D EvalA

Evaluates training-free, one-shot instance segmentation in dense, cluttered, and occluded scenes. It probes the model's ability to localize and segment specific target instances using few exemplars and point prompts, while handling high object density and overlapping objects. Use when the user wants to benchmark on PerSense-D, COCO-20i, COCO-20d, LVIS-92i, LVIS-92d, or asks about evaluating this task. Reports mIoU.

researchpythonperformance
0
3
Perrecbench EvalA

Evaluates whether LLMs can capture true personalized user preferences by ranking items or users in groups, while explicitly controlling for confounding factors like user rating bias and item quality. It probes the model's ability to perform comparative reasoning rather than simple rating prediction. Use when the user wants to benchmark on PerRecBench, or asks about evaluating this task. Reports Kendall’s tau.

researchpythongo
0
3
Perr EvalA

This benchmark evaluates a model's ability to recognize the emotional relationship (e.g., intimate, hostile, neutral) between two interacting characters in drama videos. It probes multi-modal fusion capabilities by requiring the model to integrate visual, audio, and textual cues to classify pairwise interactions. Use when the user wants to benchmark on ERATO, or asks about evaluating this task. Reports Micro-F1.

researchpythongo
0
3
PerplexityA

This protocol evaluates language model memorisation and training data contamination by measuring how well the model predicts benchmark text compared to out-of-distribution baselines. Lower perplexity on benchmark passages relative to a clean baseline indicates the model has likely seen the text during training. Use when the user has predictions and gold and needs to compute perplexity.

researchpythongo
0
3
Perla 3d EvalA

Evaluates a 3D vision-language model's ability to answer questions about indoor scenes and generate dense captions for 3D instances. It probes fine-grained spatial reasoning, object attribute recognition, and scene understanding by comparing generated text against ground-truth annotations using standard natural language generation metrics. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports CiDEr.

researchpythongo
0
3
Perceptionprocessbench EvalA

Evaluates a vision-language process reward model's ability to detect step-level visual grounding and reasoning errors in structured multimodal reasoning traces. The benchmark specifically probes whether the model can distinguish between correct and subtly mutated perception steps that are designed to be challenging for automated error detection. Use when the user wants to benchmark on PerceptionProcessBench, or asks about evaluating this task. Reports step-level correctness.

researchpythongo
0
3
Perceptioncomp EvalA

This benchmark evaluates long-horizon, perception-centric video reasoning in multimodal LLMs. It requires models to gather visual evidence across temporally separated segments and integrate multiple compositional constraints (e.g., object recognition, temporal tracking, spatial inference) to answer complex questions. Use when the user wants to benchmark on PerceptionComp, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Perception Test EvalA

Evaluates multimodal video models on core perception skills and reasoning types across six computational tasks, including tracking, temporal localization, and video question answering. It probes zero-shot and few-shot generalization on real-world videos with dense annotations. Use when the user wants to benchmark on Perception Test, or asks about evaluating this task. Reports top-1 accuracy.

researchpythongo
0
3
Per Object Depth Estimation EvalA

Evaluates the accuracy of per-object depth estimation for vehicles in autonomous driving scenarios. It measures how well a model predicts the depth of individual objects given monocular RGB images and 2D bounding boxes. Use when the user wants to benchmark on Waymo Open Dataset, KITTI Detection Dataset, KITTI MOT Dataset, or asks about evaluating this task. Reports delta<1.25.

researchpythongo
0
3
Peptide Protein Interaction EvalA

Evaluates a model's ability to predict whether a given peptide-protein pair interacts (binary classification) and to localize binding residues on both the peptide and protein sequences. It also assesses the model's capacity to generate target-specific peptide sequences that improve structural binding affinity over native templates. Use when the user wants to benchmark on Test167, LEADS-PEP, Test251, or asks about evaluating this task. Reports AUROC.

researchpythonapi
0
3
Peoples Speech EvalA

Evaluates the quality and generalization capability of a large-scale, commercially licensed speech recognition dataset by training an acoustic model on it and measuring word error rate on standard read-speech benchmarks. Use when the user wants to benchmark on The People's Speech, Librispeech, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythonperformance
0
3
Pencil Puzzle Bench EvalA

Evaluates multi-step verifiable reasoning and agentic iteration on constraint-satisfaction puzzles. It probes a model's ability to plan, execute moves, check constraints step-by-step, and course-correct over long contexts. Use when the user wants to benchmark on Pencil Puzzle Bench, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Pen4rec EvalA

Evaluates a model's ability to predict the next item in a session-based recommendation task by capturing evolving user preferences and mitigating preference drift over time. Use when the user wants to benchmark on Yoochoose, Diginetica, LastFM, PHEME, or asks about evaluating this task. Reports P@20.

researchpythonperformance
0
3
Peg Insertion EvalA

Evaluates a robot policy's ability to perform contact-rich manipulation by jointly reasoning over visual and haptic feedback. It measures how well a learned representation improves sample efficiency, generalizes across peg geometries, and recovers from perturbations during peg insertion tasks. Use when the user wants to benchmark on Custom Peg Insertion Environment, or asks about evaluating this task. Reports sum of rewards achieved in an episode, normalized by the highest attainable reward.

researchpythongo
0
3
Peerprism EvalA

Evaluates the ability of various LLM text detection methods to distinguish between human-written and AI-generated peer reviews, while also assessing their robustness to hybrid (human idea + AI text) workflows. Use when the user wants to benchmark on PeerPrism, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Peer Review Toxic Detection EvalA

This benchmark evaluates the ability of models to detect toxic sentences within academic peer reviews, where toxicity manifests as subtle emotive, rhetorical, or unconstructive language rather than overt abuse. It measures alignment with human judgments and assesses whether models can revise toxic sentences while preserving the original critique. Use when the user wants to benchmark on Peer Review Toxic Detection Dataset, or asks about evaluating this task. Reports Cohen's Kappa.

researchpythongo
0
3
Peer Review Analysis EvalA

This protocol evaluates the linguistic and content-level properties of academic peer review reports to assess how LLM assistance influences review quality, complexity, and aspect coverage over time. Use when the user wants to benchmark on ICLR & NeurIPS Peer Reviews, or asks about evaluating this task. Reports aspect_mentions.

researchpythongit
0
3
Peek Robot Zero Shot EvalA

Evaluates zero-shot generalization and visual/semantic robustness of robot manipulation policies when transferred to new real-world setups, visual clutter, and unseen object configurations. Use when the user wants to benchmark on Franka Sim-to-Real Custom Setup, BRIDGE-v2, or asks about evaluating this task. Reports success rate.

researchpython
0
3
Peek Engagement Prediction EvalA

This benchmark evaluates a model's ability to predict whether a learner will engage with a specific educational video fragment based on their historical interaction sequence and the fragment's content features. It probes sequential behavior modeling and content-based recommendation in informal, self-directed learning environments. Use when the user wants to benchmark on PEEK, or asks about evaluating this task. Reports F1-measure.

researchpythongo
0
3
Pediatric Brain Tumor Seg EvalA

Evaluates deep learning architectures for multi-class segmentation of pediatric brain tumors on MRI scans. It probes the model's ability to accurately delineate tumor sub-regions (whole tumor, enhanced tumor, cystic component, edema) and assesses cross-domain generalizability to adult glioma data. Use when the user wants to benchmark on PED BraTS 2024, CBTN, BraTS Adult Glioma 2023, or asks about evaluating this task. Reports lesion-wise Dice.

researchpythontesting
0
3
Pebench EvalA

Evaluates multimodal large language models on their ability to selectively forget specific person or event concepts while preserving general knowledge. It probes cross-concept interference, unlearning efficacy, and the trade-off between forgetting targeted data and maintaining model utility. Use when the user wants to benchmark on PEBench, or asks about evaluating this task. Reports Efficacy.

researchpythongo
0
3