Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,891
skills in category
996
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,785–5,808 of 23,891 skills

Polymath EvalA

Evaluates multi-modal mathematical and cognitive reasoning capabilities on visual puzzles. It probes spatial interpretation, relational understanding, pattern recognition, and long-horizon logical reasoning using diagram-based multiple-choice questions. Use when the user wants to benchmark on POLYMATH, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Polyglot Toxicity Prompts EvalA

Evaluates the toxicity of LLM-generated continuations across 17 languages using naturally occurring prompts scraped from the web. It probes how model size, language resource availability, and instruction/preference tuning affect the generation of harmful content. Use when the user wants to benchmark on PolygloToxicityPrompts (PTP), or asks about evaluating this task. Reports AT.

researchpythongit
0
3
Poly Fever EvalA

Evaluates large language models' ability to detect hallucinations by verifying factual claims across 11 languages. It probes cross-linguistic consistency, topic-aware fact-checking, and resistance to web-resource bias in multilingual settings. Use when the user wants to benchmark on Poly-FEVER, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Polqa EvalA

Evaluates open-domain question answering in Polish by measuring both passage retrieval accuracy and answer generation quality. It probes a model's ability to retrieve relevant evidence from a large corpus and accurately extract or generate answers from those passages. Use when the user wants to benchmark on PolQA, or asks about evaluating this task. Reports fuzzy_match.

researchpythongo
0
3
PolosA

Evaluates the alignment and quality of generated image captions relative to reference captions and source images. It measures how well a learned metric correlates with human judgments, specifically probing hallucination robustness and open-vocabulary caption evaluation. Use when the user has predictions and gold and needs to compute Polos.

researchpythongo
0
3
Political Toxicity Annotation EvalA

Evaluates the ability of LLMs and API-based classifiers to accurately annotate toxicity and incivility in political protest content against a human gold standard. It probes zero-shot classification performance, threshold sensitivity, and output reproducibility across different model sizes and temperatures. Use when the user wants to benchmark on Political protest content dataset, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Polish Nlp EvalA

Evaluates Polish language understanding, summarization, and question answering capabilities of text-to-text models. It probes how well encoder-decoder and decoder-only architectures generalize from multilingual pre-training to monolingual Polish tasks using exact-match generation and ROUGE-based metrics. Use when the user wants to benchmark on KLEJ benchmark, Allegro Articles, Polish Summaries Corpus, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Polish Medical Qa EvalA

Evaluates large language models on Polish medical licensing and specialization exams to assess cross-lingual medical knowledge transfer, domain-specific understanding, and specialty-level accuracy compared to human medical graduates. Use when the user wants to benchmark on Polish Medical Exams (LEK/LDEK/PES), or asks about evaluating this task. Reports score.

researchpythongo
0
3
Polish Asr EvalA

Evaluates the transcription accuracy of various automatic speech recognition (ASR) models on Polish-language audio, contrasting read-speech benchmarks with spontaneous, noisy medical consultations to probe domain generalization. Use when the user wants to benchmark on Mozilla Common Voice (MCV) Polish, Multilingual LibriSpeech (MLS) Polish, Medical interview corpus, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythondatabase
0
3
Policy Selection EvalA

Evaluates the model's ability to retrieve relevant coordination policies that guide task planning based on a high-level progress summary of the current state. Use when the user wants to benchmark on Policy Selection Test Suite, or asks about evaluating this task. Reports F1 Score.

researchpythonrust
0
3
Polaris EvalA

Evaluates the ability of machine learning models to distinguish between reference stars and circumstellar exoplanetary disks in high-contrast polarimetric imaging data. It probes representation learning quality through downstream supervised classification and unsupervised clustering tasks. Use when the user wants to benchmark on POLARIS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Polar Detection EvalA

This benchmark evaluates a model's ability to detect online polarization in social media text by classifying statements as polarized or non-polarized. It specifically probes the model's capacity to produce interpretable, structured reasoning alongside binary predictions while handling class imbalance and reducing false negatives. Use when the user wants to benchmark on POLAR @ SemEval-2026, or asks about evaluating this task. Reports macro-F1.

researchpythongo
0
3
Pokegym EvalA

PokeGym evaluates vision-language models' ability to perform long-horizon planning and spatial reasoning in a complex 3D open-world game using only raw RGB observations. It specifically probes visual grounding, autonomous goal decomposition, and physical deadlock recovery, revealing whether models can navigate cluttered environments, interact with objects, and recover from entrapment without explicit state feedback. Use when the user wants to benchmark on PokeGym, or asks about evaluating thi...

researchpythongo
0
3
Pokec N Fairness EvalA

Evaluates the ability of Graph Neural Networks to perform node classification while mitigating bias related to a protected attribute (Region). It probes the trade-off between predictive accuracy and group fairness across different GNN architectures. Use when the user wants to benchmark on Pokec-n, or asks about evaluating this task. Reports F1 score.

researchpythonnode
0
3
Poirot Alignment EvalA

Evaluates the ability to detect cyber attack campaigns by aligning threat intelligence query graphs with system provenance graphs derived from kernel audit logs. It probes structural pattern matching, causal dependency reasoning, and robustness against malware mutations and benign system noise. Use when the user wants to benchmark on DARPA TC Dataset, Public Malware Reports, or asks about evaluating this task. Reports alignment score.

researchpythongo
0
3
Points Long Video Image EvalA

Evaluates multimodal large language models on fine-grained image understanding and long-form video comprehension tasks. It measures the trade-off between visual token compression efficiency and task accuracy across diverse benchmarks. Use when the user wants to benchmark on MVBench, Video-MME, MLVU, LongVideoBench, MMBench, MMMU_val, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Pointgat Molecule C10 EvalA

Evaluates a hybrid graph attention and 3D point cloud neural network's ability to predict quantum chemical properties and molecular physicochemical traits. It probes the model's capacity to integrate 2D topological graph features with 3D spatial geometry for accurate regression and classification of molecular energies and properties. Use when the user wants to benchmark on MoleculeNet, C10, or asks about evaluating this task. Reports MAE, R².

researchpythongo
0
3
Point It Out EvalA

Evaluates vision-language models' embodied reasoning and visual grounding capabilities across three hierarchical stages: referred-object localization, task-driven pointing, and multi-step visual trace prediction in real-world scenarios. Use when the user wants to benchmark on Point-It-Out (PIO), or asks about evaluating this task. Reports score.

researchpythongo
0
3
Point Adjust F1A

Evaluates anomaly detection performance by granting full credit for all points in an anomalous segment if at least one point is detected, often inflating scores for algorithms that merely hit a segment once. Use when the user has predictions and gold and needs to compute point-adjust F1.

researchpythongo
0
3
Poi EvalA

Probes a model's ability to identify privacy-sensitive objects in images by reasoning about scene context rather than relying solely on visual appearance. It evaluates whether the system can distinguish between obvious privacy leaks (e.g., faces) and context-dependent sensitive information (e.g., people in specific roles). Use when the user wants to benchmark on MOSAIC, PRIVACY1000, or asks about evaluating this task. Reports F1 Score.

researchpythongo
0
3
Poet Weather Calibration EvalA

Evaluates the ability of a hierarchical transformer (PoET) to post-process and calibrate medium-range ensemble weather forecasts for 2m temperature and precipitation, compared to a baseline method (MBM) and raw ensemble outputs. Use when the user wants to benchmark on ECMWF ensemble forecasts, or asks about evaluating this task. Reports CRPS.

researchpythonperformance
0
3
Podcastmix EvalA

Evaluates the quality of monaural music and speech source separation in podcast audio. It measures both objective signal fidelity using BSS-eval metrics and subjective perceptual quality using standardized listening tests. Use when the user wants to benchmark on PodcastMix, or asks about evaluating this task. Reports SDR.

researchpythontesting
0
3
Po Meta Dataset EvalA

Evaluates few-shot representation learning under partial observability. Models must match query image views to their underlying source images using only partial support views (≤50% coverage) and viewpoint coordinates. Use when the user wants to benchmark on PO-Meta-Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Pnm Flow EvalA

Evaluates a topology-based pore network model's ability to predict flow-permeable surface area and hydraulic conductance in granular materials from micro-CT images. Use when the user wants to benchmark on Sphere Packing & High-Explosive Micro-CT Samples, or asks about evaluating this task. Reports conductance_ratio.

researchpythontesting
0
3