Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,640
skills in category
985
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,425–5,448 of 23,640 skills

Pulseimpute EvalA

Evaluates models on imputing missing values in pulsative physiological signals (ECG and PPG) under realistic, data-driven missingness patterns. It further assesses clinical utility by measuring downstream performance on heartbeat detection and cardiac classification tasks. Use when the user wants to benchmark on ECG, PPG, or asks about evaluating this task. Reports MSE, F1 Score.

researchpythongo
0
3
Pubtables V2 EvalA

Evaluates vision-language and specialized models on page-level and document-level table structure recognition, requiring them to extract hierarchical table structures from full pages or multi-page documents. It also probes cross-page table continuation prediction by testing whether models can identify when a table spans two contiguous pages. Use when the user wants to benchmark on PubTables-v2, or asks about evaluating this task. Reports GriTS_Top.

researchpythongo
0
3
Pubmedqa EvalA

This benchmark evaluates a model's ability to perform biomedical research question answering by reasoning over structured scientific abstracts. It requires models to infer yes/no/maybe answers to questions derived from paper titles using only the non-conclusion sections of the abstract, without access to the final conclusion. Use when the user wants to benchmark on PubMedQA (PQA-L), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Pubmedfact1k EvalA

This evaluation probes a model's ability to verify scientific claims in a three-way classification setting that includes uncertainty abstention. It measures performance on Supported, Refuted, and NEI (Not Enough Information) labels, testing the model's capacity to avoid overconfident predictions when evidence is insufficient or conflicting. Use when the user wants to benchmark on PubMedFact1k, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Publish And Perish EvalA

Evaluates a dynamical systems model of scientific publishing by simulating the interplay between AI-accelerated manuscript writing and peer review throughput. It measures how queue pressure drives AI adoption in review, degrades verification quality, and ultimately impacts net scientific knowledge output over a 20-year horizon. Use when the user wants to benchmark on NeurIPS main track submissions, ICLR submissions, arXiv monthly submissions, bioRxiv annual preprints, or asks about evaluating...

researchpython
0
3
Public Defense Retrieval EvalA

Evaluates the ability of retrieval and reranking models to surface relevant appellate brief paragraphs for public defender search queries. It probes domain-specific adaptation, query expansion strategies, and the impact of synthetic data generation on legal information retrieval. Use when the user wants to benchmark on PD Dataset, NJ OPD Dataset, BarExam-QA, LePaRD, or asks about evaluating this task. Reports recall@5.

researchpythongo
0
3
Publaynet Layout EvalA

Evaluates the ability of object detection models to identify and localize document layout elements (text, title, list, table, figure) in scientific PDF pages. It also probes transfer learning capabilities by fine-tuning on out-of-domain documents and table detection tasks. Use when the user wants to benchmark on PubLayNet, or asks about evaluating this task. Reports MAP @ IOU [0.50:0.95].

researchpythongo
0
3
Pubhealth Fact Checking EvalA

This benchmark evaluates automated fact-checking systems on public health claims, measuring both their ability to predict the veracity of a claim and the quality of the generated explanations. It probes domain-specific reasoning, extractive-abstractive explanation generation, and formal coherence properties like consistency and relevance. Use when the user wants to benchmark on PUBHEALTH, or asks about evaluating this task. Reports macroF1.

researchpythongo
0
3
Pts3d Llm EvalA

Evaluates multimodal large language models on 3D scene understanding tasks, including visual grounding, dense captioning, and spatial/situated question answering. It specifically probes how different 3D token structures (point-based vs. video-based) and feature fusion strategies impact performance on indoor RGB-D scans. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports NS.

researchpythonperformance
0
3
Ptbxl Multi Label Ecg EvalA

Evaluates deep learning architectures for multi-label classification of 12-lead ECG recordings into 23 diagnostic categories. Probes the trade-off between local morphological feature extraction and sequential temporal modeling under severe class imbalance. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports Macro AUROC.

researchpythongo
0
3
Ptbxl Ecg EvalA

Evaluates the ability of foundation models to perform multi-label classification on clinical 12-lead ECG recordings. It probes robustness to class imbalance and sample efficiency by measuring performance across varying label sets and training data sizes. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports macro AUROC.

researchpythongo
0
3
Ptbxl Ecg Classification EvalA

Evaluates the ability of self-supervised pre-trained Vision Transformers to classify ECG signals across multiple diagnostic label hierarchies (e.g., all statements, ST-MEM labels, diagnostic subclasses, rhythm statements). Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports macro AUC.

researchpythonperformance
0
3
Ptbxl Ecg Anomaly Detection EvalA

This evaluation probes a model's ability to detect cardiovascular anomalies in 12-lead electrocardiogram (ECG) signals and localize the specific temporal regions where abnormalities occur. It tests the model's capacity to capture both global and local temporal dependencies in raw, unsegmented time-series data without relying on traditional R-peak detection or heartbeat segmentation. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Ptbxl Af Detection EvalA

This benchmark evaluates how ECG sampling frequency impacts deep learning models for binary atrial fibrillation detection. It probes a model's discrimination capability and the reliability of its predicted probabilities across different temporal resolutions, while enforcing strict patient-level separation to prevent data leakage. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Ptbx1 Ecg Statement Prediction EvalA

Evaluates deep learning models on 12-lead ECG time series for multi-label classification of diagnostic, rhythm, and form statements. It probes the ability of architectures to learn directly from raw signals versus traditional feature extraction, and assesses transfer learning and demographic attribute prediction capabilities. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports term-centric macro-averaged AUC.

researchpythongo
0
3
Ptb Wikitext2 Lm EvalA

Evaluates next-token prediction accuracy and long-range dependency modeling in language models, with a specific focus on handling rare and out-of-vocabulary words without expanding vocabulary size. Use when the user wants to benchmark on Penn Treebank, WikiText-2, or asks about evaluating this task. Reports perplexity.

researchpythonperformance
0
3
Pt En Scientific Abstracts EvalA

Evaluates the quality of a newly constructed Portuguese-English parallel corpus of scientific abstracts. It probes the effectiveness of automated sentence alignment algorithms and the translation performance of SMT and NMT models on domain-specific academic text. Use when the user wants to benchmark on Theses and Dissertations Abstracts Corpus, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Psyeval EvalA

Evaluates an AI model's ability to generate empathetic, principle-constrained psychological counseling dialogues in simulated multi-turn interactions. It probes competencies like accurate empathy, logical consistency, resistance handling, and ethical guidance beyond surface-level language features. Use when the user wants to benchmark on PsyEval, or asks about evaluating this task. Reports PsyEval.

researchpythonapi
0
3
Psychomotor Skill Benchmarking EvalA

Evaluates the objective quantification and benchmarking of psychomotor execution quality in sports using wearable IMU data. It maps raw 3D motion trajectories into a normalized performance space and uses unsupervised clustering to identify optimal movement patterns and detect technical deviations. Use when the user wants to benchmark on Table Tennis Forehand Stroke (IMU), or asks about evaluating this task. Reports Euclidean distance to ideal performance origin.

researchpythonperformance
0
3
Psrb EvalA

This benchmark evaluates the robustness and accuracy of automatic speech recognition (ASR) systems for the Persian language across diverse acoustic conditions, demographic groups, and linguistic domains. It specifically probes how well models handle regional accents, spontaneous or informal speech, and underrepresented demographics, while highlighting architectural and data-related performance gaps. Use when the user wants to benchmark on PSRB, or asks about evaluating this task. Reports SW-WER.

researchpythonperformance
0
3
Psp Protein Structure EvalA

Evaluates protein structure prediction models on their ability to infer accurate 3D atomic coordinates from amino acid sequences. It specifically probes topological backbone similarity and side-chain accuracy when trained on large-scale distilled protein datasets. Use when the user wants to benchmark on CASP14 test set, PSP dataset, or asks about evaluating this task. Reports TM-score.

researchpythonapi
0
3
Psp Accent EvalA

Evaluates the phonological accent fidelity and prosodic naturalness of Indic text-to-speech systems across Hindi, Telugu, and Tamil. It decomposes accent into per-phoneme dimensions (retroflex, aspiration, Tamil-zha, vowel-length) and corpus-level distributional metrics, revealing gaps between intelligibility and native-like accent. Use when the user wants to benchmark on PSP Benchmark Sets, or asks about evaluating this task. Reports FAD.

researchpythongo
0
3
Psiloqa EvalA

Evaluates the ability of models to detect span-level hallucinations in multilingual question-answering contexts. It probes cross-lingual generalization and token-level inconsistency detection between generated answers and ground truth. Use when the user wants to benchmark on PsiloQA, or asks about evaluating this task. Reports IoU.

researchpythongo
0
3
Psgg Metrics EvalA

Evaluates panoptic scene graph generation models on their ability to predict object triplets with correct masks and relations. It probes recall-based performance at different top-k limits, mean recall across predicates, pair recall, and predicate ranking accuracy. Use when the user wants to benchmark on PSG, or asks about evaluating this task. Reports Mean Recall@k.

researchpythongo
0
3