Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,785
skills in category
992
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,641–5,664 of 23,785 skills

Prior Loss Seg BenchmarkA

Evaluates the effectiveness of various prior-based loss functions (low-level boundary/distance and high-level shape/size constraints) for medical image segmentation across diverse anatomical structures and imaging modalities. Use when the user wants to benchmark on WMH, ISLES, Atrium, Colon, Spleen, Hippocampus, Prostate, ACDC, or asks about evaluating this task. Reports Dice score.

researchpythongo
0
3
Principlismqa EvalA

Evaluates large language models' ability to reason about medical ethics using the Principlism framework (autonomy, non-maleficence, beneficence, justice). It probes both theoretical knowledge of ethical principles and their practical application to complex, open-ended clinical dilemmas. Use when the user wants to benchmark on PrinciplismQA, or asks about evaluating this task. Reports Knowledge accuracy, Practice score.

researchpythongo
0
3
Principle Alignment EvalA

Probes an LLM's ability to align generated responses with a set of natural language constitutional principles without parameter fine-tuning. It measures both overall conformance quality and the reduction of critical principle violations through an inference-time self-correction pipeline. Use when the user wants to benchmark on SafeRLHF, HH-RLHF, or asks about evaluating this task. Reports 5-Point Likert Score Ranking.

researchpythonperformance
0
3
Primesrl EvalA

Evaluates the quality of Semantic Role Labeling (SRL) systems by measuring precision and recall for predicate senses and argument annotations. It specifically probes a model's step-dependent error propagation by penalizing argument scores when the associated predicate sense is incorrect, while also handling discontinuous and reference arguments. Use when the user has predictions and gold and needs to compute PriMeSRL-Eval.

researchpythongo
0
3
Prime Dp EvalA

Evaluates a pre-trained seismic model's multi-task capability on single-station waveforms, specifically phase picking (Pg, Sg, Pn, Sn), P-wave polarization classification, and seismic event type classification. The protocol tests generalization across temporal splits and transfer learning on local data to mitigate dataset imbalance. Use when the user wants to benchmark on CSNCD, or asks about evaluating this task. Reports recall.

researchpythongo
0
3
Pri Mo Mo Hpo EvalA

Evaluates multi-objective hyperparameter optimization algorithms on deep learning benchmarks, measuring their ability to find high-quality Pareto fronts of validation error and training cost under varying prior conditions and budget constraints. Use when the user wants to benchmark on Yahpo-Gym & PD1 HPO Benchmarks, or asks about evaluating this task. Reports mean dominated hypervolume.

researchpythongo
0
3
Preset Voice Matching EvalA

Evaluates a privacy-regulated speech-to-speech translation framework that replaces voice cloning with matching to pre-consented preset voices. Probes classifier robustness, cross-lingual speech naturalness, and inference efficiency across multilingual scenarios. Use when the user wants to benchmark on RAVDESS, CGDD, CAFE, EmoDB, CREMA-D, or asks about evaluating this task. Reports NISQA.

researchpythongo
0
3
Preposition Sense Disambiguation EvalA

Evaluates a model's ability to classify the sense of a preposition in context. It tests cross-lingual context representation and semi-supervised learning for fine-grained lexical disambiguation. Use when the user wants to benchmark on Web-reviews corpus, SemEval corpus, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Prefeval EvalA

Measures an agent's capability to retain and adhere to user preferences during long, multi-turn conversations. It tests whether the model can maintain consistency without external reminders or with explicit preference cues. Use when the user wants to benchmark on PrefEval, or asks about evaluating this task. Reports preference retention accuracy.

researchpython
0
3
Preference Discerning EvalA

This benchmark evaluates a model's ability to dynamically adapt to evolving user preferences by conditioning on natural language preferences inferred from interaction history. It probes recommendation accuracy, fine- and coarse-grained preference steering, sentiment following, and history consolidation across multiple e-commerce and gaming datasets. Use when the user wants to benchmark on Amazon Beauty, Amazon Sports and Outdoors, Amazon Toys and Games, Steam, or asks about evaluating this ta...

researchpythongo
0
3
Precision Recall F1 TA

Evaluates a model's ability to identify and rank key moments (shots) in soccer match videos for summarization. It measures how well the model selects representative content when constrained to match the exact duration of a human-curated highlight summary. Use when the user has predictions and gold and needs to compute F1 Score@$T$.

researchpythongo
0
3
Pre Training Validation LossA

Evaluates the generalization capability of a language model during the pre-training phase by measuring the average cross-entropy loss on a held-out validation corpus. Lower values indicate that the model has better learned the underlying token distribution and converges more effectively under the given architectural and training configurations. Use when the user has predictions and gold and needs to compute pre-training validation loss.

researchpythongo
0
3
Pralekha EvalA

Evaluates cross-lingual document alignment (CLDA) techniques by measuring chunk/sentence-level alignment accuracy (intrinsic) and the resulting document-level machine translation quality (extrinsic) across English and 11 Indic languages. Use when the user wants to benchmark on Pralekha, or asks about evaluating this task. Reports F1 Score, DocCOMET.

researchpythongo
0
3
Pqa Biochem Lite EvalA

Evaluates a model's ability to answer free-form scientific questions about unseen protein sequences using zero-shot multimodal reasoning. It probes biochemical property extraction, functional annotation, and cross-modal alignment between protein embeddings and natural language. Use when the user wants to benchmark on Pika-DS, or asks about evaluating this task. Reports mw MALE.

researchpythongo
0
3
Ppo Rl Benchmark EvalA

Evaluates reinforcement learning algorithms on continuous control and pixel-based Atari tasks to measure sample efficiency, stability, and final performance. It probes the ability of policy optimization methods to learn effective control policies across diverse physics simulators and arcade games. Use when the user wants to benchmark on OpenAI Gym (MuJoCo), Roboschool, Arcade Learning Environment, or asks about evaluating this task. Reports average total reward of the last 100 episodes.

researchpythongo
0
3
Ppg To Ecg Translation EvalA

Evaluates the fidelity of synthesizing ECG signals from PPG inputs and measures the downstream utility of the generated signals for cardiac and physiological task analysis. Use when the user wants to benchmark on WESAD, CAPNO, DALIA, BIDMC, MIMIC, PPG-BP, Cuffless-BP, or asks about evaluating this task. Reports RMSE.

researchpython
0
3
Ppg Health Benchmark EvalA

Evaluates the transferability and emergent capabilities of generalist and specialist time-series foundation models on diverse physiological tasks using PPG and cross-modal signals. Probes classification (e.g., arrhythmia, mental load, activity recognition) and regression (e.g., vital signs, blood chemistry, blood pressure) performance across clinical and ambulatory settings. Use when the user wants to benchmark on Stanford AF, Simband, Real World PPG, MIMIC-III, Sleep-EDF, or asks about evalu...

researchpythongo
0
3
Ppb Affinity EvalA

Evaluates protein language model architectures for predicting binding affinity in multi-chain protein-protein complexes. It probes how well different architectural designs capture inter-chain interactions compared to simple sequence or embedding concatenation. Use when the user wants to benchmark on PPB-Affinity, or asks about evaluating this task. Reports Spearman ρ.

researchpythongit
0
3
Powrl EvalA

Evaluates reinforcement learning agents for real-time power grid topology control under adversarial attacks and dynamic loads. It probes the agent's ability to maintain grid stability, minimize operational costs, and avoid blackouts across multiple challenging scenarios. Use when the user wants to benchmark on L2RPN NeurIPS 2020 (Robustness track) Offline, L2RPN NeurIPS 2020 (Robustness track) Online, L2RPN WCCI 2020 Offline, or asks about evaluating this task. Reports survival steps, scenari...

researchpythonperformance
0
3
Power System Forecasting EvalA

Evaluates the zero-shot and fine-tuning performance of time-series foundation models and deep learning baselines on deterministic and probabilistic power system forecasting tasks. It probes capabilities including horizon sensitivity, multivariate covariate handling, and generalization to unseen geographic sites. Use when the user wants to benchmark on ARPA-E PERFORM, or asks about evaluating this task. Reports nMAE.

researchpythonexpress
0
3
Potemkin EvalA

Evaluates the adversarial robustness of tool-using agentic AI against two orthogonal attack surfaces: breadth attacks that poison retrieval results to induce epistemic drift, and depth attacks that inject structural traps into information graphs to cause navigational collapse. It also probes agent susceptibility to linguistic credibility cues and hedging. Use when the user wants to benchmark on Potemkin-S2, Potemkin-Phantoms, Potemkin-Claims, or asks about evaluating this task. Reports DR.

researchpythongo
0
3
Postersum EvalA

Evaluates multimodal large language models' ability to generate accurate, abstractive summaries from complex, visually dense scientific posters. It probes layout understanding, visual-textual integration, and hierarchical summarization capabilities by measuring how well models extract and synthesize information from combined image and text inputs. Use when the user wants to benchmark on PosterSum, or asks about evaluating this task. Reports ROUGE-L.

researchpythongit
0
3
Postercraft Text EvalA

Evaluates the ability of text-to-image models to accurately render specified textual elements within aesthetically designed posters. It measures how well generated images preserve the exact characters, words, and layout instructions from the input prompt. Use when the user wants to benchmark on PosterCraft Test Prompts, or asks about evaluating this task. Reports Text F-score.

researchpythongit
0
3
Possible Stories Ifsm EvalA

Assesses whether large language models can generate story endings that align with free-form instructions provided alongside a narrative context. It measures both instruction-following accuracy and the model's ability to produce distinct endings for different instructions. Use when the user wants to benchmark on Possible Stories, or asks about evaluating this task. Reports IFSM.

researchpythongit
0
3