Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,853
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,897–7,920 of 20,853 skills

Dta Affinity Prediction EvalA

Evaluates a model's ability to predict the binding affinity between small molecule drugs and protein targets. It probes regression accuracy, ranking consistency, and correlation strength on standardized drug-target interaction datasets. Use when the user wants to benchmark on Davis, KIBA, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3
Dt Pens EvalA

This benchmark evaluates a model's ability to generate personalized news headlines by accurately capturing user interests from implicit feedback (clicks and dwell times) while filtering out noise. It probes the system's capacity to align generated text with both lexical patterns and semantic meaning relative to ground-truth headlines tailored to specific user preferences. Use when the user wants to benchmark on DT-PENS, or asks about evaluating this task. Reports ROUGE-1.

researchpythontesting
0
3
Dstc11 Track3 EvalA

Evaluates a system's ability to track dialogue state in spoken conversations, specifically measuring robustness to ASR errors, disfluencies, and proper noun mismatches. Use when the user wants to benchmark on DSTC11 Track 3, or asks about evaluating this task. Reports JGA.

researchpythongo
0
3
Dstc11 Track2 Intent Induction EvalA

Evaluates a model's ability to automatically induce conversation intents by clustering utterances from task-oriented dialogues without prior intent labels. It measures how well the induced clusters align with ground-truth intent categories using supervised clustering metrics. Use when the user wants to benchmark on DSTC11 Track 2, or asks about evaluating this task. Reports ACC.

researchpythongo
0
3
Dstc10 Spoken EvalA

Evaluates task-oriented dialogue systems on spoken conversations to measure robustness against ASR errors and disfluencies. It probes multi-domain dialogue state tracking, knowledge-seeking turn detection, knowledge selection, and response generation capabilities under realistic speech conditions. Use when the user wants to benchmark on DSTC10, DSTC9, MultiWOZ 2.1, or asks about evaluating this task. Reports Joint Goal Accuracy.

researchpythongo
0
3
Dst Jga EvalA

Evaluates a model's ability to track dialogue state in task-oriented conversations by predicting slot-value pairs across multiple domains. It measures how well the model maintains accurate belief states over multi-turn interactions, both with its own previous predictions and with ground-truth history. Use when the user wants to benchmark on RiSAWOZ, MultiWOZ, CrossWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).

researchpythongo
0
3
Dst EvalA

Evaluates cross-lingual and zero-shot dialogue state tracking by measuring a model's ability to predict correct slot-value pairs in target languages using limited or translated training data. Use when the user wants to benchmark on Parallel MultiWoZ, Multilingual WoZ, or asks about evaluating this task. Reports Joint Goal Accuracy.

researchpythongo
0
3
Dsp EvalA

Evaluates retrieval-augmented language models on open-domain, multi-hop, and conversational question answering by testing their ability to dynamically search for evidence, bootstrap in-context demonstrations, and generate accurate answers without fine-tuning. Use when the user wants to benchmark on Open-SQuAD, HotPotQA, QReCC, or asks about evaluating this task. Reports EM.

researchpythongo
0
3
Dsin Ctr EvalA

This benchmark evaluates a model's ability to predict click-through rates (CTR) by leveraging session-aware user behavior sequences. It probes how well a system can decompose historical interactions into time-separated sessions, model cross-session interest evolution, and adaptively weight session interests relative to a target item. Use when the user wants to benchmark on Advertising Dataset, Recommender Dataset, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Dsf Gan Utility EvalA

Evaluates the predictive utility of synthetic tabular data generated by a GAN. It measures how well a downstream classifier or regressor trained on the synthetic samples performs when evaluated on a strictly held-out real validation set. Use when the user wants to benchmark on Two distinct tabular datasets (names in Appendix A), or asks about evaluating this task. Reports model performance.

researchpythonperformance
0
3
Dsd Scene Analysis EvalA

Evaluates the ability of vision-language models to generate detailed, technically accurate scene descriptions from images, leveraging high-fidelity human annotations and peer-ranked photography data. Use when the user wants to benchmark on DataSeeds.AI Sample Dataset (DSD), or asks about evaluating this task. Reports BLEU-4.

researchpythonaws
0
3
Dsbench EvalA

Evaluates Vision-Language Models' ability to perceive and reason about safety-critical scenarios in autonomous driving, covering both external environmental hazards (e.g., traffic rules, obstacles, weather) and in-cabin driver states (e.g., fatigue, distraction, emotion). It probes fine-grained hazard recognition, regulatory compliance, and multi-step safety reasoning under diverse, high-risk conditions. Use when the user wants to benchmark on DSBench, or asks about evaluating this task. Repo...

researchpythongo
0
3
Ds 1000 EvalA

This benchmark evaluates a model's ability to generate correct, executable Python code for data science tasks, specifically focusing on NumPy operations. It probes functional correctness under natural language descriptions and tests robustness against surface-form and semantic perturbations of original StackOverflow problems. Use when the user wants to benchmark on numpy-100, or asks about evaluating this task. Reports pass@1.

researchpythonapi
0
3
Drvoice EvalA

Evaluates a speech-text voice conversation model's capabilities in speech-to-text understanding, speech-to-speech generation, and overall speech quality. It probes modality alignment, reasoning, open-ended QA, and instruction following across multiple audio benchmarks. Use when the user wants to benchmark on OpenAudioBench, VoiceBench, UltraEval-Audio, Big Bench Audio, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Drugplayground EvalA

Evaluates LLMs' ability to generate accurate, chemically plausible drug property descriptions and to produce meaningful text embeddings for drug discovery. It probes descriptive accuracy, lexical/structural alignment with ground truth, and embedding similarity for downstream representation tasks. Use when the user wants to benchmark on MolTextNet, or asks about evaluating this task. Reports Normalized Total score.

researchpythongit
0
3
Drugpc EvalA

Probes multi-step therapeutic reasoning and tool-use for drug-related questions, including interactions, contraindications, and patient-specific treatment strategies. It tests the model's ability to dynamically select biomedical tools, retrieve verified knowledge, and generate evidence-grounded answers. Use when the user wants to benchmark on DrugPC, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Drugood EvalA

Evaluates the out-of-distribution (OOD) generalization and robustness of graph neural networks and sequence models on molecular binding affinity prediction tasks under various domain shifts and annotation noise levels. Use when the user wants to benchmark on DrugOOD, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Drugcareqa EvalA

Evaluates an AI system's ability to perform integrated clinical decision-making by simulating real-world online medical consultations. It probes the model's capacity to reason through patient symptoms, generate accurate diagnoses, and recommend appropriate medications within a unified workflow. Use when the user wants to benchmark on DrugCareQA, or asks about evaluating this task. Reports diagnostic and medication recommendation accuracy.

researchpythongo
0
3
Drugbank Hetionet EvalA

Evaluates the ability of matrix completion algorithms to predict missing biological interactions (drug-target or compound-disease) using sparse association matrices and side information. Use when the user wants to benchmark on DrugBank, Hetionet (Drug Repurposing), or asks about evaluating this task. Reports AUPR.

researchpythongo
0
3
Drug Target Interaction EvalA

Evaluates computational models on predicting binary drug-target interactions using standardized bioactivity data. It probes the model's ability to learn molecular and protein representations and generalize across different data splits (lenient, cold-ligand, cold-target). Use when the user wants to benchmark on Curated DTI dataset, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Drug Pair Scoring EvalA

This evaluation benchmarks deep learning architectures on predicting drug-drug interactions, polypharmacy side effects, and drug synergy. It measures how well models encode molecular graphs and combine them to score pairwise biological outcomes across multiple pharmacological domains. Use when the user wants to benchmark on TWOSIDES, Drugbank DDI, DrugComb, DrugCombDB, OncolyPharm, or asks about evaluating this task. Reports AUROC.

researchpythonperformance
0
3
Drug Discovery Benchmarks EvalA

Evaluates a multi-modal foundation model's capability across classification, regression, and generation tasks in drug discovery. It probes the model's ability to predict cell types, assess drug efficacy and safety, design antibody CDR regions, and estimate binding affinities for proteins and small molecules. Use when the user wants to benchmark on Zheng68k, MoleculeNet (BBBP/ClinTox), GDSC (Cancer-Drug Response 1-3), SAbDab, Weber TCR Benchmark, SKEMPI S1131, DTI Benchmark, or asks about eval...

researchpythongo
0
3
Drsm Certified Robustness EvalA

Evaluates the standard classification accuracy and certified robustness of a malware detector against adversarial byte perturbations. It measures how well the model maintains correct predictions under a bounded perturbation budget using a de-randomized smoothing defense with window ablation. Use when the user wants to benchmark on PACE, or asks about evaluating this task. Reports Standard Accuracy.

researchpythongit
0
3
Droughted EvalA

Evaluates time-series forecasting models on predicting U.S. drought severity across 1 to 6 week horizons using meteorological and static features. It probes both regression accuracy and multi-class classification performance for drought monitoring levels. Use when the user wants to benchmark on DroughtED, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3