Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,853
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,849–7,872 of 20,853 skills

E Care EvalA

Evaluates a model's ability to perform commonsense causal reasoning by predicting the reasonableness of causal facts, and to generate conceptually grounded natural language explanations for those causal relationships. Use when the user wants to benchmark on e-CARE, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Dzen EvalA

Evaluates foundation models' ability to answer multiple-choice academic questions in both English and Dzongkha across varying grade levels and scientific subjects. It specifically probes factual recall, procedural application, and multi-step reasoning capabilities in a low-resource multilingual setting. Use when the user wants to benchmark on DZEN, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dysql Bench EvalA

Evaluates a model's ability to perform dynamic, multi-turn Text-to-SQL interactions that support full CRUD operations. It probes contextual reasoning, adaptability to evolving user requests, and error recovery within stateful database dialogues. Use when the user wants to benchmark on DySQL-Bench, or asks about evaluating this task. Reports state-equivalence accuracy.

researchpythongo
0
3
Dysarthric Asr EvalA

Evaluates the ability of ASR and LLM-enhanced decoding models to accurately transcribe dysarthric speech across varying severity levels and domains. It probes robustness to phonetic distortions, grammatical consistency, and cross-dataset generalization. Use when the user wants to benchmark on TORGO, UASpeech, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Dynamic Urban View Synthesis EvalA

Evaluates novel view synthesis and image reconstruction quality in dynamic urban environments containing fast-moving objects and varying environmental conditions. It measures how well a model can render unseen viewpoints and reconstruct training views while handling dynamic geometry and pose drift. Use when the user wants to benchmark on Argoverse 2, KITTI, VKITTI2, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Dynamic Unlearning EvalA

Evaluates the effectiveness and robustness of LLM unlearning methods by measuring residual knowledge retrieval across dynamically generated single-hop, multi-hop, and alias-based queries, alongside the retention of adjacent and general knowledge. Use when the user wants to benchmark on RWKU, TOFU, or asks about evaluating this task. Reports Multi-hop Forgetting Criterion.

researchpythongo
0
3
Dynamic Topic Quality EvalA

Evaluates dynamic topic models by measuring topic coherence and diversity across chronological time slices, and assesses the utility of learned document-topic distributions via downstream text classification and clustering tasks. Use when the user wants to benchmark on NeurIPS, ACL, UN, NYT, WHO, or asks about evaluating this task. Reports Topic Coherence (TC).

researchpythongo
0
3
Dynamic Superb Phase2 EvalA

Evaluates instruction-based universal speech and audio models across 180 tasks spanning speech, music, and environmental audio. It probes capabilities like automatic speech recognition, emotion recognition, speaker verification, and audio classification using a unified instruction-following framework. Use when the user wants to benchmark on Dynamic-SUPERB Phase-2, or asks about evaluating this task. Reports relative_score.

researchpythongo
0
3
Dynamic Superb EvalA

Evaluates instruction-tuned speech models on their ability to perform diverse speech and audio tasks using natural language instructions. It probes zero-shot generalization by testing performance on seen versus unseen tasks and instructions across six dimensions: content, speaker, semantics, degradation, paralinguistics, and audio. Use when the user wants to benchmark on Dynamic-SUPERB, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dynamic Nerf Soccer EvalA

Evaluates the ability of dynamic NeRF models to perform photorealistic novel view synthesis in large-scale, dynamic sports environments. It probes spatiotemporal modeling capabilities, specifically how well models handle fast-moving small objects (like a soccer ball) and scale variations across different camera configurations. Use when the user wants to benchmark on Synthetic Soccer Scenes (Multi-Camera), or asks about evaluating this task. Reports PSNR.

researchpython
0
3
Dynamic House Simulator EvalA

Evaluates an agent's ability to perform temporal link prediction and object search in partially observable, dynamic environments by predicting object locations, ranking location likelihoods, and navigating to objects sequentially. Use when the user wants to benchmark on Dynamic House Simulator, or asks about evaluating this task. Reports NDCG.

researchpythongo
0
3
Dynamic Audio Visual Nav EvalA

Evaluates an embodied agent's ability to navigate towards and catch a moving, previously unheard sound source in unmapped 3D environments using only audio and visual observations. It probes spatial reasoning, temporal memory, and robustness to noisy or distractor audio scenarios. Use when the user wants to benchmark on Replica, Matterport3D, or asks about evaluating this task. Reports DSPL.

researchpythongo
0
3
Dynamath EvalA

Evaluates the robustness of Vision-Language Models in mathematical reasoning by measuring performance across dynamically generated variants of seed questions. It probes how well models handle numerical, geometric, and contextual perturbations while maintaining consistent logical deduction. Use when the user wants to benchmark on DynaMath, or asks about evaluating this task. Reports average-case accuracy.

researchpythongo
0
3
Dyad Arch EvalA

Evaluates a block-sparse linear layer approximation (DYAD) against dense baselines across standard NLP and vision benchmarks, measuring accuracy preservation and computational efficiency. Use when the user wants to benchmark on BLIMP, OPENLLM, GLUE+, MNIST, or asks about evaluating this task. Reports BLIMP accuracy.

researchpythongo
0
3
Dyabd Segmentation EvalA

This benchmark evaluates the segmentation capabilities of deep learning models on dynamic abdominal MRI scans. It specifically probes how well models handle extreme anatomical variability caused by real-time muscle motion during breathing and Valsalva maneuvers, across few-shot, prompt-based, and fully automatic inference settings. Use when the user wants to benchmark on DyABD, or asks about evaluating this task. Reports Dice.

researchpythongit
0
3
Dy Meter EvalA

Evaluates online anomaly detection models under concept drift by testing their ability to adapt to evolving data distributions without retraining. It probes instance-level sensitivity to context-dependent anomalies across continuous and discrete streaming scenarios. Use when the user wants to benchmark on Ionosphere, Pima, Satellite, Mammography, BGL, NSL-KDD, KDD99, Activity Recognition, Internal Bleeding, NASA, GaitPhase, EPG, ECG, Machine temperature, CPU utilization, INSECTS-Abr, INSECTS-...

researchpythongo
0
3
Dw Bench EvalA

Evaluates LLMs' ability to reason about data warehouse graph topologies, specifically focusing on foreign key path enumeration, data lineage impact analysis, and multi-hop graph traversal. It probes whether models can perform structural graph reasoning versus relying on lexical cues, using heterogeneous schema graphs with foreign key and lineage edges. Use when the user wants to benchmark on DW-Bench, or asks about evaluating this task. Reports Micro-EM.

researchpythongo
0
3
Dvqa EvalA

This benchmark evaluates a model's ability to perform visual reasoning and information extraction on bar chart data visualizations. It specifically probes whether systems can accurately read chart-specific labels, handle out-of-vocabulary terms, and answer natural language questions about quantitative relationships and chart structure. Use when the user wants to benchmark on DVQA, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Dvfs Latency Energy EvalA

Evaluates the accuracy of a data-driven DVFS-aware latency model for DNN inference on GPUs against a traditional FLOPs-based benchmark. It probes the model's ability to predict real-world inference time and energy consumption under varying frequency settings, deadlines, and cooperative offloading scenarios. Use when the user wants to benchmark on CIFAR10, or asks about evaluating this task. Reports inference time (ms).

researchpythonperformance
0
3
Dvd Dst EvalA

Evaluates a model's ability to track visual objects and their attributes across turns in video-grounded dialogues. It probes long-term cross-modal dependency resolution and precise state decoding under controlled, bias-free synthetic dialogue conditions. Use when the user wants to benchmark on DVD-DST, or asks about evaluating this task. Reports Joint Acc.

researchpythongo
0
3
Dvbench EvalA

Evaluates Vision Large Language Models' ability to understand safety-critical driving videos across a hierarchical taxonomy of 25 abilities, including perception, temporal-spatial reasoning, and risk assessment. Use when the user wants to benchmark on DVBench, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Dutch Financial Benchmark EvalA

Evaluates LLMs on domain-specific financial tasks in Dutch, including sentiment analysis, named entity recognition, relation extraction, query answering, and headline classification. It also tests cross-lingual adaptability by benchmarking the Dutch model on English financial data. Use when the user wants to benchmark on Dutch Financial Benchmark, English Financial Benchmark, or asks about evaluating this task. Reports zero-shot performance.

researchpythongo
0
3
Dutch Book Review Sentiment EvalA

Evaluates the effectiveness of Universal Language Model Fine-tuning (ULMFiT) versus traditional SVM classifiers on low-resource, domain-specific sentiment classification tasks using small training sets. Use when the user wants to benchmark on Dutch book reviews, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dutch Ade Corpus EvalA

This benchmark evaluates transformer and Bi-LSTM models for detecting adverse drug events (ADEs) in Dutch clinical free text. It probes named entity recognition for drugs and disorders, relation classification for ADE and prescribing indication pairs, and document-level ADE detection. The protocol emphasizes handling class imbalance and evaluating performance across strict/lenient entity matching and single vs. grouped ADE relations. Use when the user wants to benchmark on Dutch ADE corpus, I...

researchpythonperformance
0
3