Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,861
skills in category
870
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,065–8,088 of 20,861 skills

Dermabench EvalA

Evaluates vision-language models on dermatological visual question answering and clinical reasoning. It probes the model's ability to understand skin lesions across diverse Fitzpatrick skin types, answer structured diagnostic questions, and reason about morphology and distribution. Use when the user wants to benchmark on DermaBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Depthcues EvalA

Evaluates whether large vision models inherently understand human monocular depth cues (e.g., occlusion, perspective, texture gradient) through classification tasks, and measures their downstream monocular depth estimation performance on standard datasets. Use when the user wants to benchmark on DepthCues, NYUv2, DIW, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Depth EvalA

Evaluates the effectiveness of a hierarchically pre-trained encoder-decoder model (DEPTH) against a standard T5 baseline on discourse understanding, natural language inference, sentiment analysis, grammar checking, and instruction following. Use when the user wants to benchmark on MNLI, SST2, CoLA, DiscoEval, Natural Instructions, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Depth Anything Ac EvalA

Evaluates zero-shot monocular relative depth estimation robustness under complex environmental conditions such as low light, adverse weather (rain, fog, snow), and synthetic noise. It probes the model's ability to recover fine-grained spatial relationships and object boundaries from degraded inputs without fine-tuning. Use when the user wants to benchmark on DA-2K (multi-condition), NuScenes-night, Robotcar-night, Driving-Stereo, KITTI-C, KITTI, NYU-D, Sintel, ETH3D, DIODE, or asks about eval...

researchpython
0
3
Depression Diagnosis Chat EvalA

Evaluates a model's ability to conduct depression-diagnosis-oriented dialogues by tracking psychological states, generating appropriate responses, summarizing patient symptoms, and classifying depression/suicide severity. It also assesses conversational qualities like fluency, empathy, and doctor-likeness through human evaluation. Use when the user wants to benchmark on MedDialog, or asks about evaluating this task. Reports BLEU-2, Average weighted F1.

researchpythongo
0
3
Dental Triagebench EvalA

Evaluates multimodal clinical reasoning by requiring models to integrate radiographic images (OPGs) and patient complaints to predict hierarchical dental triage labels. It probes the model's ability to perform precise, multi-label treatment referrals and broad specialty-level routing in a zero-shot clinical setting. Use when the user wants to benchmark on Dental-TriageBench, or asks about evaluating this task. Reports Macro-F1.

researchpythongo
0
3
Densepose Coco EvalA

Evaluates a model's ability to perform dense human pose estimation by predicting per-pixel body part labels and UV coordinates on a 3D surface model. It measures how well the model handles real-world variations in scale, pose, occlusion, and background clutter. Use when the user wants to benchmark on COCO-DensePose, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
Densemarks EvalA

Evaluates a model's ability to learn dense, pose-robust 3D canonical embeddings for human head images. It probes geometric fidelity in point matching, semantic consistency across identities, and robustness to occlusions and extreme poses. Use when the user wants to benchmark on CelebV-HQ, Nersemble, or asks about evaluating this task. Reports MAE.

researchpython
0
3
Demon EvalA

Evaluates a model's ability to comprehend and follow complex, interleaved multimodal instructions that require inferring missing visual details and reasoning across multiple images and text turns. It probes reasoning-aware detail comprehension, image-text alignment, and sensitivity to visual context order. Use when the user wants to benchmark on DEMON, MME, OwlEval, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Delucionqa EvalA

This benchmark evaluates a model's ability to detect hallucinations in domain-specific question answering systems that use retrieval-augmented generation. It probes whether models can correctly identify when a generated answer contradicts or goes beyond the provided retrieved context, often due to over-reliance on pre-trained knowledge or incomplete retrieval. Use when the user wants to benchmark on DelucionQA, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Delta Chi2A

Evaluates the detectability and orbital parameter recovery of long-period exoplanets using simulated Gaia astrometric observations. It quantifies detection significance by comparing the goodness-of-fit between a standard astrometric model and a full orbital model. Use when the user has predictions and gold and needs to compute Δχ².

researchpythongo
0
3
Deferredseg EvalA

Evaluates a pixel-wise deferral framework for medical image segmentation that dynamically routes uncertain pixels to synthetic or real experts. It measures how well the collaborative system improves segmentation accuracy over strong baselines like MedSAM across diverse organs and imaging modalities. Use when the user wants to benchmark on PROMISE12, LiTS, AMOS22, Chaksu, or asks about evaluating this task. Reports DSC.

researchpythonrust
0
3
Defect Spectrum EvalA

Evaluates industrial defect segmentation models by measuring their ability to accurately localize and classify multiple defect types within complex manufacturing images. It also assesses how well models trained on refined, granular annotations generalize compared to those trained on coarse original annotations, using both pixel-level and image-level quality control metrics. Use when the user wants to benchmark on Defect Spectrum, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Defame Fact Checking EvalA

This evaluation probes a model's ability to dynamically retrieve and reason over multimodal evidence to verify open-domain claims. It tests zero-shot multimodal fact-checking across text-only, text-image, and out-of-context scenarios, measuring both classification accuracy and the quality of generated justifications. Use when the user wants to benchmark on AVeriTeC, MOCHEG, VERITE, ClaimReview2024+, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Deepvision 103k EvalA

Evaluates the multimodal mathematical reasoning and general multimodal reasoning capabilities of vision-language models. It probes visual perception, step-by-step logical deduction, and cross-domain generalization on K12-level math and broader visual tasks. Use when the user wants to benchmark on Multimodal Math & General Reasoning Benchmarks (WeMath, MathVerse_vision, MathVision, LogicVista, MMMU_VAL, MMMU_Pro_full, M^3CoT), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Deepurban Trajectory EvalA

Evaluates trajectory prediction and planning capabilities in high-density urban environments with significant vehicle-to-vulnerable-road-user interactions. It measures prediction accuracy and safety compliance using displacement errors and collision scores. Use when the user wants to benchmark on DeepUrban, or asks about evaluating this task. Reports ADE.

researchpythongo
0
3
Deeptoponet EvalA

Evaluates a deep learning model's ability to reconstruct subglacial bed topography by fusing sparse radar ice thickness measurements with high-resolution surface elevation and ice dynamics data. It probes spatial interpolation accuracy, structural preservation of terrain features, and robustness in data-scarce glacial regions. Use when the user wants to benchmark on Upernavik Isstrøm, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3
Deeptilebars EvalA

Evaluates neural information retrieval models on ad-hoc web search and benchmark datasets by measuring how well they rank relevant documents using segment-level matching and discourse structure modeling. Use when the user wants to benchmark on TREC 2010-2012 Web Track, LETOR 4.0 MQ2008, or asks about evaluating this task. Reports nDCG@20.

researchpython
0
3
Deepspeech Wer EvalA

This evaluation probes an end-to-end speech recognition system's ability to accurately transcribe conversational telephone speech and robustly handle background noise without phoneme-level modeling or explicit speaker adaptation. It measures transcription accuracy against ground truth references using standard error rates. Use when the user wants to benchmark on Switchboard Hub5’00 (LDC2002S23), Custom Noisy Speech Test Set, or asks about evaluating this task. Reports word error rate (WER).

researchpythonapi
0
3
Deepscientist EvalA

Evaluates an autonomous AI research system's ability to progressively advance state-of-the-art methods across three distinct AI tasks: agent failure attribution, LLM inference acceleration, and AI text detection. It also assesses the scientific quality of the AI-generated research papers through automated and human peer review. Use when the user wants to benchmark on Who&When benchmark, MBPP, AI Text Detection dataset, or asks about evaluating this task. Reports Accuracy, AUROC.

researchpythongo
0
3
Deeprecsys EvalA

Evaluates the throughput, tail latency, and power efficiency of a dynamic scheduling system for at-scale neural recommendation inference across various industry models, hardware platforms, and tail-latency constraints. Use when the user wants to benchmark on Industry-representative recommendation models (DLRM-RMC1/2/3, WND, MT-WND, NCF, DIN, DIEN), or asks about evaluating this task. Reports QPS.

researchpythongo
0
3
Deepprotein Benchmark EvalA

Evaluates deep learning models on a comprehensive suite of protein sequence learning tasks, including function prediction, subcellular localization, protein-protein interaction, epitope/paratope prediction, antibody developability, CRISPR repair outcomes, and protein structure prediction. Use when the user wants to benchmark on Fluorescence, Stability, β-lactamase, Solubility, Subcellular, Binary, PPI Affinity, Yeast, Human PPI, IEDB, PDB-Jespersen, SAbDab-Liberis, TAP, SAbDab-Chen, CRISPR-Le...

researchpythongo
0
3
Deepmind Control Suite EvalA

Evaluates continuous control reinforcement learning agents on a standardized suite of physics-based simulation tasks. It probes sample efficiency, stability, and performance over long training horizons using uniform action, observation, and reward structures. Use when the user wants to benchmark on DeepMind Control Suite, or asks about evaluating this task. Reports return.

researchpythonperformance
0
3
Deepmath 103k EvalA

Evaluates mathematical reasoning capabilities on a curated, decontaminated dataset of challenging problems, measuring performance across standardized math competitions and academic benchmarks. Use when the user wants to benchmark on DeepMath-103K, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3