Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,753
skills in category
990
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,097–6,120 of 23,753 skills

Navsim Epdms EvalA

Evaluates closed-loop end-to-end autonomous driving performance in safety-critical and diverse real-world scenarios. It measures collision avoidance, rule compliance, progress, and comfort under reactive traffic conditions. Use when the user wants to benchmark on navhard, navtest, or asks about evaluating this task. Reports EPDMS.

researchpythongo
0
3
Nautilus Voice Cloning EvalA

Evaluates a voice cloning system's ability to generate high-quality, speaker-similar speech using minimal untranscribed or transcribed target speech. It probes both text-to-speech (TTS) and voice conversion (VC) capabilities, focusing on naturalness, speaker similarity, and accent preservation across native and non-native speakers. Use when the user wants to benchmark on VCC2018 SPOKE task, VCTK & EMIME, or asks about evaluating this task. Reports MOS.

researchpythonperformance
0
3
Nautilus EvalA

Evaluates large multimodal models on underwater scene understanding across eight tasks, including coarse/fine classification, image/region captioning, grounding, detection, VQA, and object counting. It probes the model's robustness to severe underwater image degradation (light scattering, absorption, color casts) and its ability to generalize to unseen underwater domains. Use when the user wants to benchmark on NautData, IOCfish5k, MarineInst20M, or asks about evaluating this task. Reports ac...

researchpythongo
0
3
Naturalvoices Vc EvalA

Evaluates the ability of voice conversion models to preserve speaker identity, intelligibility, and emotional expression when converting spontaneous, in-the-wild podcast speech. It benchmarks both standard and emotion-aware conversion across multiple architectures and data scales. Use when the user wants to benchmark on NaturalVoices, ESD, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Naturalspeech EvalA

Evaluates the perceptual quality and generation speed of an end-to-end text-to-speech system. It measures how closely synthesized speech matches human recordings and outperforms prior cascaded or flow-based TTS baselines. Use when the user wants to benchmark on LJSpeech, or asks about evaluating this task. Reports CMOS.

researchpython
0
3
Naturalreasoning EvalA

Evaluates the zero-shot reasoning capabilities of models trained via knowledge distillation or self-training on the NaturalReasoning dataset. It measures performance across diverse mathematics and science benchmarks to assess scaling efficiency and generalization. Use when the user wants to benchmark on MATH, GPQA, GPQA-Diamond, MMLU-Pro, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Natural Instructions EvalA

Evaluates a model's ability to generalize to unseen NLP tasks by leveraging crowdsourced natural language instructions alongside training data. It measures how well instruction-based learning transfers across different task categories, datasets, and individual tasks compared to data-only training. Use when the user wants to benchmark on Natural Instructions, or asks about evaluating this task. Reports ROUGE-L.

researchpythongo
0
3
Nash Pruning EvalA

Evaluates the generation quality and inference efficiency of structuredly pruned encoder-decoder language models across abstractive QA, summarization, classification, and instruction-following tasks. Use when the user wants to benchmark on TweetQA, XSum, SAMSum, CNN/DailyMail, GLUE/SuperGLUE (RTE, BoolQ, CB), Databricks-dolly-15k, Self-Instruct, Vicuna Evaluation, or asks about evaluating this task. Reports ROUGE-L.

researchpythongo
0
3
Nasadat Covid Severity EvalA

Evaluates the ability of deep learning and geometric deep learning models to forecast county-level COVID-19 hospitalizations using satellite-derived atmospheric variables (AOD, temperature, humidity) alongside baseline features. It probes spatio-temporal forecasting capabilities and the conditional predictive utility of environmental risk factors on disease severity. Use when the user wants to benchmark on NASAdat, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Narrativeqa EvalA

Evaluates English long-document retrieval on complex, narrative-style questions, probing deep comprehension and information extraction from lengthy texts. Use when the user wants to benchmark on NarrativeQA, or asks about evaluating this task. Reports nDCG@10.

researchpythonperformance
0
3
Narrasum EvalA

This benchmark evaluates a model's ability to perform abstractive and extractive summarization on long-form narrative texts (movie/TV plot descriptions). It probes the model's capacity to capture event causality, character motivations, temporal dynamics, and overall narrative coherence while maintaining faithfulness to the source document. Use when the user wants to benchmark on NarraSum, or asks about evaluating this task. Reports ROUGE F1.

researchpythongo
0
3
Narm Session Rec EvalA

Evaluates session-based recommendation models by predicting the next item a user will click based on their sequential interaction history within a session. It probes the model's ability to capture both sequential behavior and session-level intent/purpose. Use when the user wants to benchmark on YOOCHOOSE 1/64, YOOCHOOSE 1/4, DIGINETICA, or asks about evaluating this task. Reports Recall@20.

researchpythongo
0
3
Nanoknow EvalA

Evaluates how pre-training data exposure and external context influence closed-book and open-book question answering accuracy. It probes the model's reliance on parametric knowledge versus retrieved evidence, and measures the impact of answer frequency and distractors. Use when the user wants to benchmark on Natural Questions, SQuAD, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Naijas2st EvalA

Evaluates speech-to-text and speech-to-speech translation capabilities across low-resource Nigerian languages (Hausa, Igbo, Yorùbá, Nigerian Pidgin) and English. It specifically probes how well cascaded, end-to-end, and AudioLLM architectures handle multi-accent variations and bidirectional translation directions. Use when the user wants to benchmark on NaijaS2ST, or asks about evaluating this task. Reports SSA-COMET.

researchpython
0
3
Nacsp EvalA

Evaluates a model's ability to predict discrete neural audio codec parameters (quantizers, sampling rate, bits per second) from audio samples, enabling fine-grained source attribution of AI-generated speech. The protocol frames open-set attribution as a multi-task regression problem rather than binary classification, requiring the model to generalize across both seen and unseen codec configurations. Use when the user wants to benchmark on ST-Codecfake, CodecFake, or asks about evaluating this...

researchpythongo
0
3
Nab Yahoo Anomaly Detection EvalA

Evaluates unsupervised anomaly detection algorithms on streaming time-series data. It probes the model's ability to identify point, contextual, and collective anomalies in highly imbalanced real-world and synthetic datasets without labeled training data. Use when the user wants to benchmark on Numenta Anomaly Benchmark, Yahoo Anomaly Dataset, or asks about evaluating this task. Reports F-measure.

researchpythongo
0
3
Nab EvalA

Evaluates real-time anomaly detection algorithms on streaming time-series data, measuring their ability to detect natural and synthetic anomalies while penalizing false alarms and delayed detections. Use when the user wants to benchmark on NAB 1.0, or asks about evaluating this task. Reports NAB Score.

researchpythongo
0
3
Naamapadam EvalA

Evaluates Named Entity Recognition (NER) capabilities across 11 Indic languages. It probes a model's ability to identify and classify PERSON, LOCATION, and ORGANIZATION entities in low-resource and multilingual settings using projection-based and fine-tuned approaches. Use when the user wants to benchmark on Naamapadam, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
N2c2 Concept Relation EvalA

Evaluates clinical NLP models on extracting medical concepts and their relations from clinical notes. It probes the model's ability to handle nested/overlapped concepts and assesses cross-institutional generalization across different benchmark years. Use when the user wants to benchmark on n2c2 2018, n2c2 2022, n2c2 cross-institution (MIMIC-train/UW-test), or asks about evaluating this task. Reports strict micro-averaged F1-score.

researchpythongo
0
3
Myriadal EvalA

Evaluates active few-shot learning frameworks for histopathology image classification under extremely tight annotation budgets (1, 5, and 10 labeled samples). It probes how well uncertainty-based diversity sampling and self-supervised contrastive pretraining can reduce sample redundancy and improve classification performance compared to standard few-shot and active learning baselines. Use when the user wants to benchmark on NCT-CRC-HE-100K, BreaKHis, or asks about evaluating this task. Report...

researchpythongo
0
3
Mxnet Framework Benchmark EvalA

Evaluates the raw execution speed, memory footprint, and distributed scalability of the MXNet deep learning framework against Torch7, Caffe, and TensorFlow. It measures how efficiently the library handles standard convolutional neural network architectures and large-scale image classification tasks across single and multiple GPU nodes. Use when the user wants to benchmark on convnet-benchmarks, ILSVRC12, or asks about evaluating this task. Reports forward-backward performance.

researchpythongo
0
3
Mwp Value Accuracy EvalA

Evaluates mathematical reasoning and robustness on single-equation math word problems. It probes a model's ability to parse linguistic variations, ignore irrelevant information, and solve inverted or structurally complex problems. Use when the user wants to benchmark on MAWPS, SVAMP, PARAMAWPS, or asks about evaluating this task. Reports Value accuracy.

researchpythongo
0
3
Mwp Localization EvalA

Evaluates LLMs' ability to solve math word problems after socio-cultural localization of entities into low-resource languages. Probes whether models maintain reasoning accuracy when cultural context shifts, focusing solely on final answer correctness rather than step-by-step reasoning. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Mwlp Storm Repair EvalA

Evaluates an algorithm's ability to optimally partition repair targets among multiple crews and route them to minimize total weighted latency (average wait time) while balancing workload distribution across crews in post-disaster urban scenarios. Use when the user wants to benchmark on Random Environments, Champaign Case Study, or asks about evaluating this task. Reports wait.

researchpythongo
0
3