Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,821
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,457–6,480 of 20,821 skills

Lingoqa EvalA

Evaluates vision-language models on autonomous driving video question answering, testing their ability to understand temporal visual context, describe scenes, predict actions, and justify answers based on driving scenarios. Use when the user wants to benchmark on LingoQA, or asks about evaluating this task. Reports Ling-Judge.

researchpythongo
0
3
Lingoly EvalA

This benchmark evaluates large language models' ability to perform multi-step linguistic reasoning and deductive puzzle solving in low-resource and extinct languages. It probes out-of-domain grammatical inference and instruction-following under conditions of minimal pre-training exposure, requiring models to extract and apply novel rules from provided context rather than relying on memorized knowledge. Use when the user wants to benchmark on LINGOLY, or asks about evaluating this task. Report...

researchpythongo
0
3
Lingo Space Grounding EvalA

Evaluates a model's ability to ground natural language spatial instructions to specific 2D pixel locations in RGB-D tabletop scenes. It probes both single-relation grounding and incremental/compositional grounding where multiple spatial predicates must be satisfied sequentially or simultaneously. Use when the user wants to benchmark on CLIPort Benchmark, ParaGon Benchmark, SREM Benchmark, LINGO-Space Benchmark, Composite Instruction Task, or asks about evaluating this task. Reports success sc...

researchpythongo
0
3
Linglanmidian EvalA

Evaluates LLMs on Traditional Chinese Medicine (TCM) knowledge recall, multi-hop clinical reasoning, information extraction, and clinical decision-making. It probes synonym-tolerant clinical labeling, robustness on curated hard subsets, and performance across diverse TCM-specific task formats including QA, NER, and dosage prediction. Use when the user wants to benchmark on LingLanMiDian, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Limitgen EvalA

Evaluates whether LLMs can accurately identify and articulate critical limitations in scientific research papers across methodological, experimental, analytical, and literature-related dimensions. The benchmark probes the model's ability to ground critiques in domain-specific best practices and produce actionable, substantive feedback rather than superficial presentation critiques. Use when the user wants to benchmark on LimitGen, or asks about evaluating this task. Reports Limitation Quality.

researchpythongo
0
3
Lime Mmt47 EvalA

Evaluates the ability of lightweight Mixture of Experts (MoE) parameter-efficient fine-tuning methods to generalize across diverse multimodal tasks. It probes how well shared PEFT modules with expert modulation vectors capture task-specific specialization without learned routing parameters. Use when the user wants to benchmark on MMT-47, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Lightweight Action Recognition EvalA

Evaluates the real-world efficiency (training and inference latency, VRAM footprint) of video action recognition models across desktop GPUs and mobile devices, alongside their classification accuracy on standard benchmarks. Use when the user wants to benchmark on EK100, SSV2, K400, or asks about evaluating this task. Reports relative latency.

researchpython
0
3
Lightgcn Rec EvalA

Evaluates the ability of graph-based collaborative filtering models to rank relevant items for users based on sparse user-item interaction graphs. It probes how well neighborhood aggregation and embedding smoothing capture latent preferences without relying on node semantic features. Use when the user wants to benchmark on Gowalla, Yelp2018, Amazon-Book, or asks about evaluating this task. Reports recall@20.

researchpythongo
0
3
Lifelong Rl EvalA

Evaluates the ability of reinforcement learning agents to sequentially learn multiple tasks while retaining prior knowledge, generalizing to unseen environments, and leveraging forward transfer from previous tasks. It probes parameter isolation, knowledge composition, and robustness across discrete and continuous action spaces with varying reward and input distributions. Use when the user wants to benchmark on ProcGen, CT-graph, Minigrid, Continual World, or asks about evaluating this task. R...

researchpythonperformance
0
3
Libritts Tts EvalA

Evaluates the naturalness and quality of synthesized speech from text-to-speech models trained on the LibriTTS corpus. It probes how audio sampling rate, text normalization, and sentence-level splitting affect human-perceived speech naturalness compared to the original LibriSpeech dataset. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Libritts Ssd EvalA

Evaluates the zero-shot speaker adaptation capability of an autoregressive speech synthesis model on unseen speakers. It measures content accuracy, voice cloning fidelity, and audio quality, while also quantifying inference speedup over standard autoregressive decoding. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Libritts Selfvc Watermark EvalA

Evaluates the robustness of neural audio watermarking systems against self voice conversion attacks and transmission channel distortions, while measuring speaker identity preservation, linguistic content integrity, and perceptual quality. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports bitwise extraction accuracy.

researchpython
0
3
Libritts R EvalA

This protocol evaluates the audio quality and naturalness of a restored multi-speaker TTS corpus (LibriTTS-R) compared to the original LibriTTS dataset. It measures both ground-truth speech fidelity and the downstream impact on multi-speaker TTS model generation quality using human subjective listening tests. Use when the user wants to benchmark on LibriTTS-R, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Librispeech Wer EvalA

Evaluates the ability of a semantic-aware speech-to-text transmission system to accurately reconstruct text from speech signals under noisy communication channels (AWGN and Rayleigh). It probes semantic feature extraction, redundancy removal, and robustness to channel noise. Use when the user wants to benchmark on Librispeech, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Librispeech Voice Search Wer EvalA

Evaluates the word error rate of streaming speech recognition systems using a first-pass RNN-T model followed by a second-pass rescorer. It probes the ability of parallel Transformer rescoring to improve transcription accuracy while maintaining low-latency streaming constraints on-device. Use when the user wants to benchmark on Librispeech, Google Voice Search, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Librispeech Sr Semantic Comm EvalA

Evaluates the ability of semantic communication systems to transmit speech spectra over noisy wireless channels (AWGN and Rayleigh) and accurately recover text transcriptions, comparing performance against traditional speech and text transceivers. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports Character Error Rate (CER).

researchpythontesting
0
3
Librispeech Pc EvalA

Evaluates the ability of end-to-end automatic speech recognition (ASR) models to correctly predict punctuation marks and word capitalization in transcribed speech. It specifically isolates punctuation-specific errors to enable fine-grained comparison between cascade and end-to-end architectures. Use when the user wants to benchmark on LibriSpeech-PC, or asks about evaluating this task. Reports Punctuation Error Rate (PER).

researchpythongo
0
3
Librispeech EvalA

Evaluates speech recognition performance under varying amounts of labeled data (1h, 10h, 100h) and different model sizes. It probes the ability of self-supervised speech models to adapt to downstream transcription tasks with limited supervision. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythongo
0
3
Librispeech Edit EvalA

Evaluates speech editing models on their ability to accurately modify target words while preserving speaker identity, acoustic quality, and temporal alignment of unedited regions. Use when the user wants to benchmark on LibriSpeech-Edit, or asks about evaluating this task. Reports WER.

researchpython
0
3
Librispeech Chime4 Asr EvalA

Evaluates the robustness of automatic speech recognition (ASR) models under various noise conditions and signal-to-noise ratios (SNRs) using simulated and real-world noisy speech datasets. Use when the user wants to benchmark on LibriSpeech, CHiME-4, or asks about evaluating this task. Reports WER.

researchpythontesting
0
3
Librispeech Asv EvalA

Evaluates the robustness of speaker recognition models against adversarial audio perturbations by measuring how effectively masked energy attacks disrupt speaker verification while preserving perceptual audio quality. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports EER (%).

researchpythongo
0
3
Librispeech Asr Correction EvalA

Evaluates confidence-based filtering strategies for applying LLMs to post-hoc correction of ASR transcripts, measuring how well the system reduces transcription errors in low-confidence segments while preserving accurate outputs. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports WER.

researchpython
0
3
Libris2s Tts EvalA

Evaluates the acoustic quality and prosodic fidelity of generated German-English speech-to-speech translation audio. It measures perceived naturalness via an automated MOS approximation, and quantifies pitch and energy accuracy against ground truth references. Use when the user wants to benchmark on LibriS2S (Frankenstein subset), or asks about evaluating this task. Reports MOSNet score.

researchpython
0
3
Libriquote EvalA

Probes the ability of zero-shot text-to-speech systems to generate expressive, character-specific utterances while preserving reference speaker timbre. It evaluates cross-sentence generation where a narration clip guides the synthesis of a fictional quotation, testing prosodic variability, emotional expressiveness, and speech intelligibility. Use when the user wants to benchmark on LibriQuote, or asks about evaluating this task. Reports WER.

researchpythongo
0
3