Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,859
skills in category
995
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,145–6,168 of 23,859 skills

Music Audio Representation EvalA

Evaluates the quality of pre-trained audio embeddings for downstream music understanding tasks including tagging, genre classification, mood prediction, pitch/instrument detection, key classification, and emotion recognition. It tests whether frozen embeddings can be effectively probed with simple MLP classifiers to achieve competitive performance without fine-tuning the backbone model. Use when the user wants to benchmark on MSDS, MSD50, MSD100, MSD500, AMM, MuMu, MTT, NSynthP, NSynthI, GTZA...

researchpythongo
0
3
Mushroom Segmentation EvalA

This benchmark evaluates the zero-shot instance segmentation capability of models trained on synthetic data when applied to real-world agricultural scenes. It probes the model's ability to generalize across domain gaps, handling varying lighting, mushroom densities, and developmental stages without fine-tuning on real annotations. Use when the user wants to benchmark on Real-world on-field dataset, M18K, or asks about evaluating this task. Reports F1 score.

researchpythonperformance
0
3
Musebench EvalA

Evaluates multimodal language models' ability to perform fine-grained, interactive reasoning over symbolic music scores and expressive performance audio. It probes capabilities in score–audio alignment, performance error detection, and expressive deviation analysis across text, audio, and image modalities. Use when the user wants to benchmark on MuseBench, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Muse EvalA

Evaluates an LLM-based planning framework's ability to decompose natural language queries into correct task selections, logical execution flows, and valid final multimodal outputs. It probes constraint-aware model orchestration and multi-modal task routing across heterogeneous AI services. Use when the user wants to benchmark on MuSE, or asks about evaluating this task. Reports Task Selection (TS).

researchpythongo
0
3
Musdb18 Separation EvalA

Probes a model's ability to isolate specific audio sources from a mixture using a provided query signal. It evaluates how well the model handles continuous latent-space conditioning and separates arbitrary or subclass instruments beyond standard training labels. Use when the user wants to benchmark on MUSDB18, or asks about evaluating this task. Reports SDR.

researchpythongit
0
3
Musdb18 Sdr EvalA

Evaluates the quality of separated audio sources (vocals, drums, bass, other) from mixed music tracks using source-to-distortion ratio, probing the model's ability to perform supervised music source separation. Use when the user wants to benchmark on MUSDB18, or asks about evaluating this task. Reports SDR.

researchpythongo
0
3
Musdb18 Hq EvalA

Evaluates the perceptual quality of multi-stem music source separation (vocals, drums, bass, other) generated by a discrete token modeling framework. It measures how well the model separates audio tracks compared to discriminative baselines, focusing on perceptual audio quality and vocal intelligibility/naturalness. Use when the user wants to benchmark on MUSDB18-HQ, or asks about evaluating this task. Reports ViSQOL.

researchpython
0
3
Musdb Sdr EvalA

Evaluates the ability of waveform-to-waveform models to separate individual musical instruments (drums, bass, other, vocals) from a mixed audio track. It probes the model's capacity to isolate sources while minimizing contamination and artifacts, measured against ground-truth stems. Use when the user wants to benchmark on MusDB, or asks about evaluating this task. Reports SDR.

researchpythonperformance
0
3
Musdb EvalA

Evaluates a model's ability to isolate individual musical stems (vocals, drums, bass, other) from mixed audio recordings, testing long-range context modeling and cross-domain attention capabilities in source separation. Use when the user wants to benchmark on MUSDB, or asks about evaluating this task. Reports SDR.

researchpythontesting
0
3
Musciclaims EvalA

Multimodal scientific claim verification, requiring models to read complex figures and captions to determine if a scientific claim is supported, neutral, or contradicted. It also probes evidence localization, basic visual understanding, cross-modal aggregation, and epistemic sensitivity. Use when the user wants to benchmark on MuSciClaims, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Muscat EvalA

Evaluates multilingual automatic speech recognition (ASR) systems on spontaneous scientific conversations, focusing on their ability to handle code-switching, transcribe domain-specific technical terms, and maintain accuracy across varying audio recording devices and segmentation methods. Use when the user wants to benchmark on MUSCAT, or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Musals EvalA

Evaluates the computational efficiency and alignment quality of a multiple sequence alignment algorithm across genomic and protein datasets. It probes the trade-off between runtime scalability and evolutionary accuracy metrics like distance distortion and gap percentage. Use when the user wants to benchmark on Greengenes 12.10, Greengenes 13.5, PDB, PFam-10k, PFam-100k, PFam-1M, or asks about evaluating this task. Reports runtime, distance distortion.

researchpythongo
0
3
Mus EvalA

Evaluates large language models' ability to perform multi-step, commonsense-rich reasoning over long natural language narratives. It probes whether models can follow complex, implicit logical chains (e.g., murder motives, object spatial reasoning, team skill matching) without relying on simple keyword heuristics or rule-based shortcuts. Use when the user wants to benchmark on MuSR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Muri 101 EvalA

Evaluates multilingual instruction-following capabilities across Natural Language Understanding (NLU) and open-ended generation (NLG) tasks, specifically probing performance on low-resource and multilingual settings using translated and native benchmarks. Use when the user wants to benchmark on Multilingual MMLU, TranslatedDolly, Taxi1500, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Murgat Attribution EvalA

Evaluates multimodal large language models' ability to generate verifiable, fact-level citations grounded in video and audio inputs. It probes whether models can correctly decompose reasoning into atomic claims and align them with precise temporal and modality-specific evidence without hallucinating references. Use when the user wants to benchmark on Video-MMMU, WorldSense, or asks about evaluating this task. Reports MURGAT-S.

researchpythongo
0
3
Mumo Instruct EvalA

Evaluates a model's ability to perform multi-objective molecular lead optimization by modifying a starting molecule to simultaneously improve multiple conflicting pharmacological properties while retaining structural similarity. Use when the user wants to benchmark on MuMO-Instruct, or asks about evaluating this task. Reports Success Rate (SR).

researchpythonperformance
0
3
Multiwoz2.1 EvalA

Evaluates a model's ability to track dialogue state (domain, slot, value triplets) across conversation turns, specifically probing its robustness to user mind-changes or 'turnback' utterances that modify previously stated intentions. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint goal accuracy.

researchpythongo
0
3
Multiwoz EvalA

Evaluates end-to-end task-oriented dialogue systems on their ability to track user goals, fulfill multi-domain requests, and generate contextually appropriate responses. It specifically probes how well models maintain conversation state and achieve user objectives without relying on full historical dialogue context. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports inform rate, success rate.

researchpythongo
0
3
Multiwoz Dst EvalA

Evaluates a model's ability to track and predict dialogue states across multiple domains in a conversation. It measures how accurately the system maintains slot-value pairs as the user's goals evolve and switches between domains like restaurant, hotel, and taxi. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).

researchpythongo
0
3
Multiwoz Dialogue EvalA

Evaluates the quality, diversity, and goal adherence of task-oriented dialogue generation models. It measures how well a model generates natural, diverse responses while correctly incorporating specified dialogue goals and slot values. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports BLEU-4.

researchpythongo
0
3
Multiwoz Dialog Summarization EvalA

Evaluates abstractive summarization models on their ability to preserve critical semantic slots, entities, and domain consistency across multi-domain dialog conversations. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports ROUGE.

researchpythongo
0
3
Multiwoz 2.1 Jga EvalA

Evaluates a model's ability to perform Dialogue State Tracking (DST) by predicting the correct values and statuses for all requested slots across multi-domain conversations. It specifically probes robustness to long-range contextual noise and class imbalance in slot status prediction. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint-goal-accuracy (JGA).

researchpythongo
0
3
Multiwoz 2.1 EvalA

Evaluates a model's ability to track and predict the complete set of user intent slots (dialogue state) across multiple domains in a multi-turn conversation. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports slot accuracy.

researchpythongo
0
3
Multivox EvalA

Evaluates voice assistants' ability to jointly ground visual and paralinguistic speech cues (e.g., pitch, emotion, volume, background sounds) in context-aware responses. It specifically tests robustness against confounding samples that flip speech properties to prevent overreliance on unimodal priors. Use when the user wants to benchmark on MultiVox, or asks about evaluating this task. Reports visual grounding and non-verbal speech signals.

researchpythonperformance
0
3