Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,815
skills in category
993
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,121–6,144 of 23,815 skills

Mwp Localization EvalA

Evaluates LLMs' ability to solve math word problems after socio-cultural localization of entities into low-resource languages. Probes whether models maintain reasoning accuracy when cultural context shifts, focusing solely on final answer correctness rather than step-by-step reasoning. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Mwlp Storm Repair EvalA

Evaluates an algorithm's ability to optimally partition repair targets among multiple crews and route them to minimize total weighted latency (average wait time) while balancing workload distribution across crews in post-disaster urban scenarios. Use when the user wants to benchmark on Random Environments, Champaign Case Study, or asks about evaluating this task. Reports wait.

researchpythongo
0
3
Mvtec EvalA

Unsupervised anomaly detection and pixel-level localization on industrial defect data. It probes the model's ability to distinguish normal from defective samples and precisely segment defect regions without using labeled anomalies during training. Use when the user wants to benchmark on MVTec, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Mvtec Ad EvalA

Evaluates an image classification and segmentation model's ability to detect and localize defects in industrial products without seeing anomalous examples during training. It probes the model's capacity to learn nominal feature distributions and identify deviations at both image and pixel levels. Use when the user wants to benchmark on MVTec AD, Magnetic Tile Defects (MTD), Mini Shanghai Tech Campus (mSTC), or asks about evaluating this task. Reports AUROC.

researchpythontesting
0
3
Mvsl Biomedical Fewshot EvalA

Evaluates a vision-language model's few-shot classification capability on diverse biomedical images, testing cross-modal alignment, generalization to unseen disease categories, and robustness across multiple imaging modalities and anatomical regions. Use when the user wants to benchmark on CTKidney, DermaMNIST, Kvasir, RETINA, LC25000, CHMNIST, BTMRI, OCTMNIST, BUSI, COVID-QU-Ex, KneeXray, or asks about evaluating this task. Reports classification accuracy (%).

researchpythongo
0
3
Mvpaint T2t EvalA

This evaluation probes a model's ability to generate high-quality, multi-view consistent 3D textures on arbitrary meshes conditioned on text instructions. It measures visual fidelity, distributional similarity to ground truth, and cross-view consistency through both automated generative metrics and human preference studies. Use when the user wants to benchmark on Objaverse T2T benchmark, GSO T2T benchmark, or asks about evaluating this task. Reports FID.

researchpythongo
0
3
Mvn Text Classification EvalA

Evaluates text classification models on sentiment analysis and news categorization tasks. It probes the model's ability to aggregate diverse feature views (word-level and n-gram) to predict fine-grained sentiment categories and news topics. Use when the user wants to benchmark on Stanford Sentiment Treebank, AG News, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mvl Sib EvalA

Evaluates cross-modal and text-only topical matching capabilities of vision-language models across 205 languages. It probes whether models can correctly associate images with semantically related texts (or vice versa) in a multilingual multiple-choice setting. Use when the user wants to benchmark on MVL-SIB, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mvbench EvalA

Evaluates multi-modal large language models' ability to understand video content, with a strong focus on temporal perception and static-to-dynamic task transformation across 20 diverse categories ranging from basic perception to complex reasoning. Use when the user wants to benchmark on MVBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mv Adapter Multi View EvalA

Evaluates a diffusion adapter's ability to generate geometrically consistent multi-view images conditioned on text prompts or reference images with camera parameters. It measures visual fidelity, image-text alignment, and multi-view structural similarity against ground-truth 3D scans. Use when the user wants to benchmark on Objaverse, Google Scanned Objects (GSO), or asks about evaluating this task. Reports FID.

researchpythongo
0
3
Muvienefr EvalA

Evaluates a model's ability to perform multi-task view synthesis by predicting multiple scene properties (RGB, surface normals, shading, edges, keypoints, semantic segmentation) from novel viewpoints, given a set of source-view annotations and camera poses. Use when the user wants to benchmark on Replica, SceneNet RGB-D, or asks about evaluating this task. Reports RGB.

researchpythontesting
0
3
Must Rag EvalA

This evaluation probes a model's ability to answer music-specific factual and contextual questions using retrieval-augmented generation. It measures accuracy on both in-domain artist metadata and out-of-domain music knowledge across multiple-choice formats. Use when the user wants to benchmark on ArtistMus, TrustMus, or asks about evaluating this task. Reports accuracy.

researchpythonrust
0
3
Muspike EvalA

Evaluates the quality of symbolic music generation by spiking neural networks across multiple datasets. It assesses both objective statistical properties (pitch, rhythm, harmony) and subjective cognitive/perceptual dimensions (fluency, emotion, impression, autobiographical association). Use when the user wants to benchmark on JSB Chorales, POP909, Lakh MIDI, EMOPIA, XMIDI, or asks about evaluating this task. Reports Personal preference.

researchpythongo
0
3
Musictheorybench EvalA

Evaluates a model's ability to reason about music theory concepts and understand symbolic music representations, alongside general language knowledge and structured music generation capabilities. Use when the user wants to benchmark on MusicTheoryBench, MMLU, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3
Musicsem EvalA

Evaluates multimodal models on their ability to understand, generate, and retrieve music based on semantically rich, context-aware natural language descriptions. It probes fine-grained musical semantics beyond technical attributes, including atmospheric, situational, and contextual cues. Use when the user wants to benchmark on MusicSem, or asks about evaluating this task. Reports BLEU.

researchpythontesting
0
3
Musicscore EvalA

Evaluates the ability of text-to-image generative models to produce visually coherent and structurally plausible music score images conditioned on textual descriptions of musical attributes like instrumentation, key, and composer. It benchmarks visual fidelity and distribution matching against ground-truth sheet music. Use when the user wants to benchmark on MusicScore-400, MusicScore-14k, MusicScore-200k, or asks about evaluating this task. Reports FID.

researchpythongo
0
3
Musicgen EvalA

Evaluates the capability of text-to-music generation models to produce high-fidelity, controllable audio that aligns with textual descriptions and matches human perceptual quality standards. Use when the user wants to benchmark on MusicCaps, or asks about evaluating this task. Reports FAD.

researchpythongo
0
3
Musiccaps EvalA

Evaluates a model's ability to generate high-fidelity, long-form music from complex text descriptions. It probes both audio quality/plausibility and the model's adherence to specific textual constraints such as genre, mood, tempo, and instrumentation. Use when the user wants to benchmark on MusicCaps, or asks about evaluating this task. Reports FAD.

researchpythongo
0
3
Music Tagging EvalA

Evaluates a model's ability to predict multiple audio tags (e.g., genre, mood, instruments) from short audio segments. It probes long-range temporal dependency modeling and robustness to class imbalance in user-generated music metadata. Use when the user wants to benchmark on MagnaTagATune (MTAT), Million Song Dataset (MSD), or asks about evaluating this task. Reports AUPR.

researchpythongo
0
3
Music Sep EvalA

Evaluates zero-shot language-queried audio source separation on musical instrument classes. The benchmark tests the model's ability to isolate a target instrument from a mixed audio mixture using text labels. Use when the user wants to benchmark on MUSIC, or asks about evaluating this task. Reports SDRi.

researchpythongo
0
3
Music Plagiarism Detection EvalA

Evaluates a model's ability to detect plagiarized or remixed segments within audio tracks by computing segment-level musical similarity and attributing similarities to specific elements like melody, chords, and vocals. Use when the user wants to benchmark on Similar Music Pair, or asks about evaluating this task. Reports similarity score.

researchpythongo
0
3
Music Controlnet EvalA

This evaluation probes a diffusion-based music generation model's ability to precisely follow time-varying control signals (melody, dynamics, rhythm) and global style tags (genre/mood). It measures how faithfully the generated audio adheres to these inputs while maintaining overall audio realism and diversity. Use when the user wants to benchmark on In-domain test set, MusicCaps, MusicCaps+ChatGPT, Created Controls dataset, or asks about evaluating this task. Reports Melody accuracy.

researchpythonc#
0
3
Music Autotagging EvalA

Evaluates audio representation models on music autotagging tasks, measuring how well they predict categorical genre/instrument/mood tags and continuous musical features from audio input. It compares performance across generic tag datasets and expert-annotated continuous features to highlight limitations in current evaluation practices. Use when the user wants to benchmark on MagnaTagATune, MTG-Jamendo, MGPHot-tag, MGPHot-reg, or asks about evaluating this task. Reports MAP.

researchpythongo
0
3
Music Audio Tagging EvalA

This evaluation probes a model's ability to perform large-scale music audio tagging by predicting a fixed set of semantic labels (e.g., genre, mood, instrumentation) from raw 30-second audio clips. It measures how well architectures generalize across varying dataset sizes and label granularities. Use when the user wants to benchmark on MagnaTagATune (MTT), Million Song Dataset (MSD), Private Dataset, or asks about evaluating this task. Reports top-50 tag prediction.

researchpythongo
0
3