Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,818
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,289–6,312 of 20,818 skills

Magicanime EvalA

Evaluates generative models on cartoon animation tasks including audio-driven facial animation, face reenactment, image-to-video generation, and frame interpolation. It probes a model's ability to produce stylized, temporally consistent video with accurate facial details and cross-modal alignment. Use when the user wants to benchmark on MagicAnime-Bench, or asks about evaluating this task. Reports VSR.

researchpythonperformance
0
3
Maggn Rule Extraction EvalA

Evaluates the predictive performance of mean-aggregation GNNs with non-negative weights on link prediction and node classification tasks, while assessing the soundness, monotonicity, and logical complexity of the extracted explanatory rules. Use when the user wants to benchmark on WN18RRv1, FB237v1, NELLv1, LUBM, LogInfer-WN-hier, LogInfer-WN-sym, LogInfer-WN-hier_nmhier, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Maggie Instance Matting EvalA

Evaluates the capability of models to generate high-fidelity alpha mattes for multiple human instances in images and videos, focusing on detail preservation, instance separation, and temporal consistency across frames. Use when the user wants to benchmark on HIM2K+M-HIM2K, V-HIM60, or asks about evaluating this task. Reports MAD.

researchpython
0
3
Maestro EvalA

This evaluation protocol assesses the transfer learning capability of self-supervised and supervised vision models on multimodal, multitemporal, and multispectral Earth observation data. It probes downstream performance on tree species classification and agricultural/land cover segmentation tasks across varying dataset scales and fusion strategies. Use when the user wants to benchmark on TreeSatAI-TS, PASTIS-HD, FLAIR#2, FLAIR-HUB, or asks about evaluating this task. Reports weighted F1 score...

researchpythongo
0
3
Madrigal Drug Comb EvalA

Evaluates a multimodal AI model's ability to predict clinical outcomes and adverse reactions for drug combinations from preclinical data. It probes robustness to missing modalities and generalization to novel drugs under strict hold-out splits. Use when the user wants to benchmark on TWOSIDES, DrugBank, or asks about evaluating this task. Reports AUROC.

researchpythonreact
0
3
Madlad 400 Mt EvalA

Evaluates the multilingual machine translation and zero-shot/few-shot translation capabilities of models trained on the MADLAD-400 dataset. It probes cross-lingual generalization, low-resource language handling, and the impact of data auditing on translation quality across multiple benchmarks and language pairs. Use when the user wants to benchmark on WMT, Flores-200, NTREX, GATONES, or asks about evaluating this task. Reports BLEU.

researchpython
0
3
Made EvalA

Evaluates end-to-end autonomous materials discovery pipelines by measuring how effectively different policies (planners, generators, selectors, and agentic orchestrators) can find thermodynamically stable compounds under constrained oracle query budgets. It probes the trade-offs between discovery efficiency, structural diversity, and adaptivity as chemical complexity and stability thresholds increase. Use when the user wants to benchmark on MADE Benchmark Environments, or asks about evaluatin...

researchpythongit
0
3
Machiavelli Safeguard EvalA

Evaluates an LLM agent's susceptibility to unethical steering prompts in a text-based adventure game environment. It probes the capability of anomaly detection systems to classify agent trajectories as ethical or unethical based on their interaction traces. Use when the user wants to benchmark on MACHIAVELLI, or asks about evaluating this task. Reports AUPRC.

researchpythonapi
0
3
Macedonian Benchmarks EvalA

Evaluates a low-resource language model's capability on standard commonsense reasoning, reading comprehension, and factual knowledge tasks adapted to Macedonian. It measures how well continued pretraining and instruction tuning improve performance on these benchmarks compared to multilingual baselines. Use when the user wants to benchmark on Macedonian Benchmarks (ARC Easy, ARC Challenge, BoolQ, HellaSwag, OpenBookQA, PIQA, WinoGrande), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Macbench EvalA

Evaluates vision-language models' ability to perform multimodal scientific reasoning in chemistry and materials research. It probes capabilities across data extraction, experimental understanding, and data interpretation, specifically testing spatial reasoning, cross-modal synthesis, and multi-step inference. Use when the user wants to benchmark on MaCBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mac Cvr EvalA

This benchmark evaluates a model's ability to predict conversion rates for ad clicks under multiple attribution mechanisms. It probes ranking capability by measuring how well predicted probabilities distinguish positive from negative samples, both globally and per user. The task treats conversion prediction as a weighted binary classification problem where continuous attribution weights act as sample importance weights. Use when the user wants to benchmark on MAC, or asks about evaluating thi...

researchpythongo
0
3
Ma4div Diversity EvalA

Evaluates a model's ability to rerank search results to maximize topic diversity, ensuring that top-k results cover multiple relevant subtopics for a given query rather than just maximizing single-topic relevance. Use when the user wants to benchmark on TREC 2009~2012 Web Track, DU-DIV, or asks about evaluating this task. Reports α-NDCG@10.

researchpythonperformance
0
3
Ma Snn EvalA

Evaluates the effectiveness and energy efficiency of Multi-dimensional Attention (MA) modules integrated into Spiking Neural Networks (SNNs) for event-based action recognition and static image classification. Use when the user wants to benchmark on DVS128 Gesture, DVS128 Gait, ImageNet-1K, or asks about evaluating this task. Reports Top-1 Accuracy (%).

researchpythongo
0
3
M4 Rag EvalA

Evaluates vision-language models on multilingual, multicultural, and multimodal retrieval-augmented generation tasks. It measures how different retrieval strategies, language alignment, and model scale impact accuracy on culturally diverse image-question pairs. Use when the user wants to benchmark on CVQA, WorldCuisines, or asks about evaluating this task. Reports macro-averaged accuracy.

researchpythongo
0
3
M3retrieve EvalA

Evaluates the ability of multimodal retrieval models to accurately rank relevant medical documents in response to text-and-image queries. It probes domain-specific alignment, handling of complex clinical terminology, and cross-specialty generalization in safety-critical healthcare settings. Use when the user wants to benchmark on M3Retrieve, or asks about evaluating this task. Reports nNDCG@10.

researchpythongo
0
3
M3it EvalA

Evaluates a vision-language model's ability to follow multi-modal instructions, answer knowledge-based visual questions, and generalize to unseen languages and video tasks. It probes cross-modal alignment, cross-lingual transfer, and the model's conversational response quality. Use when the user wants to benchmark on M^3IT, OK-VQA, A-OKVQA, ViQuAE, Flickr-8k-CN, FM-IQA, Chinese-FoodNet, MSRVTT, iVQA, ActivityNet-QA, MSRVTT-QA, MSVD-QA, or asks about evaluating this task. Reports ROUGE-L.

researchpythongo
0
3
M3finmeeting EvalA

Probes long-context financial meeting understanding across three languages (EN, ZH, JA) and 11 GICS sectors. It evaluates a model's ability to condense lengthy transcripts into structured summaries, extract relevant question-answer pairs, and localize precise answers within designated sections while ignoring noise. Use when the user wants to benchmark on M3FinMeeting, or asks about evaluating this task. Reports compression ratio.

researchpythongit
0
3
M3da EvalA

This benchmark evaluates unsupervised domain adaptation (UDA) methods for 3D medical image segmentation across eight practical domain shifts, including inter-modality changes (MRI-CT), scanner parameters, contrast presence, and radiation dose. It measures how well models trained on a source domain can segment target domain volumes without target labels, highlighting the robustness of adaptation techniques to real-world imaging variability. Use when the user wants to benchmark on AMOS, LIDC, B...

researchpythongo
0
3
M3cotbench EvalA

This benchmark evaluates the Chain-of-Thought reasoning capabilities of multimodal large language models on medical image understanding tasks. It probes whether models can generate transparent, step-by-step diagnostic pathways that align with clinical ground truth, rather than just producing correct final answers. Use when the user wants to benchmark on M3CoTBench, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
M3cot EvalA

Evaluates vision-language models' ability to perform multi-step, multi-modal chain-of-thought reasoning across diverse domains like science, commonsense, and mathematics. It probes the model's capacity to integrate visual information with textual reasoning steps and produce accurate final answers under various prompting and fine-tuning setups. Use when the user wants to benchmark on M3CoT, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
M3cad EvalA

Evaluates cooperative autonomous driving capabilities across perception, mapping, motion forecasting, occupancy prediction, and path planning. It probes whether multi-vehicle cooperation and realistic, non-straight trajectories improve ego-vehicle performance compared to single-vehicle baselines. Use when the user wants to benchmark on M3CAD, or asks about evaluating this task. Reports AMOTA.

researchpythongo
0
3
M3av EvalA

Evaluates multimodal academic lecture understanding across speech recognition, speech synthesis, and slide/script generation. It probes models' ability to handle complex academic language, rare words, multimodal alignment, and knowledge comprehension. Use when the user wants to benchmark on M3AV, or asks about evaluating this task. Reports BWER, ROUGE-1/2/L.

researchpython
0
3
M2voc 2021 EvalA

Evaluates few-shot voice cloning systems on their ability to preserve speaker identity and transfer speaking styles using only 100 or 5 reference samples per speaker. It probes low-data robustness, style disentanglement, and naturalness in synthetic speech generation. Use when the user wants to benchmark on M2VoC 2021 Test Set, or asks about evaluating this task. Reports MOS (Quality, Speaker Similarity, Style Similarity).

researchpython
0
3
M2rag Multimodal EvalA

Evaluates multimodal retrieval-augmented generation systems across open-domain question answering, image captioning, and fact verification. It measures how effectively a system selects and utilizes retrieved multimodal evidence to improve generation quality and factual accuracy. Use when the user wants to benchmark on M2RAG, or asks about evaluating this task. Reports CIDEr.

researchpythonperformance
0
3