Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,289–6,312 of 20,818 skills
Evaluates generative models on cartoon animation tasks including audio-driven facial animation, face reenactment, image-to-video generation, and frame interpolation. It probes a model's ability to produce stylized, temporally consistent video with accurate facial details and cross-modal alignment. Use when the user wants to benchmark on MagicAnime-Bench, or asks about evaluating this task. Reports VSR.
Evaluates the predictive performance of mean-aggregation GNNs with non-negative weights on link prediction and node classification tasks, while assessing the soundness, monotonicity, and logical complexity of the extracted explanatory rules. Use when the user wants to benchmark on WN18RRv1, FB237v1, NELLv1, LUBM, LogInfer-WN-hier, LogInfer-WN-sym, LogInfer-WN-hier_nmhier, or asks about evaluating this task. Reports accuracy.
Evaluates the capability of models to generate high-fidelity alpha mattes for multiple human instances in images and videos, focusing on detail preservation, instance separation, and temporal consistency across frames. Use when the user wants to benchmark on HIM2K+M-HIM2K, V-HIM60, or asks about evaluating this task. Reports MAD.
This evaluation protocol assesses the transfer learning capability of self-supervised and supervised vision models on multimodal, multitemporal, and multispectral Earth observation data. It probes downstream performance on tree species classification and agricultural/land cover segmentation tasks across varying dataset scales and fusion strategies. Use when the user wants to benchmark on TreeSatAI-TS, PASTIS-HD, FLAIR#2, FLAIR-HUB, or asks about evaluating this task. Reports weighted F1 score...
Evaluates a multimodal AI model's ability to predict clinical outcomes and adverse reactions for drug combinations from preclinical data. It probes robustness to missing modalities and generalization to novel drugs under strict hold-out splits. Use when the user wants to benchmark on TWOSIDES, DrugBank, or asks about evaluating this task. Reports AUROC.
Evaluates the multilingual machine translation and zero-shot/few-shot translation capabilities of models trained on the MADLAD-400 dataset. It probes cross-lingual generalization, low-resource language handling, and the impact of data auditing on translation quality across multiple benchmarks and language pairs. Use when the user wants to benchmark on WMT, Flores-200, NTREX, GATONES, or asks about evaluating this task. Reports BLEU.
Evaluates end-to-end autonomous materials discovery pipelines by measuring how effectively different policies (planners, generators, selectors, and agentic orchestrators) can find thermodynamically stable compounds under constrained oracle query budgets. It probes the trade-offs between discovery efficiency, structural diversity, and adaptivity as chemical complexity and stability thresholds increase. Use when the user wants to benchmark on MADE Benchmark Environments, or asks about evaluatin...
Evaluates an LLM agent's susceptibility to unethical steering prompts in a text-based adventure game environment. It probes the capability of anomaly detection systems to classify agent trajectories as ethical or unethical based on their interaction traces. Use when the user wants to benchmark on MACHIAVELLI, or asks about evaluating this task. Reports AUPRC.
Evaluates a low-resource language model's capability on standard commonsense reasoning, reading comprehension, and factual knowledge tasks adapted to Macedonian. It measures how well continued pretraining and instruction tuning improve performance on these benchmarks compared to multilingual baselines. Use when the user wants to benchmark on Macedonian Benchmarks (ARC Easy, ARC Challenge, BoolQ, HellaSwag, OpenBookQA, PIQA, WinoGrande), or asks about evaluating this task. Reports accuracy.
Evaluates vision-language models' ability to perform multimodal scientific reasoning in chemistry and materials research. It probes capabilities across data extraction, experimental understanding, and data interpretation, specifically testing spatial reasoning, cross-modal synthesis, and multi-step inference. Use when the user wants to benchmark on MaCBench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to predict conversion rates for ad clicks under multiple attribution mechanisms. It probes ranking capability by measuring how well predicted probabilities distinguish positive from negative samples, both globally and per user. The task treats conversion prediction as a weighted binary classification problem where continuous attribution weights act as sample importance weights. Use when the user wants to benchmark on MAC, or asks about evaluating thi...
Evaluates a model's ability to rerank search results to maximize topic diversity, ensuring that top-k results cover multiple relevant subtopics for a given query rather than just maximizing single-topic relevance. Use when the user wants to benchmark on TREC 2009~2012 Web Track, DU-DIV, or asks about evaluating this task. Reports α-NDCG@10.
Evaluates the effectiveness and energy efficiency of Multi-dimensional Attention (MA) modules integrated into Spiking Neural Networks (SNNs) for event-based action recognition and static image classification. Use when the user wants to benchmark on DVS128 Gesture, DVS128 Gait, ImageNet-1K, or asks about evaluating this task. Reports Top-1 Accuracy (%).
Evaluates vision-language models on multilingual, multicultural, and multimodal retrieval-augmented generation tasks. It measures how different retrieval strategies, language alignment, and model scale impact accuracy on culturally diverse image-question pairs. Use when the user wants to benchmark on CVQA, WorldCuisines, or asks about evaluating this task. Reports macro-averaged accuracy.
Evaluates the ability of multimodal retrieval models to accurately rank relevant medical documents in response to text-and-image queries. It probes domain-specific alignment, handling of complex clinical terminology, and cross-specialty generalization in safety-critical healthcare settings. Use when the user wants to benchmark on M3Retrieve, or asks about evaluating this task. Reports nNDCG@10.
Evaluates a vision-language model's ability to follow multi-modal instructions, answer knowledge-based visual questions, and generalize to unseen languages and video tasks. It probes cross-modal alignment, cross-lingual transfer, and the model's conversational response quality. Use when the user wants to benchmark on M^3IT, OK-VQA, A-OKVQA, ViQuAE, Flickr-8k-CN, FM-IQA, Chinese-FoodNet, MSRVTT, iVQA, ActivityNet-QA, MSRVTT-QA, MSVD-QA, or asks about evaluating this task. Reports ROUGE-L.
Probes long-context financial meeting understanding across three languages (EN, ZH, JA) and 11 GICS sectors. It evaluates a model's ability to condense lengthy transcripts into structured summaries, extract relevant question-answer pairs, and localize precise answers within designated sections while ignoring noise. Use when the user wants to benchmark on M3FinMeeting, or asks about evaluating this task. Reports compression ratio.
This benchmark evaluates unsupervised domain adaptation (UDA) methods for 3D medical image segmentation across eight practical domain shifts, including inter-modality changes (MRI-CT), scanner parameters, contrast presence, and radiation dose. It measures how well models trained on a source domain can segment target domain volumes without target labels, highlighting the robustness of adaptation techniques to real-world imaging variability. Use when the user wants to benchmark on AMOS, LIDC, B...
This benchmark evaluates the Chain-of-Thought reasoning capabilities of multimodal large language models on medical image understanding tasks. It probes whether models can generate transparent, step-by-step diagnostic pathways that align with clinical ground truth, rather than just producing correct final answers. Use when the user wants to benchmark on M3CoTBench, or asks about evaluating this task. Reports F1.
Evaluates vision-language models' ability to perform multi-step, multi-modal chain-of-thought reasoning across diverse domains like science, commonsense, and mathematics. It probes the model's capacity to integrate visual information with textual reasoning steps and produce accurate final answers under various prompting and fine-tuning setups. Use when the user wants to benchmark on M3CoT, or asks about evaluating this task. Reports accuracy.
Evaluates cooperative autonomous driving capabilities across perception, mapping, motion forecasting, occupancy prediction, and path planning. It probes whether multi-vehicle cooperation and realistic, non-straight trajectories improve ego-vehicle performance compared to single-vehicle baselines. Use when the user wants to benchmark on M3CAD, or asks about evaluating this task. Reports AMOTA.
Evaluates multimodal academic lecture understanding across speech recognition, speech synthesis, and slide/script generation. It probes models' ability to handle complex academic language, rare words, multimodal alignment, and knowledge comprehension. Use when the user wants to benchmark on M3AV, or asks about evaluating this task. Reports BWER, ROUGE-1/2/L.
Evaluates few-shot voice cloning systems on their ability to preserve speaker identity and transfer speaking styles using only 100 or 5 reference samples per speaker. It probes low-data robustness, style disentanglement, and naturalness in synthetic speech generation. Use when the user wants to benchmark on M2VoC 2021 Test Set, or asks about evaluating this task. Reports MOS (Quality, Speaker Similarity, Style Similarity).
Evaluates multimodal retrieval-augmented generation systems across open-domain question answering, image captioning, and fact verification. It measures how effectively a system selects and utilizes retrieved multimodal evidence to improve generation quality and factual accuracy. Use when the user wants to benchmark on M2RAG, or asks about evaluating this task. Reports CIDEr.