Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,313–6,336 of 20,816 skills
Evaluates a model's ability to generate fluent, semantically accurate, and hallucination-free natural language responses grounded in video and audio inputs. It probes multimodal fusion, knowledge grounding, and dialogue generation capabilities across diverse question types and modalities. Use when the user wants to benchmark on AVSD10, NExT-OE, MUSIC-AVQA, or asks about evaluating this task. Reports CIDEr.
Evaluates graph neural networks on predicting material properties from 3D crystal and molecular structures. It probes model performance across diverse material types, physical properties, and realistic data partitioning strategies. Use when the user wants to benchmark on OMDB, QMOF, MP, ISMETAL, EDOS, PDOS, DF2D, Perovskites, EFORM, Phonons, Dielectric, LOG_GVRH, LOG_KVRH, OC20, QM9, or asks about evaluating this task. Reports MAE / Accuracy.
Evaluates vision-language models' ability to correctly identify true statements about images while rejecting culturally plausible but visually incorrect counterfactual statements. It specifically probes grounding failures and cultural reasoning biases across multiple languages and dialects. Use when the user wants to benchmark on M²CQA, or asks about evaluating this task. Reports CFHR.
Evaluates a model's ability to verify scientific claims by cross-referencing textual assertions with provided multimodal evidence (figures/diagrams). It probes cross-modal reasoning, spatial/anatomical understanding, and the generation of factually grounded explanations. Use when the user wants to benchmark on M2-Verify-Med, M2-Verify-Gen, or asks about evaluating this task. Reports Macro-F1.
Evaluates multi-step spatial and physical reasoning in multimodal language models by requiring them to generate or validate chain-of-thought plans for solving Portal 2-inspired puzzle maps. The benchmark probes the model's ability to integrate visual map layouts with textual instructions to produce physically sound, multi-step traversal strategies. Use when the user wants to benchmark on M-Portal, or asks about evaluating this task. Reports F1 score.
Evaluates the effectiveness and refinement efficiency of neural network architecture search and selection methods on graph datasets. It measures how well a method can find near-optimal models within a limited search budget and how quickly it reaches a target performance level across diverse graph topologies and tasks. Use when the user wants to benchmark on Graph Architecture Search Benchmark (22 datasets), or asks about evaluating this task. Reports classification accuracy / AUC-ROC.
Evaluates 3D spatial reasoning and combinatorial planning by requiring models to assemble jigsaw-style pieces into a 5x5x5 cube under physical constraints. The benchmark probes the model's ability to extract geometric patterns from rendered views and logically arrange pieces without gaps or overlaps. Use when the user wants to benchmark on M-Cube, or asks about evaluating this task. Reports binary_evaluation.
Evaluates multimodal information retrieval models across eight heterogeneous query-to-candidate modalities (text, image, image-text pairs) using instruction-tuned and fine-tuned vision-language models. Probes zero-shot generalization, cross-modality alignment, and the impact of instruction tuning on retrieval accuracy in large-scale candidate pools. Use when the user wants to benchmark on M-BEIR, or asks about evaluating this task. Reports Recall@5.
Probes inflexible reasoning and medical abstraction in LLMs by presenting adversarial, long-tail clinical scenarios designed to trigger the Einstellung effect. It evaluates whether models can apply deductive logic and uncertainty estimation rather than relying on rote pattern matching or memorization from pretraining data. Use when the user wants to benchmark on M-ARC, or asks about evaluating this task. Reports accuracy.
Evaluates multilingual aspect-based sentiment analysis (ABSA) models on triplet extraction (aspect term, category, sentiment) and pairwise extraction (aspect term, sentiment) across 21 languages and 7 domains. Probes cross-lingual transfer, cross-domain adaptation, and zero-shot LLM prompting capabilities. Use when the user wants to benchmark on M-ABSA, or asks about evaluating this task. Reports Micro-F1.
This benchmark evaluates multimodal large language models' ability to perform timestamp-aware summarization of long videos. It probes temporal grounding, instruction adherence regarding length constraints, and cross-modal consistency between visual/audio content and generated text descriptions. Use when the user wants to benchmark on LVSum, or asks about evaluating this task. Reports Kendall's tau & Spearman's rho.
Evaluates Large Vision-Language Models on their ability to generate factually consistent outputs aligned with visual input, specifically measuring the reduction of object hallucinations in open-ended generation while preserving general multimodal reasoning and visual grounding capabilities. Use when the user wants to benchmark on POPE, CHAIR, HallusionBench, AMBER, VizWiz, MME, LLaVA-Wild, MM-Vet, or asks about evaluating this task. Reports CHAIR (object hallucination score).
This evaluation probes the demographic fairness of large vision-language models (LVLMs) by measuring how accurately they classify occupations and predict demographic attributes (gender, race, age, skin tone) across different prompt formats. It specifically quantifies performance gaps between demographic groups to identify persistent biases in model predictions. Use when the user wants to benchmark on FACET, UTKFace, or asks about evaluating this task. Reports recall.
This benchmark evaluates multimodal models' ability to comprehend extreme-length videos (averaging ~70 minutes) by testing six core temporal understanding capabilities. It probes long-term memory, multi-hop reasoning, and instruction-following across diverse video categories like sports, documentaries, and TV shows. Use when the user wants to benchmark on LVBench, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of models to perform pixel-level semantic segmentation on large-scale, diverse image collections without human annotations. It probes unsupervised representation learning, category discovery, and fine-grained mask prediction capabilities. Use when the user wants to benchmark on ImageNet-S, ImageNet-S50, ImageNet-S300, or asks about evaluating this task. Reports mIoU.
This benchmark evaluates 3D lung tumor segmentation accuracy in CT imaging, comparing traditional CNN architectures against foundation models under standard, few-shot, and prompt-based inference regimes. It probes model robustness to varying training data sizes and input prompting strategies in a medical imaging context. Use when the user wants to benchmark on NSCLC-Radiomics (Lung1), Task06 (Medical Segmentation Decathlon), or asks about evaluating this task. Reports Dice Score.
Evaluates a deep neural network's capability to classify network traffic packets as normal or specific attack types. It probes spatial-temporal feature extraction, handling of class imbalance, and robustness against overlapping attack signatures in intrusion detection systems. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, or asks about evaluating this task. Reports Detection Rate (DR%).
Evaluates the quality, contextual variation isolation, aesthetic appeal, and identity preservation of the Lunara Aesthetic II image variation dataset compared to other web-scale datasets. Use when the user wants to benchmark on Lunara-II-Variations, Lunara-I, CC3M, LAION-2B-Aesthetic, WIT, or asks about evaluating this task. Reports LAION Aesthetics v2 score.
Evaluates a model's ability to generate physically plausible, temporally coherent HDR video from standard dynamic range (SDR) inputs. It probes reconstruction fidelity in perceptually uniform HDR spaces, temporal stability across frames, and the model's capacity to recover clipped radiance details using learned visual priors. Use when the user wants to benchmark on ARRI Cinema Footage, UPIQ, or asks about evaluating this task. Reports PU21-PSNR.
Evaluates deep learning models on multi-vendor full-field digital mammography for breast cancer diagnosis, BI-RADS classification, and breast density prediction. It specifically probes the model's robustness to domain shifts induced by different imaging vendors and X-ray energies. Use when the user wants to benchmark on LUMINA, or asks about evaluating this task. Reports AUC.
Evaluates the computational efficiency and human-readability of two General Game Playing systems (Ludii and RBG) by measuring their playout throughput and the token count required to define game rules. Use when the user wants to benchmark on Unspecified game rule suite, or asks about evaluating this task. Reports playouts.
Evaluates Luxembourgish (LTZ) language understanding across eight diverse NLU tasks, including classification, sequence labeling, and textual entailment. It probes encoder models and prompted LLMs on their ability to handle low-resource language nuances, structural complexity, and label sensitivity. Use when the user wants to benchmark on ltzGLUE, or asks about evaluating this task. Reports macro-F1.
This evaluation probes the accuracy-latency-efficiency trade-off of streaming voice agents under realistic ASR conditions. It measures how well a system maintains reasoning quality while minimizing computational overhead and response delays when processing natural speech with disfluencies, misrecognitions, and non-uniform speaking rates. Use when the user wants to benchmark on VERA (AIME and GPQA-Diamond), Spoken-MQA, BigBenchAudio, Pause-and-Repair Benchmark, or asks about evaluating this ta...
Evaluates the vision-language understanding and domain generalization capabilities of Mixture-of-Experts (MoE) models. It tests how well a long-tailed distribution-aware router preserves routing variance for vision tokens while maintaining load balancing for language tokens, impacting both accuracy and inference efficiency. Use when the user wants to benchmark on GQA, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MM-Vet, PACS, VLCS, Office-Home, DomainNet, or asks about evaluating this task. Re...