Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,705–7,728 of 20,849 skills
Evaluates vision-language models on their ability to interpret visual encodings (position, length, area, color, shape) and perform chart-specific analytic tasks (e.g., value retrieval, anomaly detection, correlation estimation). It probes fine-grained visual perception, reasoning under different encoding constraints, and whether model capabilities scale with size or prompting strategies. Use when the user wants to benchmark on EncQA, or asks about evaluating this task. Reports accuracy.
This evaluation protocol assesses the capability of decoder-based language models adapted into encoder-only architectures to perform diverse downstream tasks, including text classification, scoring, and information retrieval. It specifically probes how architectural modifications like bidirectional attention, pooling strategies, and dropout affect performance on standard benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, MS MARCO, or asks about evaluating this task. Reports ...
Evaluates machine translation quality and computational efficiency on a curated set of English-to-French sentences spanning simple, technical, and complex domains. It measures linguistic accuracy alongside inference latency and hardware resource consumption under consumer-grade GPU constraints. Use when the user wants to benchmark on Custom EN-FR Test Set, or asks about evaluating this task. Reports BLEU score.
Evaluates English-to-Catalan machine translation quality in the biomedical domain using a two-stage cascade pivot strategy (English→Spanish→Catalan) versus direct translation, measuring lexical overlap and fluency via BLEU scores on domain-specific test sets. Use when the user wants to benchmark on WMT Biomedical test set, El Periódico test set, or asks about evaluating this task. Reports BLEU.
Evaluates a multimodal model's capability to generate images from text prompts and edit existing images based on natural language instructions. It probes semantic alignment, fine-grained text rendering accuracy, and instruction-following fidelity across diverse visual tasks. Use when the user wants to benchmark on GenEval, DPG-bench, OneIG-Bench, TIIF-Bench mini, LeX-Bench, CVTG-2K, LongText-Bench, ImgEdit, GEdit-Bench, OmniContext, ICE-Bench, or asks about evaluating this task. Reports Word ...
Evaluates image editing capability by measuring instruction following (text similarity) and source image preservation (image similarity). It tests the model's ability to modify images based on textual instructions while maintaining relevant visual elements. Use when the user wants to benchmark on EMU-Edit, or asks about evaluating this task. Reports CLIP-T.
Evaluates a model's ability to map clinical questions to structured logical forms and to extract precise answer spans or predict answer classes from unstructured electronic medical records. It probes complex clinical reasoning, including temporal, arithmetic, and multi-sentence contextual understanding. Use when the user wants to benchmark on emrQA, or asks about evaluating this task. Reports Exact Match (EM).
This benchmark evaluates a model's ability to generate or retrieve empathetic, relevant, and fluent responses to emotionally grounded personal stories. It probes how well dialogue systems can acknowledge and react to a speaker's feelings in open-domain conversations. Use when the user wants to benchmark on EmpatheticDialogues, or asks about evaluating this task. Reports Human Empathy/Relevance/Fluency.
Evaluates an omni-modal language model's ability to engage in end-to-end spoken dialogue with vivid emotional control. It probes the model's dialogue quality, text generation accuracy under different input modalities, and its capability to control and classify speech styles/emotions. Use when the user wants to benchmark on EMOVA-EmotionDialogue-Test, or asks about evaluating this task. Reports end-to-end spoken dialogue score.
Evaluates multimodal LLMs on understanding, reasoning, and predicting fine-grained emotion transitions in short video clips. It probes capabilities across four progressive tasks: detecting whether an emotion change occurs, identifying before/after emotion states, generating evidence-grounded reasoning for transitions, and predicting the next emotion state. Use when the user wants to benchmark on EmoTrans, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to recognize and aggregate facial emotions across multiple individuals in a crowd scene. It probes the system's capacity to capture complex spatial dependencies among overlapping facial expressions to predict the dominant group-level emotion. Use when the user wants to benchmark on EmotiW2018, GECV, or asks about evaluating this task. Reports mean accuracy (mAC).
Evaluates multimodal emotion recognition and sentiment analysis capabilities across unimodal and fused modalities, alongside emotional speaker style captioning. It probes how well models capture discrete emotions, continuous sentiment, and fine-grained speaking styles from Chinese dyadic dialogues. Use when the user wants to benchmark on EmotionTalk, or asks about evaluating this task. Reports ACC.
Evaluates large language models' emotional intelligence and empathy by testing their ability to recognize key events, mixed events, implicit emotions, and user intent from real-world emotional scenarios, and to generate appropriate empathetic responses. Use when the user wants to benchmark on EmotionQueen, or asks about evaluating this task. Reports PASS rate.
Evaluates vision models' ability to classify emotions in images across multiple affective categories. It also measures the alignment between emotions expressed in text prompts and those visually present in generated images. Use when the user wants to benchmark on EmoSet, or asks about evaluating this task. Reports macro-averaged F1-score.
Probes a model's ability to perform context-aware, multi-dimensional emotion understanding by predicting Plutchik’s 8 basic emotions from rich textual scenarios. It specifically tests zero-shot multi-label emotion prediction and evaluates whether models can capture emotional entanglement (co-occurrence) rather than treating emotion dimensions as independent. Use when the user wants to benchmark on EmoScene, or asks about evaluating this task. Reports Macro F1.
Evaluates machine learning models' ability to classify brief, nonverbal vocal bursts into discrete emotional states. It probes the limits of audio classification on highly ambiguous, short-duration human vocalizations across 30 affective categories. Use when the user wants to benchmark on EmoGator, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to recognize and classify handwritten characters (digits and letters) from standardized 28x28 grayscale images. It probes robustness to case variations, class overlap, and imbalanced distributions across multiple dataset configurations. Use when the user wants to benchmark on EMNIST, or asks about evaluating this task. Reports classification accuracy.
This benchmark evaluates embodied mobile manipulation agents in open environments, testing their ability to execute long-horizon, language-conditioned tasks that require interleaved high-level planning and low-level continuous navigation/manipulation. It specifically probes reasoning fidelity, execution success, adaptability to failures, and path efficiency compared to expert trajectories. Use when the user wants to benchmark on EMMOE-100, or asks about evaluating this task. Reports PLWSR.
This benchmark evaluates multimodal large language models on their ability to perform integrated visual-textual reasoning across mathematics, physics, chemistry, and coding. It probes capabilities such as fine-grained spatial simulation, multi-hop visual inference, and cross-modal problem solving under both direct and chain-of-thought prompting conditions. Use when the user wants to benchmark on EMMA-mini, or asks about evaluating this task. Reports accuracy.
Evaluates massively multilingual language models on intrinsic next-word prediction, machine translation, text classification, math reasoning, and code generation across dozens of languages, with a specific focus on low-resource language performance and cross-lingual transfer capabilities. Use when the user wants to benchmark on Glot500-c, Parallel Bible Corpus (PBC), FLORES-200, SIB-200, Taxi-1500, MGSM, or asks about evaluating this task. Reports BLEU.
Evaluates the effectiveness of the Emilia dataset for Text-to-Speech generation by comparing models trained on Emilia versus MLS. It probes intelligibility, speaker similarity, and naturalness across formal and spontaneous speaking styles in both English and multilingual settings. Use when the user wants to benchmark on LibriSpeech-Test, Emilia-Test, Aishell-3, Common Voice, or asks about evaluating this task. Reports WER.
Human-subject validation of cross-modal emotional alignment between music and images. It probes whether pairing audio and visual stimuli based on a 13-dimensional emotional coordinate space yields higher perceptual matching accuracy compared to semantic-only or random matching. Use when the user wants to benchmark on EMID, or asks about evaluating this task. Reports accuracy.
Evaluates a multimodal emotion recognition system that fuses facial, posture, and gait cues with situational knowledge to classify emotional states. It probes the model's ability to generalize across posed and wild settings, different data modalities, and subject-independent splits. Use when the user wants to benchmark on FER-2013, CAER-S, FABO, EWalk, GroupWalk, GEMEP, or asks about evaluating this task. Reports Accuracy.
Evaluates text-to-speech models on six complex linguistic and prosodic dimensions: emotions, paralinguistics, syntactic complexity, questions, foreign words, and complex pronunciation. It uses a model-as-a-judge framework to measure pairwise preference against a baseline, alongside automated speech quality metrics. Use when the user wants to benchmark on EmergentTTS-Eval, or asks about evaluating this task. Reports win-rate.