
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the ability of vision-language models to generate unified multimodal embeddings for diverse tasks including classification, visual question answering, retrieval, and visual grounding. It probes zero-shot generalization to unseen datasets and the model's capacity to follow task-specific instructions for cross-modal alignment. Use when the user wants to benchmark on MMEB, or asks about evaluating this task. Reports Precision@1.
Evaluates GUI agents across four hierarchical levels: content understanding, element grounding, single-app task automation, and multi-app task collaboration. It probes visual grounding, cross-platform generalization, and long-horizon planning capabilities while measuring both task success and step efficiency. Use when the user wants to benchmark on MMBench-GUI, or asks about evaluating this task. Reports exact-match accuracy.
This protocol evaluates the sample quality and distributional alignment of Generative Adversarial Networks across standard image datasets. It probes the generator's ability to produce high-fidelity, diverse images that match real data distributions, using both classical feature-space metrics and a novel kernel-based distance measure. Use when the user wants to benchmark on MNIST, CIFAR-10, LSUN, CelebA, or asks about evaluating this task. Reports KID.
Evaluates the geometric alignment between visual and textual attention key vectors in multimodal LLMs. It probes whether visual inputs occupy an out-of-distribution subspace relative to the text-centric key space learned during pretraining. Use when the user has predictions and gold and needs to compute MMD.
Evaluates multimodal deep research agents on iterative retrieval, citation-grounded reasoning, and long-form report synthesis. It probes how well models align textual claims with visual evidence, maintain citation discipline, and produce high-quality structured reports under multimodal constraints. Use when the user wants to benchmark on MMDeepResearch-Bench, or asks about evaluating this task. Reports Overall MMDR-Bench Score.
Evaluates large vision-language models on fine-grained visual document understanding by testing both answer prediction and multi-granularity visual grounding (region localization) across diverse document types like tables, charts, and infographics. Use when the user wants to benchmark on MMDocBench, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates multimodal retrieval systems on long documents by measuring their ability to retrieve relevant pages and fine-grained layout elements given a natural language query. Use when the user wants to benchmark on MMDocIR, or asks about evaluating this task. Reports similarity scores.
This benchmark evaluates a model's ability to handle complex, multi-turn visually-grounded dialogue and follow intricate instructions. It probes sustained contextual understanding, visual entity tracking across turns, and multi-step reasoning depth in dynamic multi-modal interactions. Use when the user wants to benchmark on MMDR-Bench, or asks about evaluating this task. Reports average human evaluation ratings.
Evaluates the quality, robustness, and efficiency of Chain-of-Thought reasoning in Large Multimodal Models. It probes whether models can generate accurate intermediate reasoning steps, maintain performance consistency between direct and CoT prompting, and produce relevant, non-redundant reasoning traces. Use when the user wants to benchmark on MME-CoT, or asks about evaluating this task. Reports F1 score.
This benchmark evaluates the emotional intelligence of multimodal large language models (MLLMs) by testing their ability to recognize emotions, perform causal reasoning about emotional triggers, and generate structured chain-of-thought explanations. It probes fine-grained sentiment analysis, multimodal fusion capabilities, and reasoning depth across diverse video scenarios. Use when the user wants to benchmark on MME-Emotion, or asks about evaluating this task. Reports CoT-S.
Evaluates multimodal large language models' ability to understand and reason over financial charts, tables, and documents. It probes fine-grained visual perception, spatial reasoning, numerical calculation, and complex financial decision-making in a domain-specific context. Use when the user wants to benchmark on MME-Finance, or asks about evaluating this task. Reports LLM-based score (0-5).
Evaluates multi-modal large language models' spatial awareness and general visual perception capabilities. It probes position reasoning, object detection, and scene understanding through both binary QA pairs and open-ended prompts. Use when the user wants to benchmark on MME, MM-Vet, or asks about evaluating this task. Reports accuracy+accuracy+, GPT-4 score.
Evaluates multimodal large language models' perception and reasoning capabilities on high-resolution, real-world images across five domains: optical character recognition, remote sensing, diagrams/tables, monitoring, and autonomous driving. Use when the user wants to benchmark on MME-RealWorld, or asks about evaluating this task. Reports Avg.
Evaluates multimodal large language models on real-world visual perception and reasoning tasks, as well as fine-grained visual-language understanding across multiple dimensions. Use when the user wants to benchmark on MME-Realworld, MMBench, or asks about evaluating this task. Reports Overall score.
This evaluation probes a model's ability to perform high-resolution visual reasoning and fine-grained grounding on complex, real-world images. It specifically tests whether the model can accurately localize relevant visual regions and correctly answer multiple-choice questions without explicit grounding supervision. Use when the user wants to benchmark on MME-Realworld, V* Bench, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models on scientific reasoning across four disciplines (math, physics, chemistry, biology) and five languages. It probes cross-lingual consistency, modality robustness (text-only vs. image-only vs. image-text), and fine-grained domain knowledge under varying visual complexity. Use when the user wants to benchmark on MME-SCI, or asks about evaluating this task. Reports accuracy.
Evaluates unified multimodal large language models (U-MLLMs) on their ability to handle mixed-modality tasks that combine visual understanding, text generation, and sequential reasoning. It probes capabilities such as interleaved image-text generation, visual chain-of-thought reasoning, and image editing with explanations. Use when the user wants to benchmark on MME-Unify, or asks about evaluating this task. Reports Acc.
Evaluates the cross-modal alignment and generalization capabilities of multimodal embedding models across classification, VQA, retrieval, and visual grounding tasks. Use when the user wants to benchmark on MMEB, or asks about evaluating this task. Reports Precision@1.
This benchmark evaluates multimodal large language models on high-resolution real-world image perception and complex reasoning tasks. It probes the models' ability to extract fine-grained details from large images and perform logical inference across diverse domains like autonomous driving, remote sensing, and document understanding. Use when the user wants to benchmark on MME-RealWorld, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates multimodal large language models on emotion recognition and emotion reasoning across diverse video clips. It probes the model's ability to extract and fuse audio, visual, and textual cues to predict categorical emotion labels and generate structured, modality-grounded explanations. Use when the user wants to benchmark on MMEVerse-Bench, EMER, or asks about evaluating this task. Reports Avg-18.
Evaluates multimodal face generation models conditioned on text and spatial inputs (semantic masks or sketches). It probes the model's ability to balance structural priors from spatial conditions with nuanced textual descriptions while maintaining photorealism and semantic alignment. Use when the user wants to benchmark on CelebA-HQ + FFHQ, or asks about evaluating this task. Reports FID.
Evaluates multimodal reasoning capabilities across STEM, puzzles, general VQA, and document understanding domains. Probes how well vision-language models perform on complex visual reasoning tasks under strict greedy decoding and high-resolution inference settings. Use when the user wants to benchmark on MMMU_val, MathVista_mini, MathVision_test, MathVerse_mini, Dynamath, LogicVista, VisuLogic, ScienceQA, RealWorldQA, MMBench-EN, MMStar_test, AI2D_test, CharXiv_reas, CharXiv_desc, or asks abou...
Evaluates multimodal models on human behavior understanding in autonomous driving contexts, covering motion prediction, text-to-motion generation, and behavior question-answering. It probes a model's ability to interpret video frames, predict future human motion, generate plausible driving-scene motions from text, and answer safety-critical behavior questions. Use when the user wants to benchmark on MMHU, or asks about evaluating this task. Reports Accuracy.
Evaluates whether physically informed audiovisual feedback improves spatial perception and task performance in medical imaging. Specifically, it measures how well users learn auditory-visual anatomical mappings and their accuracy in localizing brain tumors within a VR environment compared to unimodal baselines. Use when the user wants to benchmark on Medical imaging volumes (unspecified), or asks about evaluating this task. Reports accuracy.
Evaluates multi-modal fusion and missing data imputation strategies for predicting 12-month survival in clear cell renal cell carcinoma (ccRCC) patients. It probes a model's ability to integrate heterogeneous clinical, genomic, and imaging data while handling severe modality missingness and class imbalance. Use when the user wants to benchmark on MMIST-ccRCC, or asks about evaluating this task. Reports BAcc.
Evaluates scientific reasoning in vision-language models using bilingual (English/Hindi) multimodal questions from India's JEE Advanced exam. It probes cross-domain concept integration, meta-cognitive self-correction, and cross-lingual consistency under exam-style constraints. Use when the user wants to benchmark on mmJEE-Eval, or asks about evaluating this task. Reports Pass@1 accuracy.
Evaluates how large multimodal models (LMMs) handle factual knowledge conflicts between their internal parametric knowledge and external multimodal evidence. It probes both behavioral alignment (whether models follow internal knowledge or external context) and conflict detection capabilities across coarse- and fine-grained settings. Use when the user wants to benchmark on MMKC-Bench, or asks about evaluating this task. Reports Detection Accuracy.
Evaluates the ability of LLMs and MLLMs to perform multimodal language analysis across six high-level semantic dimensions: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. It probes cross-modal reasoning and cognitive-level semantic understanding in conversational contexts using aligned text and video utterances. Use when the user wants to benchmark on MIntRec, MIntRec2.0, MELD, IEMOCAP, MOSI, CH-SIMS v2.0, UR-FUNNY-v2, MUStARD, or asks about evaluating th...
Evaluates long-context vision-language models across five diverse tasks: visual retrieval-augmented generation, needle-in-a-haystack retrieval/reasoning, many-shot in-context learning, document summarization, and long-document VQA. It specifically probes how model performance scales with standardized context lengths from 8K to 128K tokens when processing interleaved text and images. Use when the user wants to benchmark on MMLongBench, or asks about evaluating this task. Reports substring exac...
Evaluates how fine-tuning large language models on syntactically or semantically perturbed instructions impacts downstream generalization across factual knowledge, complex reasoning, and mathematical problem-solving. It also measures potential side effects on model toxicity and truthfulness under varying noise levels during both training and evaluation. Use when the user wants to benchmark on MMLU, BBH, GSM8K, ToxiGen, TruthfulQA, or asks about evaluating this task. Reports average test accur...
Evaluates instruction-tuned language models on factual knowledge, complex reasoning, and instruction-following capabilities. The protocol measures how different data selection methods and model sizes impact performance under strict compute budgets. Use when the user wants to benchmark on MMLU, BBH, IFEval, or asks about evaluating this task. Reports 5-shot accuracy, 3-shot exact match score, 0-shot accuracy.
Evaluates broad language understanding and reasoning capabilities across multiple academic and professional domains using multiple-choice questions. It tests the model's ability to process and answer questions in a few-shot setting. Use when the user wants to benchmark on MMLU, or asks about evaluating this task. Reports macro_avg/acc_char.
Evaluates whether language models can strategically underperform on capability assessments by emulating a lower educational level (high school) on subject-specific questions, and measures how prompting strategies (zero-shot vs. chain-of-thought) affect this emulation. Use when the user wants to benchmark on MMLU, or asks about evaluating this task. Reports accuracy.
This benchmark stress-tests the reasoning capability of large language models by replacing key terms in multiple-choice questions and answers with arbitrary dummy words and their definitions. It probes whether models rely on genuine conceptual understanding or merely on lexical memorization of pre-trained vocabulary. Performance is measured across three substitution variants to isolate the impact of context modification on reasoning robustness. Use when the user wants to benchmark on MMLU-SR,...
Evaluates cross-lingual and cross-modal embedding alignment for image-text retrieval, classification, visual question answering, and visual grounding. Probes whether multilingual adaptation preserves semantic consistency across languages and modalities without degrading English performance. Use when the user wants to benchmark on MMMEB, or asks about evaluating this task. Reports P@1.
This benchmark evaluates text-to-image reasoning capabilities by requiring models to generate domain-specific diagrams, charts, and mindmaps from vague prompts. It probes factual fidelity against annotated knowledge graphs and visual clarity across six educational tiers, revealing deficits in compositional planning and abstract reasoning. Use when the user wants to benchmark on MMMG, or asks about evaluating this task. Reports MMMG-Score.
Evaluates expert-level multimodal understanding and reasoning across college-level disciplines. It probes a model's ability to interpret complex, domain-specific visual inputs combined with text, and apply specialized knowledge to solve multiple-choice or open-ended questions. Use when the user wants to benchmark on MMMU, or asks about evaluating this task. Reports micro-averaged accuracy.
This benchmark evaluates multimodal models' ability to perform robust, multi-discipline reasoning by forcing them to integrate visual and textual information without relying on shortcuts. It specifically probes resistance to guessing strategies through augmented multiple-choice options and tests true vision-text integration by embedding questions directly within images. Use when the user wants to benchmark on MMMU-Pro, or asks about evaluating this task. Reports accuracy.
Evaluates large vision-language models' ability to interpret panoramic dental X-rays. It probes fine-grained anatomical recognition, pathology detection, and clinical report generation across multiple question types. Use when the user wants to benchmark on MMOral-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal mathematical and logical reasoning capabilities of vision-language models. It probes complex multi-step problem solving, visual reasoning, logical deduction, and chart-based understanding across five diverse benchmarks. Use when the user wants to benchmark on MathVerse, MathVista, MathVision, LogicVista, ChartQA, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal street-level visual place recognition by classifying pedestrian-view locations into graph-based spatial units (nodes, edges, or combined). It probes the model's ability to fuse image, video, and textual metadata for robust geolocalization in complex urban environments. Use when the user wants to benchmark on MMS-VPR, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to localize specific 3D objects or regions within a large-scale scene based on complex natural language prompts. It probes spatial reasoning, attribute understanding, and multi-target grounding capabilities in 3D point cloud environments. Use when the user wants to benchmark on MMScan (3D Visual Grounding), or asks about evaluating this task. Reports gTop-k.
This benchmark evaluates multimodal large language models' ability to perform multi-image spatial reasoning. It probes capabilities such as tracking object and camera motion, reconstructing scenes from multiple views, and inferring spatial logic across image sequences. Use when the user wants to benchmark on MMSI-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates Large Vision-Language Models (LVLMs) on six core capabilities (coarse perception, fine-grained perception, instance reasoning, logical reasoning, science & technology, and mathematics) using a human-curated benchmark designed to enforce strict visual dependency and minimize data leakage. Use when the user wants to benchmark on MMStar, or asks about evaluating this task. Reports accuracy.
Evaluates the quality of multilingual text embeddings across diverse tasks and languages. It probes capabilities like semantic similarity, classification, retrieval, and multilingual alignment. Use when the user wants to benchmark on MTEB(Multilingual), MTEB(Europe), MTEB(Indic), or asks about evaluating this task. Reports Borda count.
Evaluates multilingual and cross-lingual text embedding capabilities across 131 tasks spanning 250+ languages. Probes performance on diverse NLP tasks including retrieval, classification, clustering, and semantic textual similarity using instruction-tuned embeddings. Use when the user wants to benchmark on MTEB Multilingual (MMTEB), or asks about evaluating this task. Reports Borda count.
Evaluates the dense retrieval capability of multilingual embedding models across multiple languages and document-query pairs. It measures how well compact models can match or exceed larger baselines on standardized multilingual retrieval benchmarks. Use when the user wants to benchmark on MTEB (Multilingual), or asks about evaluating this task. Reports MMTEB(Retrieval).
Evaluates Multimodal Large Language Models' ability to reconstruct masked text from visual context without explicit prompts. It probes layout understanding, visual grounding, and world knowledge integration by requiring models to infer missing content from surrounding text, charts, and multi-page evidence. Use when the user wants to benchmark on MMTR-Bench, or asks about evaluating this task. Reports exact-match / semantic-similarity.
Evaluates RAG systems in a live, user-centric arena setting by routing queries to appropriate retrieval pipelines and measuring response quality through direct human feedback. It probes capabilities like retrieval grounding, synthesis coherence, response relevance, and appropriate verbosity in realistic deep-research scenarios. Use when the user wants to benchmark on MMU-RAG Competition / RAG Arena, or asks about evaluating this task. Reports Preference Ratio.
Evaluates the visual reasoning and grounding capabilities of multimodal large language models (MLLMs) on basic visual patterns such as orientation, counting, viewpoint, and feature presence. It specifically probes whether models fail due to limitations in their visual encoders (e.g., CLIP) rather than language model hallucinations. Use when the user wants to benchmark on MMVP, or asks about evaluating this task. Reports accuracy.