All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,211 views
Multilingual Lang Prof EvalA

Evaluates large language models' multilingual capabilities across 100–200 languages by aggregating performance on translation, question answering, mathematics, and reasoning tasks. It tracks proficiency trends over time and correlates them with language speaker counts, GDP, and data availability. Use when the user wants to benchmark on Aggregated Multilingual Tasks (Translation, QA, Math, Reasoning), or asks about evaluating this task. Reports language proficiency scores.

researchpythongo
0
3
Multilingual Llm Downstream EvalA

Evaluates the downstream capabilities of multilingual LLMs trained on filtered pretraining data. It probes reading comprehension, general knowledge, natural language understanding, common-sense reasoning, and generative tasks across multiple languages. Use when the user wants to benchmark on FineTasks, SmolLM tasks suite, or asks about evaluating this task. Reports average rank.

researchpythonperformance
0
3
Multilingual Medical Benchmarks EvalA

Evaluates multilingual text-to-text models on medical argument mining (sequence labeling) and abstractive question answering across English, Spanish, French, and Italian. Use when the user wants to benchmark on AbstRCT, BioASQ 6B, or asks about evaluating this task. Reports sequence-level F1.

researchpythongo
0
3
Multilingual Safety EvalA

Evaluates a parameter-efficient multilingual safety guardrail's ability to classify content as safe or unsafe across high-resource and low-resource languages. It probes cross-lingual generalization and robustness against diverse harm categories using cluster-guided transfer. Use when the user wants to benchmark on Aegis-Content-Safety-2.0-Test (Aegis-CS2), HarmBench, Redteam2k, JBB-Behaviors, StrongReject, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multilingual Sts B EvalA

Evaluates the semantic similarity and cross-lingual transfer capabilities of pixel-based sentence representations by measuring how well the model captures semantic continuity across 10 languages and handles out-of-distribution text perturbations. Use when the user wants to benchmark on multilingual STS-b, Natural Questions, or asks about evaluating this task. Reports STS-b correlation.

researchpythongo
0
3
Multilingual Sts EvalA

Evaluates the ability of sentence embedding models to capture semantic similarity across monolingual and cross-lingual sentence pairs. It probes how well vector spaces are aligned across different languages and whether fine-tuning on English NLI/STS data generalizes to other languages. Use when the user wants to benchmark on STS 2017, or asks about evaluating this task. Reports Spearman's rank correlation (ρ).

researchpythongo
0
3
Multilingual Tedx EvalA

Evaluates automatic speech recognition, machine translation, and speech translation capabilities on a multilingual corpus of TEDx talks. It probes model robustness to lower-resource conditions, cross-lingual transfer, and the effectiveness of cascaded versus end-to-end modeling paradigms. Use when the user wants to benchmark on Multilingual TEDx Corpus, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Multilingual Tot Sim EvalA

Evaluates the fidelity of synthetic Tip-of-the-Tongue (ToT) queries by measuring how well they reproduce the relative ranking of retrieval systems compared to real human-authored ToT queries across four languages. Use when the user wants to benchmark on Multilingual ToT Test Collection, or asks about evaluating this task. Reports Kendall's tau & Pearson's r.

researchpythongit
0
3
Multilingual Toxicity EvalA

Evaluates multilingual toxicity detection capabilities of text classification models across multiple languages, focusing on production readiness, adversarial robustness, and handling of code-switching and obfuscation. Use when the user wants to benchmark on Production-Multilingual, Jigsaw Multilingual Toxic Comments Challenge, or asks about evaluating this task. Reports AUC-ROC.

researchpythonapi
0
3
Multilingual Toxicity Mitigation EvalA

Probes language models' ability to generate non-toxic continuations across nine languages and five scripts. It compares fine-tuning versus retrieval-based mitigation under static and continual learning settings, measuring cross-lingual transfer and the efficacy of translated training data. Use when the user wants to benchmark on HolisticBias, or asks about evaluating this task. Reports Expected Maximum Toxicity (EMT).

researchpythontesting
0
3
Multilingual Transfer EvalA

Probes cross-lingual transfer and multilingual representation learning across natural language inference, named entity recognition, question answering, and English text classification. The protocol evaluates how well a model trained on English data generalizes to 14 other languages, while also measuring per-language and multilingual fine-tuning performance. Use when the user wants to benchmark on XNLI, CoNLL-2002/2003, MLQA, GLUE, or asks about evaluating this task. Reports F1 score, Accuracy.

researchpythongo
0
3
Multilingual Translation EvalA

Evaluates large language models' multilingual instruction-following and non-English-centric translation capabilities across diverse language pairs and prompting strategies. The protocol tests how model performance varies when prompts are provided in different languages (e.g., Chinese, Finnish, English) versus automatically translated prompts. Use when the user wants to benchmark on NTREX-128, or asks about evaluating this task. Reports ChrF.

researchpythonbackend
0
3
Multilingual Vlm Bench EvalA

Evaluates vision-language models on translated benchmarks to measure cross-lingual transfer and check for performance degradation on English. It probes the model's ability to understand images and answer multiple-choice or yes/no questions in multiple European languages (DE, ES, FR, IT) while maintaining English proficiency. Use when the user wants to benchmark on MMBench (translated), ScienceQA (translated), MME (translated), POPE (translated), AI2D (translated), or asks about evaluating thi...

researchpythongo
0
3
Multilogbench EvalA

Evaluates LLMs' ability to generate appropriate logging statements for code callables across six programming languages, testing both snapshot-based code understanding and revision-history-based code evolution contexts. Use when the user wants to benchmark on MultiLogBench, or asks about evaluating this task. Reports exact-match accuracy.

developmentpythongo
0
3
Multiloko EvalA

Evaluates LLM multilingual knowledge and instruction-following across 31 languages using locally sourced, language-specific questions, while comparing performance on original versus machine-translated data. Use when the user wants to benchmark on MultiLoKo, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Multimed Asr EvalA

Evaluates multilingual automatic speech recognition (ASR) performance on medical domain audio across five languages. It probes the model's ability to accurately transcribe spoken medical terminology under diverse recording conditions, accents, and speaking roles. Use when the user wants to benchmark on MultiMed, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Multimed EvalA

Evaluates multimodal medical understanding across 11 diverse tasks including disease classification, imaging analysis, genomics, proteomics, and medical VQA. Probes cross-modal integration, generalization to out-of-distribution organs and cell types, and robustness to few-shot and zero-shot scenarios. Use when the user wants to benchmark on MultiMed, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimed St EvalA

Evaluates the capability of speech translation models to accurately convert medical speech across five languages (English, Vietnamese, German, French, Mandarin Chinese) into text. It probes both end-to-end and cascaded architectures, as well as the impact of multilingual vs. bilingual training and code-switching handling in a specialized medical domain. Use when the user wants to benchmark on MultiMed-ST, or asks about evaluating this task. Reports BLEU, BERTScore.

researchpythongo
0
3
Multimedbench EvalA

Evaluates a generalist biomedical AI model's ability to process multimodal clinical data across diverse tasks. It probes in-distribution performance on standard biomedical benchmarks, zero-shot generalization to unseen medical concepts like tuberculosis, and the clinical applicability of generated radiology reports. Use when the user wants to benchmark on MultiMedBench, Montgomery County Chest X-ray, MIMIC-CXR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Benchmarks EvalA

Evaluates multimodal perception and reasoning capabilities across image, video, and audio understanding tasks. It probes the model's ability to process heterogeneous modalities and answer complex questions or transcribe speech accurately. Use when the user wants to benchmark on AI2D, MMMU, MMStar, OCRBench, MMVet, Mathvista, LongVideoBench, DiDeMo, AVQA, MVBench, Video-MME, Aishell1, LibriSpeech, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Cot EvalA

Evaluates large vision-language models on multimodal chain-of-thought reasoning across mathematical, commonsense, visual grounding, and fine-grained identification tasks. It probes whether explicit intermediate visual representations improve cross-modal reasoning and transformer information flow. Use when the user wants to benchmark on IsoBench, MMVP, V*Bench, M3CoT-Commonsense, CoMT, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Emotion Recognition EvalA

Evaluates a model's ability to recognize emotions in conversational video clips using text, audio, and visual modalities. It probes how well identity-preserving representations and state-space fusion capture emotion-relevant acoustic and facial dynamics across different dataset configurations. Use when the user wants to benchmark on MELD, IEMOCAP, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Multimodal Eval Suite EvalA

Evaluates multimodal large language models across six domains: reasoning & math, text-rich image understanding, multi-image comprehension, general VQA, hallucination mitigation, and multilingual capability. It probes the model's ability to process visual contexts, perform complex reasoning, extract text, and answer questions accurately across diverse real-world scenarios. Use when the user wants to benchmark on Multimodal Benchmark Suite (32 datasets), or asks about evaluating this task. Repo...

researchpythongo
0
3
Multimodal Grounding EvalA

Evaluates a model's ability to localize text phrases and referring expressions within images by generating bounding box coordinates, and conversely to generate text descriptions from given bounding boxes. It probes spatial reasoning, text-image alignment, and zero-shot generalization across different expression types. Use when the user wants to benchmark on Flickr30k Entities, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports R@1.

researchpythongo
0
3
Multimodal Instruction EvalA

This protocol evaluates how concept- versus skill-targeted instruction selection strategies improve vision-language model performance under strict data budget constraints. It probes the model's zero-shot generalization across diverse tasks including VQA, OCR, spatial reasoning, and scientific understanding by aligning training data with the benchmark's dominant cognitive demand. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA (SQA-I), TextVQA, POPE, MME, MMBench (en), LL...

researchpythongo
0
3
Multimodal Math Reasoning EvalA

Evaluates the ability of multimodal large language models to solve mathematical problems that require interpreting visual diagrams alongside textual prompts. It probes complex reasoning capabilities across diverse difficulty levels and languages (English and Chinese). Use when the user wants to benchmark on MathVista, MathVerse, MathVision, OlympiadBench, WeMath, MMK12-test, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Medical Seg EvalA

Evaluates the ability of self-supervised multimodal pretraining to learn modality-agnostic representations for downstream medical image segmentation and survival prediction tasks, particularly under conditions of non-registered data and low-data regimes. Use when the user wants to benchmark on BraTS (Multimodal Brain Tumor Image Segmentation Benchmark), Medical Segmentation Decathlon (Prostate), CHAOS (Liver), or asks about evaluating this task. Reports Dice coefficient (segmentation), Concor...

researchpythontesting
0
3
Multimodal Medical Stress Test EvalA

This evaluation probes the robustness and genuine multimodal reasoning capabilities of large language models in clinical settings. It measures how model accuracy degrades when visual inputs are removed, answer options are perturbed, or distractors are replaced, revealing reliance on textual shortcuts and memorization rather than true visual-textual integration. Use when the user wants to benchmark on NEJM, JAMA, VQA-RAD, OmniMedVQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Mt EvalA

Evaluates a model's ability to leverage visual context to improve machine translation, particularly for ambiguous verbs or long sentences with irrelevant text. It also assesses the quality of learned joint visual-text embeddings through an image retrieval task. Use when the user wants to benchmark on Multi30K, Ambiguous COCO, IKEA, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Multimodal Mt Rl EvalA

Evaluates a multimodal sequence-to-sequence model's ability to generate accurate translations conditioned on both source text and image features. It specifically probes whether reinforcement learning with BLEU-based rewards can mitigate exposure bias and improve translation quality over standard supervised maximum likelihood estimation. Use when the user wants to benchmark on WMT17 multimodal machine translation shared task, or asks about evaluating this task. Reports BLEU.

researchpythonperformance
0
3
Multimodal Multiplication EvalA

This benchmark evaluates the arithmetic computation capabilities of multimodal LLMs by testing their ability to multiply numbers presented across different input modalities (text, images, audio) and representations (numerical vs. alphabetic). It isolates computational difficulty from perceptual factors by systematically varying digit length and sparsity. Use when the user wants to benchmark on HDS Benchmark, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Multimodal Oil Gas Framing EvalA

This benchmark evaluates vision-language models' ability to perform multi-label framing classification on real-world oil and gas advertising videos. It probes multimodal understanding of implicit strategic communication, cultural context, and greenwashing detection across different video lengths and geographic regions. Use when the user wants to benchmark on Multimodal Oil & Gas Advertising Benchmark, or asks about evaluating this task. Reports F-score.

researchpythongo
0
3
Multimodal Ood EvalA

Evaluates a model's ability to correctly classify in-distribution (ID) classes while detecting and segmenting out-of-distribution (OOD) objects across multiple modalities (RGB, LiDAR, video, optical flow). It probes robustness to distribution shift and mitigates overconfidence in uncertainty-based methods. Use when the user wants to benchmark on SemanticKITTI, nuScenes, CARLA-OOD, HMDB51, UCF101, Kinetics-600, HAC, EPIC-Kitchens, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Multimodal Reasoning EvalA

This evaluation probes the ability of multimodal large language models to perform complex reasoning across diverse domains (mathematics, science, diagram comprehension, and creative tasks) by requiring them to explicitly ground their reasoning in visual and textual evidence before producing a final answer. Use when the user wants to benchmark on MMMU, MathVista, AI2D, EMMA, Creation-MMBench, Creation-MMBench-TO, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Rec BenchmarkA

Evaluates the performance of multimodal deep learning recommender systems across five Amazon product categories. It probes both standard recommendation accuracy and beyond-accuracy dimensions such as novelty, diversity, popularity bias, and catalog coverage. Use when the user wants to benchmark on Amazon (Office, Toys, Beauty, Sports, Clothing), or asks about evaluating this task. Reports Recall@k.

researchpythongo
0
3
Multimodal Rec EvalA

Benchmarks classical and multimodal recommender systems by evaluating how different visual and textual feature extractors impact recommendation performance. Probes the trade-off between extractor complexity and recommendation accuracy across diverse e-commerce domains. Use when the user wants to benchmark on Office Products, Digital Music, Baby, Toys & Games, Beauty, or asks about evaluating this task. Reports nDCG.

researchpythongo
0
3
Multimodal Red Teaming EvalA

Evaluates the safety and harm susceptibility of multimodal large language models (MLLMs) when exposed to adversarial prompts across different input modalities (text-only vs. image-text). It measures how effectively these prompts bypass safety filters and the severity of the resulting harmful outputs. Use when the user wants to benchmark on Multimodal Adversarial Benchmark, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpython
0
3
Multimodal Retrieval EvalA

Evaluates fine-grained and coarse-grained cross-modal retrieval capabilities across text, image, and video modalities. It probes a model's ability to align and retrieve relevant visual or textual content given a query from a different modality, including instruction-based queries. Use when the user wants to benchmark on CaReBench, ShareGPT4V, Urban1K, DOCCI, WebVid-CoVR, MMEB, Flickr30K, MSR-VTT, MSVD, DiDeMo, or asks about evaluating this task. Reports Recall@1.

researchpythongo
0
3
Multimodal Reward Benchmarks EvalA

Evaluates multimodal reward models on preference ranking tasks across image and video domains, measuring how well they score or rank candidate responses compared to ground-truth preferences. It compares multi-response scoring against single-response baselines and generative judges, while also assessing inference efficiency and downstream policy optimization stability. Use when the user wants to benchmark on VL-RewardBench, Multimodal RewardBench, MM-RLHF RewardBench, MR2Bench-Image, VideoRewa...

researchpythongo
0
3
Multimodal Safety EvalA

Evaluates the ability of vision-language models to avoid generating unsafe outputs when given benign multimodal inputs (Safe Image + Safe Text → Unsafe Output). It measures safety alignment and task effectiveness under intent-aware prompting across multiple benchmarks. Use when the user wants to benchmark on SIUO, HoliSafe-Bench (SSU subset), MM-SafetyBench (Tiny version), or asks about evaluating this task. Reports Safety Rate.

ai-agentspythonperformance
0
3
Multimodal Tabular Automl EvalA

Evaluates automated machine learning strategies for supervised learning on multimodal tabular datasets containing text, numeric, and categorical features. It probes how well different featurization methods, neural backbones, and ensemble aggregation techniques handle mixed data types and extract predictive signal from text fields. Use when the user wants to benchmark on Multimodal Tabular Benchmark (18 datasets), or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Multimodal Tool Use EvalA

Evaluates the ability of agentic multimodal models to perform visual perception, document understanding, and mathematical reasoning. It specifically probes whether models can strategically decide when to invoke external tools (e.g., image cropping, web search, Python code execution) versus answering directly, balancing task accuracy with tool efficiency. Use when the user wants to benchmark on V-Bench, HRBench-4K/8K, TreeBench, MME-RealWorld, SEEDBench2-Plus, CharXiv, MathVista_mini, MathVers...

researchpythongo
0
3
Multimodal Visual Reasoning EvalA

This evaluation probes a model's ability to perform long-chain, multi-modal reasoning on mathematical problems that require deep visual understanding. It measures how well the model integrates image evidence with textual reasoning steps to arrive at correct answers across diverse math benchmarks. Use when the user wants to benchmark on MathVista, MathVision, MathVerse, Dynamath, OlympiadBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Vqa EvalA

Evaluates the zero-shot and few-shot visual question answering capabilities of multimodal large language models (MLLMs). It probes scene and spatial understanding, OCR capabilities, commonsense knowledge reasoning, and multimodal in-context learning across diverse benchmarks. Use when the user wants to benchmark on GQA, VQA-v2, VizWiz, TextVQA, OKVQA, POPE, MMMU (Val), MMBench (Dev), MMStar, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Understanding EvalA

Evaluates a model's ability to understand and reason over diverse visual inputs, including general VQA, document/chart understanding, OCR, and hallucination robustness. Use when the user wants to benchmark on MMMU(Val), MMStar, MME, OCRBench, HallB(Avg), MMB(Dev En V1.1), TextVQA, DoCVQA, InfoVQA, AI2D, ChartQA, RWQA, or asks about evaluating this task. Reports VLMEvalKit score.

researchpythongo
0
3
Multimodalqa EvalA

Evaluates complex question answering capabilities that require joint reasoning across text, tables, and images. It probes multi-hop reasoning, cross-modal inference, and the ability to align and process structured and unstructured data to produce correct answer lists. Use when the user wants to benchmark on MultiModalQA, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Multinrc EvalA

This benchmark evaluates LLMs' ability to perform multi-step reasoning in native non-English languages (French, Spanish, Chinese) across linguistic, wordplay, cultural/tradition, and culturally-grounded math categories. It specifically probes whether models rely on translation bias or possess deep cultural and linguistic contextual knowledge required for accurate problem-solving. Use when the user wants to benchmark on MultiNRC, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multioff Hateful Meme EvalA

Binary classification of memes as offensive or non-offensive. It probes a model's ability to detect hate speech in multimodal content by leveraging serialized scene graphs and knowledge graph entities alongside raw text. Use when the user wants to benchmark on MultiOFF, or asks about evaluating this task. Reports F1 score (offensive class).

researchpythongo
0
3
MultioutputwrapperA

Compute the MultioutputWrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultioutputWrapper, or asks how to score with MultioutputWrapper.

documentationpython
0
3
Multiparadetox EvalA

Evaluates text detoxification models across Russian, Ukrainian, and Spanish by measuring how effectively they transform toxic input into neutral output. The benchmark probes a model's ability to remove offensive language while preserving the original semantic content and maintaining grammatical fluency in the target language. Use when the user wants to benchmark on MultiParaDetox, or asks about evaluating this task. Reports STA.

researchpythongo
0
3