
Claude Skills by qhjqhj00
github.com/qhjqhj00Compute the MulticlassRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassRecallAtFixedPrecision, or asks how to score with MulticlassRecallAtFixedPrecision.
Compute the MulticlassROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassROC, or asks how to score with MulticlassROC.
Compute the MulticlassSensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassSensitivityAtSpecificity, or asks how to score with MulticlassSensitivityAtSpecificity.
Compute the MulticlassSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassSpecificity, or asks how to score with MulticlassSpecificity.
Compute the MulticlassSpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassSpecificityAtSensitivity, or asks how to score with MulticlassSpecificityAtSensitivity.
Compute the MulticlassStatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassStatScores, or asks how to score with MulticlassStatScores.
Evaluates a model's ability to compose multiple source images into a single coherent output while following textual instructions, maintaining image quality, and preserving facial consistency in human-object interaction scenarios. Use when the user wants to benchmark on MultiCom-Bench, or asks about evaluating this task. Reports VIEScore.
Evaluates a model's ability to perform fine-grained named entity recognition and entity linking across multiple languages. It probes whether external knowledge retrieval improves classification of ambiguous or low-frequency entities compared to context-only baselines. Use when the user wants to benchmark on MultiCoNER2, or asks about evaluating this task. Reports macro-F1.
This benchmark evaluates the semantic coherence, acoustic fidelity, and audio-visual synchronization of end-to-end spoken dialogue systems that generate face-to-face conversational audio and video. It probes a model's ability to maintain contextually appropriate dialogue while producing synchronized multimodal outputs without relying on intermediate text representations. Use when the user wants to benchmark on MultiDialog, or asks about evaluating this task. Reports PPL.
This benchmark evaluates the out-of-domain generalization and robustness of Retrieval-Augmented Generation (RAG) systems across diverse domains, answer formats, and context-criticality levels. It probes whether models can correctly extract and synthesize information from noisy or specialized document collections when internal knowledge is insufficient. Use when the user wants to benchmark on BioASQ, CovidQA, SearchQA, ParaphraseRC, SyllabusQA, TechQA, RobustQA, or asks about evaluating this t...
Evaluates the ability of abstractive summarization models to generate coherent, factually accurate, and semantically aligned summaries across general news, conversational, and financial domains. It probes content selection, hallucination reduction, and domain-specific adaptation by leveraging sentence-level salience signals during generation. Use when the user wants to benchmark on CNN/Dailymail, SAMSum, Financial-news based Event-Driven Trading (EDT), or asks about evaluating this task. Repo...
Evaluates large language models on financial reasoning, comprehension, and generation across multiple modalities (text, vision, audio), languages (English, Chinese, Japanese, Spanish, Greek), and task types (IE, QA, summarization, etc.), using a difficulty-aware selection framework to ensure balanced and discriminative assessment. Use when the user wants to benchmark on IESC, FinRED, FINER-ORD, Headlines, TATSA, BRL-Math, FinQA, TATQA, CECTSUM, TGEDTSUM, RMCCF, BigData22, MDSFT, RRE, AIE, LNE...
Evaluates retrieval-augmented generation (RAG) systems on multi-hop queries that require retrieving and reasoning across multiple evidence sources. It probes both the retrieval component's ability to find relevant text chunks and the generation component's ability to synthesize accurate answers from retrieved or ground-truth evidence. Use when the user wants to benchmark on MultiHop-RAG, or asks about evaluating this task. Reports Accuracy.
Probes a vision-language model's ability to perform multi-hop compositional spatial reasoning and precise visual grounding. It tests whether models can correctly answer complex, multi-step spatial queries while simultaneously localizing the target object with high bounding box accuracy. Use when the user wants to benchmark on MultihopSpatial, or asks about evaluating this task. Reports Acc@50IoU.
Evaluates image generation models on their ability to produce identity-consistent portraits while maintaining controllability over pose, expression, and lighting. It specifically probes the trade-off between accurate identity preservation and the generation of copy-paste artifacts from reference images. Use when the user wants to benchmark on MultiID-Bench, or asks about evaluating this task. Reports face similarity (Sim(G)).
Compute the MultilabelAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAccuracy, or asks how to score with MultilabelAccuracy.
Compute the MultilabelAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAUROC, or asks how to score with MultilabelAUROC.
Compute the MultilabelAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelAveragePrecision, or asks how to score with MultilabelAveragePrecision.
Compute the MultilabelConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelConfusionMatrix, or asks how to score with MultilabelConfusionMatrix.
Compute the MultilabelCoverageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelCoverageError, or asks how to score with MultilabelCoverageError.
Compute the MultilabelEER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelEER, or asks how to score with MultilabelEER.
Compute the MultilabelExactMatch metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelExactMatch, or asks how to score with MultilabelExactMatch.
Compute the MultilabelF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelF1Score, or asks how to score with MultilabelF1Score.
Compute the MultilabelFBetaScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelFBetaScore, or asks how to score with MultilabelFBetaScore.
Compute the MultilabelHammingDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelHammingDistance, or asks how to score with MultilabelHammingDistance.
Compute the MultilabelJaccardIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelJaccardIndex, or asks how to score with MultilabelJaccardIndex.
Compute the MultilabelLogAUC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelLogAUC, or asks how to score with MultilabelLogAUC.
Compute the MultilabelMatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelMatthewsCorrCoef, or asks how to score with MultilabelMatthewsCorrCoef.
Compute the MultilabelNegativePredictiveValue metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelNegativePredictiveValue, or asks how to score with MultilabelNegativePredictiveValue.
Compute the MultilabelPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelPrecision, or asks how to score with MultilabelPrecision.
Compute the MultilabelPrecisionAtFixedRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelPrecisionAtFixedRecall, or asks how to score with MultilabelPrecisionAtFixedRecall.
Compute the MultilabelPrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelPrecisionRecallCurve, or asks how to score with MultilabelPrecisionRecallCurve.
Compute the MultilabelRankingAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRankingAveragePrecision, or asks how to score with MultilabelRankingAveragePrecision.
Compute the MultilabelRankingLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRankingLoss, or asks how to score with MultilabelRankingLoss.
Compute the MultilabelRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRecall, or asks how to score with MultilabelRecall.
Compute the MultilabelRecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRecallAtFixedPrecision, or asks how to score with MultilabelRecallAtFixedPrecision.
Compute the MultilabelROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelROC, or asks how to score with MultilabelROC.
Compute the MultilabelSensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelSensitivityAtSpecificity, or asks how to score with MultilabelSensitivityAtSpecificity.
Compute the MultilabelSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelSpecificity, or asks how to score with MultilabelSpecificity.
Compute the MultilabelSpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelSpecificityAtSensitivity, or asks how to score with MultilabelSpecificityAtSensitivity.
Compute the MultilabelStatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelStatScores, or asks how to score with MultilabelStatScores.
Evaluates abstractive and extractive summarization models on real-world civil rights lawsuits, testing their ability to synthesize information from extremely long multi-document sources and generate summaries at three distinct length granularities (long, short, tiny). Use when the user wants to benchmark on Multi-LexSum, or asks about evaluating this task. Reports ROUGE-2 F1.
This evaluation protocol assesses the multilingual adaptation capabilities of large language models across text understanding and generation tasks. It probes how continual pre-training with bilingual translation data impacts performance on low-resource versus high-resource languages, measuring robustness, transferability, and cross-lingual competitiveness. Use when the user wants to benchmark on Flores200, SIB-200, Taxi1500, BELEBELE, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates how well CLIP-based models can assess the quality and semantic alignment of image captions across multiple languages. It measures the correlation between automated CLIPScore metrics and human quality judgments, as well as classification accuracy on foil-caption tasks. Use when the user wants to benchmark on Flickr8K-Expert, Flickr8K-CF, Composite, VICR, VALSE, XVNLI, MaRVL, or asks about evaluating this task. Reports Spearman ρ.
Evaluates multilingual language models on cross-lingual commonsense reasoning and plausibility. It probes whether models can rank correct assertions and answer multiple-choice questions across 11+ languages. Use when the user wants to benchmark on MickeyProbe, X-CODAH, X-CSQA, or asks about evaluating this task. Reports Accuracy.
Evaluates cross-lingual LLM performance across 20 European languages by translating five established benchmarks (ARC, HellaSwag, TruthfulQA, GSM8K, MMLU) and measuring task accuracy on the localized prompts. Use when the user wants to benchmark on ARC, HellaSwag, TruthfulQA, GSM8K, MMLU, or asks about evaluating this task. Reports accuracy.
Evaluates the multilingual capabilities of LLMs across understanding, generation, reasoning, and instruction-following tasks in both high- and low-resource languages. It measures how well models comprehend instructions, translate, summarize, and perform commonsense reasoning across 100+ languages. Use when the user wants to benchmark on PAWS-X, FLORES-101, XL-Sum, XCOPA, Self-Instruct*, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of machine learning and transformer models to classify multilingual financial messages as legitimate or fraudulent. It probes handling of code-mixed Bangla-English text, low-resource language features, and structural indicators like URLs and phone numbers. Use when the user wants to benchmark on Financial scams detection dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates large vision-language models' ability to accurately describe images and answer questions without hallucinating objects or attributes across 13 languages. It probes cross-lingual alignment, instruction following, and hallucination mitigation in both discriminative and generative settings. Use when the user wants to benchmark on POPE MUL, MME MUL, AMBER MUL, or asks about evaluating this task. Reports Accuracy, Precision, Recall, F1, ACC, ACC+, Total Score, CHAIR, Cover, Hal, Qualifie...
Evaluates multilingual intent classification capabilities in logistics customer service, measuring how well models route user queries to parent or leaf intent categories across seen and unseen languages. It specifically probes the performance gap between native and machine-translated queries to reveal how synthetic translation overestimates model robustness in real-world routing scenarios. Use when the user wants to benchmark on Logistics Customer Service Intent Benchmark, or asks about evalu...