
Claude Skills by qhjqhj00
github.com/qhjqhj00Probes a multimodal LLM's ability to answer visual questions that require external knowledge retrieval. It specifically tests the model's capacity to dynamically decide when to retrieve information and assess the relevance of retrieved documents using self-reflective tokens, without degrading performance on standard visual-only queries. Use when the user wants to benchmark on Encyclopedic-VQA, InfoSeek, or asks about evaluating this task. Reports BERT matching score (BEM), VQA accuracy.
Evaluates large language models' ability to systematically cover bounded knowledge universes and perform compositional set-based reasoning. It probes three failure stages: completeness (missing knowledge), awareness (failure to identify requirements), and application (incorrect execution) across multiple domains and languages. Use when the user wants to benchmark on KnowledgeBerg, or asks about evaluating this task. Reports Universe F1.
This benchmark evaluates multistep soft reasoning capabilities of LLMs in long narratives, specifically testing logical deduction, object placement tracking, and team allocation across English and Korean languages. It probes cross-lingual reasoning transfer and the impact of in-context learning strategies like Chain-of-Thought prompting and task-specific hints. Use when the user wants to benchmark on Ko-MuSR, MuSR, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy and inherent social bias of LLMs on a culturally adapted Korean multiple-choice question answering benchmark. It probes whether models rely on explicit contextual information versus ingrained cultural stereotypes when answering questions about various social groups. Use when the user wants to benchmark on KoBBQ, or asks about evaluating this task. Reports accuracy.
Evaluates the functional correctness and robustness of code generation models on diverse programming tasks, including standard algorithmic problems, external library usage, and competitive programming challenges. Use when the user wants to benchmark on HumanEval(+), MBPP(+), BigCodeBench, LiveCodeBench (V5), or asks about evaluating this task. Reports pass@1.
Evaluates conversational understanding and response generation capabilities of language models in Korean. It probes dialogue comprehension (classifying topics, emotions, relations, dialog acts, and facts) and response selection (choosing or generating appropriate next utterances across various Korean dialogue contexts). Use when the user wants to benchmark on KoDialogBench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the ability of large vision-language models to generate accurate, free-form Korean responses to image-based questions. It probes fine-grained capabilities across perception, reasoning, and safety/bias, specifically testing Korean cultural recognition, OCR, document/table/chart understanding, and hallucination robustness. Use when the user wants to benchmark on KOFFVQA, or asks about evaluating this task. Reports KOFFVQA Score.
Evaluates machine translation quality for Kokborok (a low-resource Tibeto-Burman language) in both English-to-Kokborok and Kokborok-to-English directions. It probes translation adequacy, fluency, and semantic similarity using both automatic metrics and human ratings. Use when the user wants to benchmark on SMOL Test Set, WMT Test Set (Bible domain), or asks about evaluating this task. Reports BLEU.
This benchmark evaluates large language models' ability to reason through Japanese national healthcare licensing examinations across ten medical professions. It probes domain-specific clinical knowledge, multimodal image interpretation, and high-stakes decision-making under strict, profession-specific passing criteria. Use when the user wants to benchmark on KokushiMD-10, or asks about evaluating this task. Reports accuracy.
Evaluates a training-free interpretability method (Δ-IoU) for detecting false negatives in binary industrial defect detection models. It probes whether post-hoc heatmap intersections can reliably flag 'in-distribution yet confidently wrong' predictions on surface defect datasets. Use when the user wants to benchmark on Kolektor SDD, Kolektor SDD2, or asks about evaluating this task. Reports Recall.
Evaluates vision-language models on Korean multimodal comprehension, document/table/chart understanding, and open-ended generation capabilities using translated and newly curated benchmarks. Use when the user wants to benchmark on K-MMBench, K-SEED, K-MMStar, K-DTCBench, K-LLaVA-W, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of cross-lingual embedding models to capture nuanced financial semantics and terminology in low-resource Korean text, specifically measuring how well they align with human judgments of sentence similarity in specialized financial contexts. Use when the user wants to benchmark on KorFinSTS, or asks about evaluating this task. Reports Spearman’s ρ.
Probes large language models' ability to answer multiple-choice questions derived from South Korean healthcare professional licensing exams. It evaluates domain-specific medical knowledge, regional clinical guideline adherence, and reasoning capabilities in Korean. Use when the user wants to benchmark on KorMedMCQA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates vision-language models on multimodal medical reasoning using questions derived from the Korean Medical Licensing Examination. It probes the models' ability to integrate textual and visual evidence across diverse clinical imaging modalities, including cross-image reasoning when multiple scans are provided. Use when the user wants to benchmark on KorMedMCQA-V, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models on a comprehensive suite of perception-language, nonverbal reasoning, OCR-free text understanding, and web page comprehension tasks. It measures zero-shot and few-shot cross-modal transfer, in-context learning, and the ability to align visual perception with language generation without external tools or fine-tuning. Use when the user wants to benchmark on MS COCO Caption, Flickr30k, VQAv2, VizWiz, Raven IQ Test, Rendered SST-2, HatefulMemes, WebSRC, ...
Evaluates machine translation performance across 41 Creole languages, testing cross-lingual transfer and the impact of data cleaning and scale on translation quality. Use when the user wants to benchmark on Kreyol-MT, or asks about evaluating this task. Reports BLEU.
Compute the kruskal metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute kruskal, or asks how to score with kruskal.
Compute the ks_1samp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute ks_1samp, or asks how to score with ks_1samp.
Compute the ks_2samp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute ks_2samp, or asks how to score with ks_2samp.
Evaluates session-based recommendation models on predicting the next item in a user's click sequence by integrating knowledge graph attributes and temporal dynamics between clicks. Use when the user wants to benchmark on Yoochoose, Diginetica, Last-fm, or asks about evaluating this task. Reports Recall@20.
Evaluates recommender system policies across three temporal granularities: request-level list-wise ranking, whole-session sequential recommendation under reinforcement learning, and cross-session user retention optimization. It measures how well simulated agents balance immediate engagement rewards with long-term user retention and list diversity. Use when the user wants to benchmark on KuaiRand, ML-1m, or asks about evaluating this task. Reports Average L-reward.
Evaluates recommendation models on live streaming data by testing their ability to rank relevant live rooms or streamers (top-K) and predict click-through probabilities (CTR), while accounting for real-time temporal dynamics and dynamic candidate pools. Use when the user wants to benchmark on KuaiLive, or asks about evaluating this task. Reports Recall@{5, 10, 20}.
Evaluates the in-context learning capabilities of a relational foundation model on multi-table predictive tasks across diverse domains. It probes the model's ability to perform binary classification, multi-class classification, and regression directly on relational database structures without flattening or fine-tuning. Use when the user wants to benchmark on RelBenchV1, RelBenchV2, SALT, 4DBInfer, or asks about evaluating this task. Reports AUROC.
Evaluates automatic speech recognition (ASR) models on spontaneous, radio-based Bambara speech containing real-world artifacts like code-switching, overlapping speakers, and background noise. It measures transcription accuracy under pragmatic normalization conditions. Use when the user wants to benchmark on Kunkado Test, Nyana-Eval, or asks about evaluating this task. Reports WER (%).
Evaluates a model's ability to distinguish between questions it can answer confidently (known) and those it cannot (unknown). It probes the model's metacognitive uncertainty articulation and calibration under varying prompt conditions. Use when the user wants to benchmark on KUQ, or asks about evaluating this task. Reports F1-score.
Evaluates the quality of learned key-value (KV) cache eviction policies in preserving long-context reasoning and generation capabilities under strict memory constraints. It measures how well different compression strategies retain critical tokens without access to query-specific attention scores during the compression phase. Use when the user wants to benchmark on RULER-4k, OASST2-4k, BoolQ, ARC-Challenge, MMLU, HellaSwag, GovReport, or asks about evaluating this task. Reports accuracy.
Probes a model's ability to perform multimodal understanding and generation on gastrointestinal endoscopic images. It evaluates capabilities in descriptive captioning, answering clinical questions about visual findings, and synthesizing anatomically plausible medical images from text prompts. Use when the user wants to benchmark on Kvasir-VQA, or asks about evaluating this task. Reports BLEU.
Evaluates multimodal vision-language models on gastrointestinal endoscopy image understanding and clinical question answering. It probes factual recall, multi-step clinical reasoning across varying complexity levels, and robustness to realistic visual perturbations like motion blur and color shifts. Use when the user wants to benchmark on Kvasir-VQA-x1, or asks about evaluating this task. Reports BERT-F1.
Evaluates text-to-image models on knowledge-intensive generation across six high-school academic subjects and two languages. It probes scientific fidelity, logical reasoning, symbolic precision, and multilingual robustness using textbook-derived prompts and atomic checklist verification. Use when the user wants to benchmark on KVBench, or asks about evaluating this task. Reports performance_score.
Evaluates knowledge-intensive visual grounding (KVG), requiring models to combine domain-specific reasoning with fine-grained visual perception to locate specific entities in images containing multiple similar objects. Use when the user wants to benchmark on KVG-Bench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates language models' ability to recognize formal game-theoretic structures (e.g., principal-agent conflict, signaling, strategic omission) in real-world knowledge work scenarios without explicit task hints. It measures the gap between a model's theoretical understanding of these concepts and its capacity for unprompted, practical problem framing. Use when the user wants to benchmark on KWBench, or asks about evaluating this task. Reports Pass Rate.
This benchmark evaluates keyword spotting models on mobile devices by measuring classification accuracy, computational cost (FLOPs, parameters), and real-time inference latency on one-second audio utterances. It specifically tests whether temporal convolutions can replace 2D convolutions to reduce computational load while maintaining or improving accuracy. Use when the user wants to benchmark on Google Speech Commands Dataset, or asks about evaluating this task. Reports accuracy.
Compute kyokote/my_metric2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kyokote/my_metric2.
Evaluates sentiment classification capability on Kyrgyz language text. It measures how effectively a model can distinguish between positive and negative sentiments using a manually annotated benchmark dataset. Use when the user wants to benchmark on kyrgyz-sst2, or asks about evaluating this task. Reports F1-score (Weighted).
Evaluates long-context language models across 20 diverse sub-tasks spanning 3k–200k tokens, probing retrieval, reasoning, summarization, and instruction understanding. It specifically tests whether models can maintain performance and follow length constraints when context length increases, highlighting the failure of traditional n-gram metrics and the need for length-instruction-enhanced evaluation. Use when the user wants to benchmark on L-Eval, or asks about evaluating this task. Reports ex...
Evaluates a streaming foreign accent conversion system's ability to neutralize non-native pronunciation while preserving speaker identity. The protocol uses self-reconstruction mode and compares synthesized outputs against offline-generated golden speaker utterances as a reference baseline. Use when the user wants to benchmark on L2-ARCTIC (Indian subset), or asks about evaluating this task. Reports non-native accent confidence.
Probes the ability to detect phoneme-level mispronunciations in second-language (L2) speech by comparing acoustic features of learner speech against a voice-cloned native reference using dynamic time warping on MFCC envelopes. Use when the user wants to benchmark on L2-Arctic, or asks about evaluating this task. Reports accuracy.
This evaluation probes the ability of an adaptive AI-to-AI deferral framework to selectively route clinical text classification tasks between domain-adapted BERT models and LLMs based on uncertainty signals. It measures whether intelligent routing improves classification accuracy while minimizing expensive LLM usage across binary and multi-class clinical NLP tasks. Use when the user wants to benchmark on ADE Corpus V2, MIMIC-IV Treatment Outcomes, or asks about evaluating this task. Reports F...
Evaluates an agent's ability to safely manage power grid topology under unexpected line failures and fluctuating renewable energy generation. It probes robustness to sudden grid attacks and adaptability to changing energy mix proportions over a full year of seasonal scenarios. Use when the user wants to benchmark on Grid2Op (NeurIPS 2020 L2RPN), or asks about evaluating this task. Reports total_reward.
Evaluates abstractive text summarization models on their ability to generate concise, fluent, and coherent Marathi news summaries from longer source articles. It benchmarks performance against both a newly curated large-scale dataset (MahaSum) and an existing multilingual benchmark (XL-Sum Marathi subset). Use when the user wants to benchmark on XLsum, MahaSum, or asks about evaluating this task. Reports ROUGE.
Evaluates LLMs on multilingual proficiency across Spanish varieties and regional languages of Spain and Latin America (Basque, Catalan, Galician). It probes capabilities in natural language inference, reasoning, question answering, summarization, and linguistic acceptability using a resource-efficient few-shot configuration. Use when the user wants to benchmark on La Leaderboard (66 datasets), or asks about evaluating this task. Reports exact-match.
This benchmark evaluates large language models' capabilities in biology research, including literature retrieval, figure/table interpretation, database querying, protocol troubleshooting, and DNA/protein sequence manipulation. It probes whether models can perform multi-step, tool-dependent scientific reasoning or rely on memorization and heuristic guesswork. Use when the user wants to benchmark on LAB-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates GPT-3's ability to predict ground-truth labels for given instances, and analyzes whether explanation quality correlates with prediction correctness across different datasets. Use when the user wants to benchmark on CommonsenseQA, SNLI, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the robustness of histopathology image classification models to both uniform and asymmetric label noise. It compares the performance of contrastive deep embeddings against non-contrastive backbones and image-based noise-robust loss functions. The protocol measures how well classifiers maintain accuracy when training labels are corrupted. Use when the user wants to benchmark on NCT-CRC-HE-100K, PatchCamelyon, BACH, MHIST, LC25000, GasHisSDB, or asks about evaluating th...
Compute the label_ranking_average_precision_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute label_ranking_average_precision_score, or asks how to score with label_ranking_average_precision_score.
Compute the label_ranking_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute label_ranking_loss, or asks how to score with label_ranking_loss.
Evaluates open-vocabulary and closed-set object detection capabilities on remote sensing imagery. It probes a model's ability to detect novel Earth-based objects without prior training on them, as well as its efficiency when fine-tuned with limited labeled data. Use when the user wants to benchmark on LAE-1M, DIOR, DOTAv2.0, LAE-80C, or asks about evaluating this task. Reports mAP.
Evaluates a language-assisted feature transformation framework for anomaly detection. It probes the model's ability to use textual prompts to define normality boundaries and selectively suppress or emphasize specific image attributes without retraining, across both semantic and industrial anomaly detection benchmarks. Use when the user wants to benchmark on Colored MNIST, Waterbirds, CelebA, MVTec AD, VisA, or asks about evaluating this task. Reports AUROC.
Evaluates the probabilistic forecasting skill and ensemble calibration of AI weather models by comparing them against a parameter-free lagged ensemble baseline. It probes whether models trained with long-lead-time objectives suffer from under-dispersion and poor variance calibration despite strong deterministic accuracy. Use when the user wants to benchmark on Atmospheric reanalysis / IFS HRES, or asks about evaluating this task. Reports CRPS.
Evaluates Chinese legal LLM capabilities across three hierarchical levels: basic legal NLP, basic legal application, and complex legal application. Probes tasks including named entity recognition, judicial summarization, case recognition, judgment prediction, legal question answering, and legal reasoning generation to measure domain-specific text processing, analysis, and reasoning skills. Use when the user wants to benchmark on LAiW Legal Evaluation Dataset (LED), or asks about evaluating th...