
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates a model's ability to predict the next item in a user's sequential interaction history using bidirectional context. It probes how well the model captures long-range sequential dependencies and user preferences from implicit feedback. Use when the user wants to benchmark on Amazon Beauty, Steam, MovieLens 1m, MovieLens 20m, or asks about evaluating this task. Reports HR@10.
Compute the BERTScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BERTScore, or asks how to score with BERTScore.
Evaluates language models' ability to classify sentiment and detect sarcasm across three distinct varieties of English (Australian, Indian, and British). It probes cross-variety generalization and the impact of domain (Google reviews vs. Reddit comments) on model performance. Use when the user wants to benchmark on BESSTIE, or asks about evaluating this task. Reports F-Score.
A meta-evaluation framework that scores AI benchmarks across four lifecycle stages to assess their quality, reproducibility, and usability. It evaluates how well benchmarks are designed, implemented, documented, and maintained for both foundation and non-foundation models. Use when the user has predictions and gold and needs to compute lifecycle_score.
Evaluates AI-generated scientific paper reviews across five dimensions: content faithfulness, argumentative alignment, focus consistency, question constructiveness, and AI-likeness. It probes whether models can replicate human-like evaluative reasoning and scoring rather than merely predicting scalar ratings. Use when the user wants to benchmark on AI Review Benchmark, or asks about evaluating this task. Reports MAE.
Evaluates an LLM's ability to generate correct function calls from natural language prompts, covering single, multiple, parallel, and parallel-multiple API invocations across different programming languages. Use when the user wants to benchmark on Berkeley Function-Calling Benchmark (BFCL), or asks about evaluating this task. Reports Overall Accuracy.
Evaluates data-driven machine learning models for medium-range weather forecasting over India. It probes the ability of models to capture spatial and temporal atmospheric dynamics across diverse Indian microclimates for variables like geopotential height, temperature, and precipitation. Use when the user wants to benchmark on BharatBench, or asks about evaluating this task. Reports RMSE.
Evaluates large language models on Southeast Asian linguistic and cultural capabilities, probing syntax, semantics, pragmatics, coreference resolution, and scalar implicatures in Indonesian and Tamil. Use when the user wants to benchmark on BHASA LINDSEA (Indonesian & Tamil), or asks about evaluating this task. Reports accuracy.
Evaluates multilingual machine translation and related sequence-to-sequence tasks across 36 Indian subcontinent languages. It probes the model's ability to handle morphological complexity, script diversity, code-mixing, and domain-specific adaptation through reference-based and reference-free metrics. Use when the user wants to benchmark on FLORES + IN22, Reserved Development Corpora, or asks about evaluating this task. Reports BLEU.
Evaluates the impact of the Bianet parallel corpus on Neural Machine Translation performance for English-Turkish and English-Kurdish language pairs in the news domain. It compares baseline models trained on existing corpora against models augmented with Bianet data, and assesses multilingual transfer learning benefits. Use when the user wants to benchmark on WMT2016, Bianet, SETIMES, Ubuntu & GNUME, or asks about evaluating this task. Reports BLEU.
This evaluation protocol assesses the effectiveness of individual and joint bias mitigation strategies across toxicity detection and word embeddings. It probes whether debiasing for one social identity correlates with or affects bias levels in others, and measures the trade-off between bias reduction and model utility. Use when the user wants to benchmark on Jigsaw Toxicity Dataset, CoNLL 2003, or asks about evaluating this task. Reports AUC.
Probes a model's ability to detect stereotypical and biased language in text, distinguishing between stereotypical and anti-stereotypical variants, and classifying sentences as biased or unbiased across specific social bias categories. Use when the user wants to benchmark on CrowS-Pairs, BABE, or asks about evaluating this task. Reports Stereotype Score (SS), F1-score.
Evaluates a pipeline for detecting and mitigating representation bias and explicit stereotypes in text corpora. It measures how well the pipeline generates attribute-specific word lists, quantifies demographic representation imbalances, and identifies stereotypical language compared to human annotations and baselines. Use when the user wants to benchmark on Small Heap, Small Heap Neutral, StereoSet (filtered), Small Heap Annotated, or asks about evaluating this task. Reports DR score.
This benchmark evaluates occupation prediction models for fairness regarding gender bias. It tests whether classifiers maintain consistent predictions when gender pronouns are swapped (individual fairness) and whether true positive rates are balanced between male and female bios across occupations (group fairness). Use when the user wants to benchmark on Bias in Bios, or asks about evaluating this task. Reports Balanced Accuracy (BA).
Evaluates whether text classification models exhibit unintended demographic bias by analyzing how model confidence scores are distributed across different identity groups. It probes the model's ability to rank toxic vs. non-toxic content fairly and detect systematic score shifts that threshold-dependent metrics might miss. Use when the user wants to benchmark on Synthetic Bias Test Set, Human-Labeled Online Comments, or asks about evaluating this task. Reports Subgroup AUC, BPSN AUC, BNSP AUC...
Evaluates how weight-activation quantization affects model capabilities, stereotypes, fairness, toxicity, and sentiment across demographic subgroups. It probes whether aggressive compression amplifies historical bias, disparate outcomes, and inter-subgroup disparities in generated text. Use when the user wants to benchmark on MMLU, RedditBias, WinoBias, DiscrimEval, DT-Fairness, BOLD, StereoSet, or asks about evaluating this task. Reports MMLU accuracy.
Measures social bias in conversational AI systems by generating targeted questions that trigger absolute and relative biases, then evaluating model responses for biased content and discriminatory preferences across social groups. Use when the user wants to benchmark on BiasAsker dataset, or asks about evaluating this task. Reports absolute bias rate.
Evaluates text-to-image models for multi-dimensional social biases across demographic attributes (sex, race, age) by measuring implicit distributional divergence, explicit instruction-following accuracy, and whether bias manifests as ignorance or discrimination. Use when the user wants to benchmark on BiasIG, or asks about evaluating this task. Reports Implicit Bias Score ($S_{sum}$).
Evaluates the robustness and sensitivity of multimodal large language models (MLLMs) to perturbations in spoken multiple-choice questions. It probes how models handle variations in language, accent, speaker gender, and answer option ordering, measuring both absolute correctness and prediction stability across conditions. Use when the user wants to benchmark on BiasInEar, or asks about evaluating this task. Reports Question Entropy.
Evaluates open-source bibliographic reference and citation parsers on their ability to extract structured metadata fields (author, source, year, volume, issue, page, organization) from raw citation text. Compares machine learning-based versus rule-based approaches, and assesses the impact of domain-specific retraining. Use when the user wants to benchmark on Unspecified bibliographic dataset, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to predict novel drug-disease associations by modeling them as a recommendation task using bidirectional behavioral sequences and prototype spaces. It probes cold-start generalization and robustness to highly sparse interaction data. Use when the user wants to benchmark on Gdataset, Cdataset, LRSSL, or asks about evaluating this task. Reports AUPRC.
Evaluates the capability of adapted causal LLMs to function as bidirectional encoders across text, vision, and audio modalities, measuring performance on downstream fine-tuning tasks and zero-shot/linear-probing embedding benchmarks. Use when the user wants to benchmark on MTEB v2 (English & Multilingual), MIRACL, CodeSearchNet, MNLI, XNLI, PAWS-X, MathShepherd, CodeComplexity, PAN-X, POS, Seahorse, MIEB lite, MAEB beta, Beaver, Safe, Aegis, e-SNLI-VE, BoolQ, or asks about evaluating this tas...
Evaluates a graph-based adversarial domain adaptation model's ability to classify Windows malware and benign binaries under concept drift using minimal labeled samples. It probes the model's robustness to real-world malware evolution by comparing performance against baselines on the Big-15 dataset. Use when the user wants to benchmark on Big-15, or asks about evaluating this task. Reports performance metrics.
Evaluates approximate nearest neighbor (ANN) indexing methods across four realistic, constrained workloads: filtered search, out-of-distribution data, sparse vectors, and streaming updates. It probes the trade-off between search accuracy (recall) and query throughput under strict memory and time constraints. Use when the user wants to benchmark on Big ANN Challenge Datasets, or asks about evaluating this task. Reports 10-recall@10.
Evaluates large language models on probabilistic reasoning over sets, including conceptual combination, semantic outlier detection, and logical deduction. It probes the model's ability to decouple knowledge retrieval from final inference and aggregate uncertain facts without relying solely on self-attention. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports accuracy.
Evaluates the predictability of LLM performance across diverse BIG-bench tasks using an MLP-based predictor. It probes how well model scale, task type, and in-context examples correlate with actual benchmark scores, and tests robustness under different holdout strategies. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports R².
Evaluates the ability of deep learning models to classify malware binaries into their respective family types using image-based representations, specifically probing performance on imbalanced class distributions. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F-Score.
Evaluates the effectiveness of multimodal visual feature fusion (grayscale, entropy graph, SimHash) using VGG16 for binary malware classification and family detection. It probes the model's ability to handle imbalanced malware datasets and detect obfuscated binaries. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F1-score.
Evaluates large language models on arithmetic reasoning across addition, subtraction, multiplication, and division tasks with varying digit lengths. It probes the model's ability to handle large-number computation, number tokenization consistency, and stepwise reasoning without relying on external tools. Use when the user wants to benchmark on BIG-bench arithmetic, Extra arithmetic tasks, or asks about evaluating this task. Reports exact string match.
Evaluates large language models' ability to generate correct, executable code for complex programming tasks requiring diverse function calls and compositional reasoning. It also probes instruction-following capabilities by comparing performance on verbose prompts versus condensed natural-language instructions. Use when the user wants to benchmark on BigCodeBench, or asks about evaluating this task. Reports Pass@1.
Evaluates vision-language models on remote sensing tasks including image captioning, binary visual question answering, multiple-choice questions, and referring expression/point detection. It probes the models' ability to understand multi-sensor (SAR + multispectral) and RGB Earth observation imagery, follow complex spatial instructions, and generate accurate land-use/land-cover descriptions or localized bounding boxes. Use when the user wants to benchmark on BigEarthNet.txt, or asks about eva...
Evaluates whether LLMs can accurately predict the time and space complexity of given code snippets and generate new code that satisfies explicit complexity constraints. It probes algorithmic reasoning and scalability awareness beyond mere syntactic or functional correctness. Use when the user wants to benchmark on BigO(Bench), or asks about evaluating this task. Reports Pass@k.
Evaluates big data processing systems under diverse workload patterns (record insertion, statistics computation, iterative graph computation) to measure throughput, latency, and execution time across relational, text, and graph data types. Use when the user wants to benchmark on BigOP Log Monitoring & PageRank Workloads, or asks about evaluating this task. Reports throughput (ops/sec).
Evaluates abstractive summarization models on patent documents, probing their ability to capture global discourse structure, maintain entity coherence, and generate novel content without excessive repetition or fabrication. Use when the user wants to benchmark on BIGPATENT, or asks about evaluating this task. Reports ROUGE-1 F1.
This benchmark evaluates large language models on multi-hop, multi-answer reasoning tasks within the biomedical domain. It probes the model's ability to perform step-by-step inference over biomedical knowledge graphs and generate multiple valid answers for one-to-many-to-many relationships. Use when the user wants to benchmark on BioHopR, or asks about evaluating this task. Reports Embedding-Based Precision.
This benchmark evaluates the ability of models to automatically generate concise and accurate summaries of complex, nested legislative texts. It probes extractive and abstractive summarization capabilities in a highly technical, domain-specific legal context. Use when the user wants to benchmark on BillSum, or asks about evaluating this task. Reports ROUGE F-Score.
Evaluates a bilingual (Arabic-English) large multimodal model's ability to understand diverse medical imaging modalities, answer visual questions, and generate or summarize clinical reports. It probes factual accuracy, clinical relevance, and linguistic quality across text-only, visual-question-answering, and report-generation tasks. Use when the user wants to benchmark on BiMed-MBench, Rad-VQA, SLAKE, Path-VQA, MIMIC-CXR, MIMIC-III, or asks about evaluating this task. Reports accuracy, F1, F...
This benchmark evaluates the reasoning interpretability of knowledge graph completion models by measuring how well their generated multi-hop paths or rules can be understood and validated. It probes whether models produce semantically reasonable explanations rather than just statistically valid paths, highlighting the gap between link prediction accuracy and actual explainability. Use when the user wants to benchmark on WD15K, FB15K-237, or asks about evaluating this task. Reports GI (Global ...
This section outlines the reward design used during reinforcement learning training, which functions as the primary evaluation metric. It probes the model's ability to perform multimodal logical reasoning and produce correctly formatted final answers. The protocol relies on a strict binary correctness check rather than partial credit for reasoning steps. Use when the user has predictions and gold and needs to compute binary_reward.
Compute the BinaryAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAccuracy, or asks how to score with BinaryAccuracy.
Compute the BinaryAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAUROC, or asks how to score with BinaryAUROC.
Compute the BinaryAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAveragePrecision, or asks how to score with BinaryAveragePrecision.
Compute the BinaryCalibrationError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryCalibrationError, or asks how to score with BinaryCalibrationError.
Compute the BinaryCohenKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryCohenKappa, or asks how to score with BinaryCohenKappa.
Compute the BinaryConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryConfusionMatrix, or asks how to score with BinaryConfusionMatrix.
Compute the BinaryEER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryEER, or asks how to score with BinaryEER.
Compute the BinaryF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryF1Score, or asks how to score with BinaryF1Score.
Compute the BinaryFairness metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryFairness, or asks how to score with BinaryFairness.
Compute the BinaryFBetaScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryFBetaScore, or asks how to score with BinaryFBetaScore.
Compute the BinaryGroupStatRates metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryGroupStatRates, or asks how to score with BinaryGroupStatRates.