All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,001 views
Bert4rec EvalA

Evaluates a model's ability to predict the next item in a user's sequential interaction history using bidirectional context. It probes how well the model captures long-range sequential dependencies and user preferences from implicit feedback. Use when the user wants to benchmark on Amazon Beauty, Steam, MovieLens 1m, MovieLens 20m, or asks about evaluating this task. Reports HR@10.

researchpython
0
3
BertscoreA

Compute the BERTScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BERTScore, or asks how to score with BERTScore.

documentationpython
0
3
Besstie EvalA

Evaluates language models' ability to classify sentiment and detect sarcasm across three distinct varieties of English (Australian, Indian, and British). It probes cross-variety generalization and the impact of domain (Google reviews vs. Reddit comments) on model performance. Use when the user wants to benchmark on BESSTIE, or asks about evaluating this task. Reports F-Score.

researchpythongo
0
3
Betterbench AssessmentA

A meta-evaluation framework that scores AI benchmarks across four lifecycle stages to assess their quality, reproducibility, and usability. It evaluates how well benchmarks are designed, implemented, documented, and maintained for both foundation and non-foundation models. Use when the user has predictions and gold and needs to compute lifecycle_score.

researchpythongo
0
3
Beyond Rating EvalA

Evaluates AI-generated scientific paper reviews across five dimensions: content faithfulness, argumentative alignment, focus consistency, question constructiveness, and AI-likeness. It probes whether models can replicate human-like evaluative reasoning and scoring rather than merely predicting scalar ratings. Use when the user wants to benchmark on AI Review Benchmark, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3
Bfcl EvalA

Evaluates an LLM's ability to generate correct function calls from natural language prompts, covering single, multiple, parallel, and parallel-multiple API invocations across different programming languages. Use when the user wants to benchmark on Berkeley Function-Calling Benchmark (BFCL), or asks about evaluating this task. Reports Overall Accuracy.

developmentjavascriptpython
0
3
Bharatbench EvalA

Evaluates data-driven machine learning models for medium-range weather forecasting over India. It probes the ability of models to capture spatial and temporal atmospheric dynamics across diverse Indian microclimates for variables like geopotential height, temperature, and precipitation. Use when the user wants to benchmark on BharatBench, or asks about evaluating this task. Reports RMSE.

datapythongit
0
3
Bhasa EvalA

Evaluates large language models on Southeast Asian linguistic and cultural capabilities, probing syntax, semantics, pragmatics, coreference resolution, and scalar implicatures in Indonesian and Tamil. Use when the user wants to benchmark on BHASA LINDSEA (Indonesian & Tamil), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Bhashaverse Translation EvalA

Evaluates multilingual machine translation and related sequence-to-sequence tasks across 36 Indian subcontinent languages. It probes the model's ability to handle morphological complexity, script diversity, code-mixing, and domain-specific adaptation through reference-based and reference-free metrics. Use when the user wants to benchmark on FLORES + IN22, Reserved Development Corpora, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3
Bianet Mt EvalA

Evaluates the impact of the Bianet parallel corpus on Neural Machine Translation performance for English-Turkish and English-Kurdish language pairs in the news domain. It compares baseline models trained on existing corpora against models augmented with Bianet data, and assesses multilingual transfer learning benefits. Use when the user wants to benchmark on WMT2016, Bianet, SETIMES, Ubuntu & GNUME, or asks about evaluating this task. Reports BLEU.

researchpythontesting
0
3
Bias Correlation Mitigation EvalA

This evaluation protocol assesses the effectiveness of individual and joint bias mitigation strategies across toxicity detection and word embeddings. It probes whether debiasing for one social identity correlates with or affects bias levels in others, and measures the trade-off between bias reduction and model utility. Use when the user wants to benchmark on Jigsaw Toxicity Dataset, CoNLL 2003, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Bias Detection EvalA

Probes a model's ability to detect stereotypical and biased language in text, distinguishing between stereotypical and anti-stereotypical variants, and classifying sentences as biased or unbiased across specific social bias categories. Use when the user wants to benchmark on CrowS-Pairs, BABE, or asks about evaluating this task. Reports Stereotype Score (SS), F1-score.

researchpythongo
0
3
Bias Detection Mitigation EvalA

Evaluates a pipeline for detecting and mitigating representation bias and explicit stereotypes in text corpora. It measures how well the pipeline generates attribute-specific word lists, quantifies demographic representation imbalances, and identifies stereotypical language compared to human annotations and baselines. Use when the user wants to benchmark on Small Heap, Small Heap Neutral, StereoSet (filtered), Small Heap Annotated, or asks about evaluating this task. Reports DR score.

researchpythongo
0
3
Bias In Bios EvalA

This benchmark evaluates occupation prediction models for fairness regarding gender bias. It tests whether classifiers maintain consistent predictions when gender pronouns are swapped (individual fairness) and whether true positive rates are balanced between male and female bios across occupations (group fairness). Use when the user wants to benchmark on Bias in Bios, or asks about evaluating this task. Reports Balanced Accuracy (BA).

researchpython
0
3
Bias Metrics EvalA

Evaluates whether text classification models exhibit unintended demographic bias by analyzing how model confidence scores are distributed across different identity groups. It probes the model's ability to rank toxic vs. non-toxic content fairly and detect systematic score shifts that threshold-dependent metrics might miss. Use when the user wants to benchmark on Synthetic Bias Test Set, Human-Labeled Online Comments, or asks about evaluating this task. Reports Subgroup AUC, BPSN AUC, BNSP AUC...

researchpythongo
0
3
Bias Quantization EvalA

Evaluates how weight-activation quantization affects model capabilities, stereotypes, fairness, toxicity, and sentiment across demographic subgroups. It probes whether aggressive compression amplifies historical bias, disparate outcomes, and inter-subgroup disparities in generated text. Use when the user wants to benchmark on MMLU, RedditBias, WinoBias, DiscrimEval, DT-Fairness, BOLD, StereoSet, or asks about evaluating this task. Reports MMLU accuracy.

researchpythongo
0
3
Biasasker EvalA

Measures social bias in conversational AI systems by generating targeted questions that trigger absolute and relative biases, then evaluating model responses for biased content and discriminatory preferences across social groups. Use when the user wants to benchmark on BiasAsker dataset, or asks about evaluating this task. Reports absolute bias rate.

documentationpythongo
0
3
Biasig EvalA

Evaluates text-to-image models for multi-dimensional social biases across demographic attributes (sex, race, age) by measuring implicit distributional divergence, explicit instruction-following accuracy, and whether bias manifests as ignorance or discrimination. Use when the user wants to benchmark on BiasIG, or asks about evaluating this task. Reports Implicit Bias Score ($S_{sum}$).

researchpythongo
0
3
Biasinear EvalA

Evaluates the robustness and sensitivity of multimodal large language models (MLLMs) to perturbations in spoken multiple-choice questions. It probes how models handle variations in language, accent, speaker gender, and answer option ordering, measuring both absolute correctness and prediction stability across conditions. Use when the user wants to benchmark on BiasInEar, or asks about evaluating this task. Reports Question Entropy.

researchpythongo
0
3
Bib Ref Parser EvalA

Evaluates open-source bibliographic reference and citation parsers on their ability to extract structured metadata fields (author, source, year, volume, issue, page, organization) from raw citation text. Compares machine learning-based versus rule-based approaches, and assesses the impact of domain-specific retraining. Use when the user wants to benchmark on Unspecified bibliographic dataset, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Bibldr Drug Repositioning EvalA

Evaluates a model's ability to predict novel drug-disease associations by modeling them as a recommendation task using bidirectional behavioral sequences and prototype spaces. It probes cold-start generalization and robustness to highly sparse interaction data. Use when the user wants to benchmark on Gdataset, Cdataset, LRSSL, or asks about evaluating this task. Reports AUPRC.

researchpythonperformance
0
3
Bidirlm EvalA

Evaluates the capability of adapted causal LLMs to function as bidirectional encoders across text, vision, and audio modalities, measuring performance on downstream fine-tuning tasks and zero-shot/linear-probing embedding benchmarks. Use when the user wants to benchmark on MTEB v2 (English & Multilingual), MIRACL, CodeSearchNet, MNLI, XNLI, PAWS-X, MathShepherd, CodeComplexity, PAN-X, POS, Seahorse, MIEB lite, MAEB beta, Beaver, Safe, Aegis, e-SNLI-VE, BoolQ, or asks about evaluating this tas...

ai-agentspythongo
0
3
Big 15 Malware Detection EvalA

Evaluates a graph-based adversarial domain adaptation model's ability to classify Windows malware and benign binaries under concept drift using minimal labeled samples. It probes the model's robustness to real-world malware evolution by comparing performance against baselines on the Big-15 dataset. Use when the user wants to benchmark on Big-15, or asks about evaluating this task. Reports performance metrics.

researchpythongo
0
3
Big Ann Competition EvalA

Evaluates approximate nearest neighbor (ANN) indexing methods across four realistic, constrained workloads: filtered search, out-of-distribution data, sparse vectors, and streaming updates. It probes the trade-off between search accuracy (recall) and query throughput under strict memory and time constraints. Use when the user wants to benchmark on Big ANN Challenge Datasets, or asks about evaluating this task. Reports 10-recall@10.

researchpythongo
0
3
Big Bench EvalA

Evaluates large language models on probabilistic reasoning over sets, including conceptual combination, semantic outlier detection, and logical deduction. It probes the model's ability to decouple knowledge retrieval from final inference and aggregate uncertain facts without relying solely on self-attention. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Big Bench Predictability EvalA

Evaluates the predictability of LLM performance across diverse BIG-bench tasks using an MLP-based predictor. It probes how well model scale, task type, and in-context examples correlate with actual benchmark scores, and tests robustness under different holdout strategies. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports R².

researchpythongo
0
3
Big2015 EvalA

Evaluates the ability of deep learning models to classify malware binaries into their respective family types using image-based representations, specifically probing performance on imbalanced class distributions. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F-Score.

researchpythongo
0
3
Big2015 Malware Detection EvalA

Evaluates the effectiveness of multimodal visual feature fusion (grayscale, entropy graph, SimHash) using VGG16 for binary malware classification and family detection. It probes the model's ability to handle imbalanced malware datasets and detect obfuscated binaries. Use when the user wants to benchmark on BIG2015, or asks about evaluating this task. Reports F1-score.

researchpythonperformance
0
3
Bigbench Arithmetic EvalA

Evaluates large language models on arithmetic reasoning across addition, subtraction, multiplication, and division tasks with varying digit lengths. It probes the model's ability to handle large-number computation, number tokenization consistency, and stepwise reasoning without relying on external tools. Use when the user wants to benchmark on BIG-bench arithmetic, Extra arithmetic tasks, or asks about evaluating this task. Reports exact string match.

researchpythongo
0
3
Bigcodebench EvalA

Evaluates large language models' ability to generate correct, executable code for complex programming tasks requiring diverse function calls and compositional reasoning. It also probes instruction-following capabilities by comparing performance on verbose prompts versus condensed natural-language instructions. Use when the user wants to benchmark on BigCodeBench, or asks about evaluating this task. Reports Pass@1.

researchpythongit
0
3
Bigearthnet Txt EvalA

Evaluates vision-language models on remote sensing tasks including image captioning, binary visual question answering, multiple-choice questions, and referring expression/point detection. It probes the models' ability to understand multi-sensor (SAR + multispectral) and RGB Earth observation imagery, follow complex spatial instructions, and generate accurate land-use/land-cover descriptions or localized bounding boxes. Use when the user wants to benchmark on BigEarthNet.txt, or asks about eva...

researchpythongo
0
3
Bigobench EvalA

Evaluates whether LLMs can accurately predict the time and space complexity of given code snippets and generate new code that satisfies explicit complexity constraints. It probes algorithmic reasoning and scalability awareness beyond mere syntactic or functional correctness. Use when the user wants to benchmark on BigO(Bench), or asks about evaluating this task. Reports Pass@k.

researchpythongo
0
3
Bigop EvalA

Evaluates big data processing systems under diverse workload patterns (record insertion, statistics computation, iterative graph computation) to measure throughput, latency, and execution time across relational, text, and graph data types. Use when the user wants to benchmark on BigOP Log Monitoring & PageRank Workloads, or asks about evaluating this task. Reports throughput (ops/sec).

businesspythongo
0
3
Bigpatent EvalA

Evaluates abstractive summarization models on patent documents, probing their ability to capture global discourse structure, maintain entity coherence, and generate novel content without excessive repetition or fabrication. Use when the user wants to benchmark on BIGPATENT, or asks about evaluating this task. Reports ROUGE-1 F1.

researchpython
0
3
Bihopr EvalA

This benchmark evaluates large language models on multi-hop, multi-answer reasoning tasks within the biomedical domain. It probes the model's ability to perform step-by-step inference over biomedical knowledge graphs and generate multiple valid answers for one-to-many-to-many relationships. Use when the user wants to benchmark on BioHopR, or asks about evaluating this task. Reports Embedding-Based Precision.

ai-agentspython
0
3
Billsum EvalA

This benchmark evaluates the ability of models to automatically generate concise and accurate summaries of complex, nested legislative texts. It probes extractive and abstractive summarization capabilities in a highly technical, domain-specific legal context. Use when the user wants to benchmark on BillSum, or asks about evaluating this task. Reports ROUGE F-Score.

researchpythonexpress
0
3
Bimedx2 Medical EvalA

Evaluates a bilingual (Arabic-English) large multimodal model's ability to understand diverse medical imaging modalities, answer visual questions, and generate or summarize clinical reports. It probes factual accuracy, clinical relevance, and linguistic quality across text-only, visual-question-answering, and report-generation tasks. Use when the user wants to benchmark on BiMed-MBench, Rad-VQA, SLAKE, Path-VQA, MIMIC-CXR, MIMIC-III, or asks about evaluating this task. Reports accuracy, F1, F...

researchpythongo
0
3
Bimr Interpretability EvalA

This benchmark evaluates the reasoning interpretability of knowledge graph completion models by measuring how well their generated multi-hop paths or rules can be understood and validated. It probes whether models produce semantically reasonable explanations rather than just statistically valid paths, highlighting the gap between link prediction accuracy and actual explainability. Use when the user wants to benchmark on WD15K, FB15K-237, or asks about evaluating this task. Reports GI (Global ...

researchpythongo
0
3
Binary RewardA

This section outlines the reward design used during reinforcement learning training, which functions as the primary evaluation metric. It probes the model's ability to perform multimodal logical reasoning and produce correctly formatted final answers. The protocol relies on a strict binary correctness check rather than partial credit for reasoning steps. Use when the user has predictions and gold and needs to compute binary_reward.

researchpythongo
0
3
BinaryaccuracyA

Compute the BinaryAccuracy metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAccuracy, or asks how to score with BinaryAccuracy.

documentationpythongit
0
3
BinaryaurocA

Compute the BinaryAUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAUROC, or asks how to score with BinaryAUROC.

documentationpythongit
0
3
BinaryaverageprecisionA

Compute the BinaryAveragePrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryAveragePrecision, or asks how to score with BinaryAveragePrecision.

documentationpythongit
0
3
BinarycalibrationerrorA

Compute the BinaryCalibrationError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryCalibrationError, or asks how to score with BinaryCalibrationError.

documentationpythongit
0
3
BinarycohenkappaA

Compute the BinaryCohenKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryCohenKappa, or asks how to score with BinaryCohenKappa.

documentationpythongit
0
3
BinaryconfusionmatrixA

Compute the BinaryConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryConfusionMatrix, or asks how to score with BinaryConfusionMatrix.

documentationpythongit
0
3
BinaryeerA

Compute the BinaryEER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryEER, or asks how to score with BinaryEER.

documentationpythongit
0
3
Binaryf1scoreA

Compute the BinaryF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryF1Score, or asks how to score with BinaryF1Score.

documentationpythongit
0
3
BinaryfairnessA

Compute the BinaryFairness metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryFairness, or asks how to score with BinaryFairness.

documentationpythongit
0
3
BinaryfbetascoreA

Compute the BinaryFBetaScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryFBetaScore, or asks how to score with BinaryFBetaScore.

documentationpythongit
0
3
BinarygroupstatratesA

Compute the BinaryGroupStatRates metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryGroupStatRates, or asks how to score with BinaryGroupStatRates.

documentationpythongit
0
3