All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs1,982 views
Adbench EvalA

Evaluates tabular anomaly detection models on their ability to identify outliers in medium- and high-dimensional datasets by measuring ranking quality and precision-recall trade-offs under a standardized semi-supervised protocol. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports ROC-AUC.

researchpython
0
3
Adcraft EvalA

Evaluates reinforcement learning agents' ability to optimize bidding strategies and budget allocation in a non-stationary, stochastic Search Engine Marketing (SEM) simulation. It probes how well policies handle sparse feedback, shifting reward landscapes, and long-term profitability constraints over a simulated campaign. Use when the user wants to benchmark on AdCraft Environment, or asks about evaluating this task. Reports NCP.

researchpythongo
0
3
Ade20k Scene Parse EvalA

Evaluates a model's ability to perform dense pixel-wise semantic segmentation across 150 common scene categories, including both discrete objects and amorphous 'stuff' classes. It probes fine-grained scene understanding and the model's capacity to handle class imbalance and varying object scales. Use when the user wants to benchmark on SceneParse150, or asks about evaluating this task. Reports Mean IoU.

researchpythongo
0
3
Adept Prosody Clone EvalA

Evaluates a zero-shot multispeaker TTS model's ability to clone both speaker voice and fine-grained prosody from untranscribed reference audio. It measures intelligibility, spectral/prosodic fidelity, and perceptual similarity against human references. Use when the user wants to benchmark on ADEPT, or asks about evaluating this task. Reports Phone Error Rate (PER).

researchpython
0
3
Ader Sr EvalA

Evaluates continual learning performance for session-based recommendation by measuring how well a model maintains prediction accuracy on historical items while adapting to new sessions over time. It probes stability-plasticity trade-offs by averaging recommendation quality across multiple sequential update cycles. Use when the user wants to benchmark on DIGINETICA, YOOCHOOSE, or asks about evaluating this task. Reports Recall@k.

researchpythongo
0
3
Adiabatic Quantum Benchmark EvalA

Evaluates the performance of adiabatic quantum optimization on complex network analysis tasks. It benchmarks quantum annealing against classical methods on Chimera Ising spin glass instances, independent set problems, planted-solution instances, and community detection. Use when the user wants to benchmark on Chimera Ising spin glass instances, Independent set problems, Planted-solution instances, Community detection, or asks about evaluating this task. Reports D-Wave run-time estimation.

researchpythongo
0
3
Adjusted Mutual Info ScoreA

Compute the adjusted_mutual_info_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute adjusted_mutual_info_score, or asks how to score with adjusted_mutual_info_score.

documentationpython
0
3
Adjusted Rand ScoreA

Compute the adjusted_rand_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute adjusted_rand_score, or asks how to score with adjusted_rand_score.

documentationpythonspring
0
3
AdjustedmutualinfoscoreA

Compute the AdjustedMutualInfoScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AdjustedMutualInfoScore, or asks how to score with AdjustedMutualInfoScore.

documentationpython
0
3
AdjustedrandscoreA

Compute the AdjustedRandScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AdjustedRandScore, or asks how to score with AdjustedRandScore.

documentationpython
0
3
Admedtagger Medical EvalA

Evaluates the ability of lightweight BERT-based models to classify Polish medical texts into five clinical categories. The benchmark tests knowledge distillation from a large LLM teacher to smaller classifiers, with ground truth curated by medical experts. Use when the user wants to benchmark on ADMEDTAGGER Physician-Validated Test Sets, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Admeood EvalA

Evaluates the out-of-distribution generalization of drug property prediction models under two specific domain shifts: noise-level-based confidence categorization (Noise Shift) and inconsistent labels across sources (Concept Conflict Drift). It probes whether models can maintain predictive performance when trained on in-distribution data and tested on molecular domains with shifted label distributions or conflicting assay results. Use when the user wants to benchmark on ADMEOOD, or asks about ...

researchpythongo
0
3
Admiere EvalA

Evaluates a model's ability to understand and represent multimodal idiomaticity by ranking images based on their alignment with a given context sentence containing a nominal compound. It probes vision-language model alignment, figurative language reasoning, and the capacity to distinguish between literal and idiomatic senses. Use when the user wants to benchmark on AdMIRe, or asks about evaluating this task. Reports Top Image Accuracy.

researchpythongo
0
3
Adni Fl EvalA

Evaluates the performance of federated learning algorithms for binary classification of Alzheimer's disease versus normal controls using structural MRI-derived features. It probes how well FL methods handle non-IID data distributions and domain shifts across different scanner parameters (1.5T vs 3.0T) while preserving data privacy. Use when the user wants to benchmark on ADNI, or asks about evaluating this task. Reports ACC.

researchpythongo
0
3
Adp EvalA

Evaluates the performance of LLM agents fine-tuned with the Agent Data Protocol (ADP) across software engineering, web browsing, OS/database tool use, and general reasoning tasks. Use when the user wants to benchmark on SWE-Bench Verified, WebArena, AgentBench, GAIA, or asks about evaluating this task. Reports unit test pass rate.

researchpythongo
0
3
Adrd Bench EvalA

Evaluates LLMs on domain-specific knowledge and clinical reasoning for Alzheimer's Disease and Related Dementias (ADRD), as well as practical daily caregiving scenarios. It probes both factual recall and error detection capabilities in a medical context. Use when the user wants to benchmark on ADRD-Bench, or asks about evaluating this task. Reports exact match accuracy.

researchpythongo
0
3
Ads Violation Cause EvalA

Evaluates an automated root-cause analysis tool for autonomous driving systems by measuring its ability to correctly identify the faulty component and the specific output message that caused a driving violation in simulation. It also measures the debugging scope reduction and computational efficiency of the tool. Use when the user wants to benchmark on ADS Violation Cause Benchmark, or asks about evaluating this task. Reports component-level success.

researchpythongo
0
3
Adte Tta EvalA

Evaluates test-time adaptation (TTA) capabilities of vision-language models under distribution shift and class imbalance. It probes how well a model can adapt to out-of-distribution and cross-domain image classification tasks without training, using adaptive entropy-based uncertainty estimation to select confident augmented views. Use when the user wants to benchmark on ImageNet & Cross-Domain Benchmarks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Adult EvalA

This benchmark evaluates income prediction models for fairness regarding demographic attributes like race and gender. It probes prediction stability under demographic perturbations (individual fairness) and measures equity in true positive rates across protected groups (group fairness). Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports Balanced Accuracy (BA).

researchpythonperformance
0
3
Advbench Asr EvalA

This benchmark probes an LLM's susceptibility to jailbreak attacks by measuring how often it generates harmful or policy-violating responses when prompted with malicious objectives. It evaluates both the raw success rate of bypassing safety filters and the relative severity of the generated harmful content through pairwise ranking. Use when the user wants to benchmark on AdvBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythonperformance
0
3
Adversarial Defense EvalA

Evaluates the robustness, seamlessness, and general utility of LLMs against adversarial inputs (jailbreaks, toxicity, hallucinations, bias) using an inference-time defense framework. Use when the user has predictions and gold and needs to compute robustness score.

researchpythongo
0
3
Adversarial Nibbler EvalA

This benchmark evaluates the robustness of text-to-image models against implicitly adversarial prompts—subtle, non-obvious text inputs that bypass automated safety filters to generate harmful images. It probes the gap between human safety perception and machine safety classification, highlighting long-tail failure modes and context-dependent vulnerabilities in generative AI. Use when the user wants to benchmark on Nibbler, or asks about evaluating this task. Reports false_negative_rate.

researchpythongo
0
3
Adversarial Nli EvalA

Evaluates natural language inference models on adversarially crafted examples designed to expose reasoning brittleness and spurious pattern reliance. It probes whether models can generalize to novel, difficult inference cases that specifically target known model weaknesses across iterative rounds of human-and-model-in-the-loop data collection. Use when the user wants to benchmark on ANLI, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Adversarial Ood Robustness EvalA

Evaluates the adversarial and out-of-distribution (OOD) robustness of LLMs across sentiment analysis, natural language inference, and domain-specific classification tasks. It measures how well models maintain performance under adversarial attacks and distribution shifts, and tests the effectiveness of prompt-based enhancement strategies (AHP and ICR). Use when the user wants to benchmark on PromptRobust (SST-2), AdvGlue++, FlipKart, DDXPlus, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Adversarial Rc EvalA

This evaluation probes a model's ability to answer reading comprehension questions under adversarial conditions, specifically testing generalization across datasets constructed by progressively stronger language models. It measures how well models trained on challenging, model-in-the-loop generated questions can handle both adversarial and standard benchmarks. Use when the user wants to benchmark on SQuAD, BiDAF-adversarial, BERT-adversarial, RoBERTa-adversarial, DROP, Natural Questions, or a...

researchpythongo
0
3
Adversarial Robustness EvalA

Evaluates the robustness of image classification models against adversarial perturbations and natural distribution shifts. It measures how well a model maintains prediction accuracy on clean data while recovering performance on out-of-distribution or adversarially attacked inputs. Use when the user wants to benchmark on MNIST, CIFAR10, ImageNet, or asks about evaluating this task. Reports Relative Robustness (RR).

datapythongo
0
3
Adversarial Text Attack EvalA

Evaluates the robustness of BERT-based text classifiers against word-level adversarial attacks by measuring how well perturbed inputs maintain semantic meaning and syntactic structure while successfully flipping model predictions. It compares three attack methods across three standard classification benchmarks to determine the optimal balance between attack success, semantic preservation, and computational efficiency. Use when the user wants to benchmark on IMDB, AG News, SST2, or asks about ...

researchpythongo
0
3
Advrace EvalA

Evaluates the robustness of machine reading comprehension models against four adversarial perturbations (AddSent, CharSwap, Distractor Extraction, and Distractor Generation) applied to passages, questions, and answer options. It measures how much model accuracy degrades when faced with label-preserving but semantically altered inputs compared to clean data. Use when the user wants to benchmark on AdvRACE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ae Nerf 3d Manipulation EvalA

Evaluates a model's ability to reconstruct 3D objects from single 2D images and disentangle/manipulate specific 3D attributes (shape, appearance, camera pose) while preserving high visual fidelity. Use when the user wants to benchmark on CARLA, Photoshapes, or asks about evaluating this task. Reports FID.

researchpythonperformance
0
3
Aepc Qa EvalA

This benchmark evaluates large language models' ability to retain and apply specialized domain knowledge in Quebec's civil law insurance sector. It probes both closed-book knowledge retention and the effectiveness of retrieval-augmented generation (RAG) pipelines under jurisdiction-specific, high-stakes regulatory scenarios. Use when the user wants to benchmark on AEPC-QA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Aeria Edge Ai EvalA

Evaluates an auction-based dynamic pricing and resource allocation mechanism for on-demand DNN inference at the edge. It probes the system's ability to jointly optimize model partitioning, pricing, and resource distribution under varying user requirements and real-world trace-driven conditions. Use when the user wants to benchmark on Multi30K, ImageNet-1K, CIFAR-100, CIFAR-10, Shanghai Telecom, or asks about evaluating this task. Reports revenue.

researchpythonperformance
0
3
Aerial D Res EvalA

This benchmark evaluates a model's ability to perform referring expression segmentation on aerial imagery, testing its capacity to localize objects or regions based on natural language instructions. It specifically probes robustness to domain-specific challenges such as densely packed targets, varying object scales, and simulated historical image degradation (monochrome, sepia, and grainy conditions). Use when the user wants to benchmark on Aerial-D, RRSIS-D, NWPU-Refer, RefSegRS, Urban1960Sa...

researchpythongo
0
3
Aeropath Airway Segmentation EvalA

This benchmark evaluates 3D medical image segmentation models on contrast-enhanced CT scans containing severe airway pathologies. It probes a model's ability to accurately segment complex, distorted bronchial trees and maintain topological completeness down to small airway generations despite anatomical anomalies like tumors and emphysema. Use when the user wants to benchmark on AeroPath, or asks about evaluating this task. Reports DSC.

researchpythongit
0
3
Aerr Continuous EvalA

Evaluates a model's ability to recognize spontaneous apparent emotional reactions from video by predicting continuous arousal and valence dimensions per frame. Use when the user wants to benchmark on SEWA, RECOLA, or asks about evaluating this task. Reports ccc.

researchpythongo
0
3
Afri Mcqa EvalA

Evaluates multimodal large language models' ability to answer visual questions about African cultural contexts in both native African languages and English, across text and audio input modalities. Use when the user wants to benchmark on Afri-MCQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
African Llm Benchmark EvalA

This evaluation probes the cross-lingual reasoning and domain knowledge capabilities of large language models across low-resource African languages. It measures how well models perform on translated benchmarks compared to English, and assesses the impact of cultural appropriateness and fine-tuning data quality on model accuracy. Use when the user wants to benchmark on Winogrande, MMLU (Clinical Sections), Belebele, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Afrimte EvalA

Evaluates machine translation quality for under-resourced African languages using human-annotated Direct Assessment (DA) scores and error-span annotations. It probes a model's ability to preserve meaning across 13 diverse language pairs. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports Direct Assessment (DA) score.

researchpythongo
0
3
Afrimteb EvalA

Evaluates text embedding models on a wide range of African language tasks, including classification, retrieval, semantic similarity, clustering, and bitext mining. It probes cross-lingual transfer, language coverage, and the ability of embeddings to capture semantic and discriminative signals across 59 African languages. Use when the user wants to benchmark on AfriMTEB, AfriMTEB-Lite, or asks about evaluating this task. Reports macro average score.

researchpythongo
0
3
Afriqa EvalA

Evaluates cross-lingual open-retrieval question answering systems across 10 African languages. It probes the pipeline's ability to translate low-resource queries, retrieve relevant passages, and accurately extract or generate answers. Use when the user wants to benchmark on AFRIQA, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Afrisenti EvalA

Evaluates multilingual and cross-lingual sentiment classification capabilities on low-resource African languages using Twitter data. It probes how well pre-trained language models handle dialectal variation, code-switching, and mixed scripts in fine-tuning and zero-shot transfer settings. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Afrisenti Sentiment EvalA

Evaluates sentiment classification capabilities across 14 low-resource African languages using Twitter data. Tests both monolingual and multilingual transfer, as well as zero-shot adaptation via parameter-efficient fine-tuning. Use when the user wants to benchmark on AfriSenti, or asks about evaluating this task. Reports weighted F1 score.

researchpythongo
0
3
Afrispeech 200 EvalA

Evaluates automatic speech recognition (ASR) models on pan-African accented English speech across clinical and general domains. It probes out-of-distribution generalization, zero-shot performance on unseen accents, and domain-specific robustness. Use when the user wants to benchmark on AfriSpeech-200, or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Afro Nmt EvalA

Evaluates neural machine translation performance across five low-resource African languages (Swahili, Amharic, Tigrigna, Oromo, Somali) paired with English. It probes model robustness to domain shifts and compares single-language, semi-supervised, transfer-learning, and multilingual training strategies. Use when the user wants to benchmark on AfroNMT, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Afromt EvalA

This benchmark evaluates machine translation capabilities across eight morphologically rich African languages translated from English. It specifically probes how well models handle complex morphosyntactic features like noun classification and verb extensions in low-resource settings. Use when the user wants to benchmark on AFROMT, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Agad Medical Auc EvalA

Evaluates a generative anomaly detection model's ability to distinguish normal from abnormal medical images using pseudo-anomaly generation and self-contrast learning. It probes robustness on fine-grained, real-world medical imaging data with limited anomaly supervision. Use when the user wants to benchmark on Alzheimer's Dataset Dubey (2019), ChestXray Kermany et al. (2018), Lung Histopathology (LC25000 subset), Retinal OCT Kermany et al. (2018), or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Age Prediction Fairness EvalA

Evaluates the accuracy and demographic fairness of deep learning models for age prediction from facial images. It probes the model's ability to generalize across diverse camera settings and ethnicities/genders while mitigating bias through distribution-aware curation and augmentation. Use when the user wants to benchmark on APPA-REAL, MORPH-2, UTKFace, Mega Asian, AFAD, CACD, or asks about evaluating this task. Reports MAE.

researchpythonapi
0
3
Agent Red Teaming EvalA

Evaluates the security and robustness of frontier AI agents against adversarial prompt injections and policy violations across multiple models and realistic deployment scenarios. It measures how effectively attacks transfer between models and whether model capability or inference compute correlates with safety. Use when the user wants to benchmark on Agent Red Teaming (ART) benchmark, or asks about evaluating this task. Reports ASR.

researchpythongo
0
3
Agent Spec EvalA

Evaluates the cross-framework portability and reusability of declarative agent specifications by executing identical agentic workflows across four different runtime frameworks (AutoGen, CrewAI, LangGraph, WayFlow) on three distinct task benchmarks. Use when the user wants to benchmark on SimpleQA Verified, BIRD-SQL, $\tau^{2}$-Bench, or asks about evaluating this task. Reports F1 score, EX%, Passˆk.

researchpythongo
0
3
Agentbench EvalA

Evaluates LLMs as autonomous agents across eight diverse, real-world environments requiring multi-turn interaction, long-term reasoning, decision-making, and strict instruction following. The benchmark measures success rates across code, game, and web-based tasks to identify performance gaps between commercial and open-source models. Use when the user wants to benchmark on AgentBench, or asks about evaluating this task. Reports overall_score.

ai-agentspythongit
0
3
Agentboard EvalA

Evaluates an agent's ability to complete multi-round interactive tasks and achieve target goals across diverse task types. It simulates real-world environments where the model must navigate sequential decision-making to reach a defined endpoint. Use when the user wants to benchmark on AgentBoard, or asks about evaluating this task. Reports target achievement rate.

researchpythongo
0
3