All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs1,997 views
Bcs Dbt Classification EvalA

Evaluates the ability of self-supervised contrastive pre-training and multi-patch fine-tuning to classify imbalanced digital breast tomosynthesis (DBT) slices and volumes as normal or abnormal. It probes the model's robustness to extreme class imbalance and its capacity to preserve spatial resolution through patch-level processing. Use when the user wants to benchmark on BCS-DBT, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Bdsaglam JerA

Compute bdsaglam/jer via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bdsaglam/jer.

developmentpython
0
3
Beads EvalA

This benchmark evaluates language models across multiple tasks to detect, quantify, and mitigate demographic and social biases. It probes classification accuracy for bias/toxicity/sentiment, token-level bias identification, demographic stereotype alignment, and the ability to generate neutral, benign text variants. Use when the user wants to benchmark on BEADs, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Beam Control EvalA

Evaluates real-time reinforcement learning for particle beam steering on edge SoCs. It probes the ability to maximize control reward while adhering to strict millisecond latency and cycle-rate constraints. Use when the user wants to benchmark on Fermilab Booster Synchrotron Dataset, or asks about evaluating this task. Reports Reward (R).

datapythongo
0
3
Beans EvalA

This benchmark evaluates machine learning models on bioacoustic animal sound recognition across 12 diverse datasets spanning birds, mammals, anurans, and insects. It probes two core capabilities: multi-label species classification and temporal sound detection, testing models' ability to generalize across species and handle varying recording conditions and class imbalances. Use when the user wants to benchmark on wtkn, bat, cbi, hbdb, dogs, dcase, enabirds, hiceas, rfcx, hainan-gibbons, esc, s...

datapythongo
0
3
Beans Zero EvalA

Evaluates zero-shot generalization of audio-language models on bioacoustic tasks, including species classification, multilabel detection, call-type prediction, lifestage classification, captioning, and individual counting across diverse taxa. Use when the user wants to benchmark on BEANS-Zero, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Beard EvalA

Evaluates the adversarial robustness of models trained on synthetically distilled datasets. It probes how well different dataset distillation methods preserve model resilience against diverse adversarial attacks across varying image-per-class (IPC) settings. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, TinyImageNet, or asks about evaluating this task. Reports Comprehensive Robustness-Efficiency Index (CREI).

researchpythongit
0
3
Bearllm Fault Diagnosis EvalA

Evaluates a multimodal LLM framework's ability to perform bearing fault diagnosis, anomaly detection, and maintenance recommendation using vibration signals and textual prompts. It probes cross-condition generalization and zero-shot transfer across diverse industrial bearing datasets. Use when the user wants to benchmark on MBHM, JUST, IMS, CWRU, XJTU, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Beat Backdoor Detection EvalA

Evaluates the ability of a black-box defense mechanism to detect backdoor-unaligned samples in LLMs by measuring changes in the model's refusal behavior when a malicious probe is concatenated to the input. Use when the user wants to benchmark on MaliciousInstruct + Advbench + UltraChat-200k, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Beat It EvalA

Evaluates a model's ability to generate 3D dance motions that are temporally synchronized with musical beats and controllable via sparse keyframes, while maintaining kinematic plausibility and motion diversity. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports BAS.

researchpythontesting
0
3
Beatv2 EvalA

Evaluates cross-dataset generalization of co-speech gesture generation on a standard English benchmark, measuring gesture quality, beat consistency, and diversity. Use when the user wants to benchmark on BEATv2, or asks about evaluating this task. Reports FGD.

researchpythonexpress
0
3
Beaver EvalA

Probes LLMs' ability to generate correct SQL queries from natural language questions over complex, enterprise-scale databases. It specifically evaluates handling of high schema complexity, multi-table joins, aggregations, and column-to-table mapping in real-world business contexts. Use when the user wants to benchmark on BEAVER, or asks about evaluating this task. Reports execution accuracy.

researchpythongo
0
3
Beavertails Moderation EvalA

This evaluation probes the safety moderation and context-comprehension capabilities of external text moderation APIs. It measures how well automated systems align with human and expert-preference labels when assessing harmfulness across specific risk categories in QA pairs. Use when the user wants to benchmark on BeaverTails Evaluation Dataset, or asks about evaluating this task. Reports agreement.

researchpythongo
0
3
Bee 8b Mllm EvalA

Evaluates the visual reasoning, factual accuracy, OCR, chart understanding, and mathematical capabilities of fully open multimodal large language models (MLLMs) against a comprehensive suite of established benchmarks. The protocol tests the model's ability to process images and text prompts, generate responses in a thinking mode, and achieve high scores across general VQA, document/chart analysis, and complex math/reasoning tasks. Use when the user wants to benchmark on Bee-8B Evaluation Benc...

researchpythongo
0
3
Beep Korean Toxic Speech EvalA

This benchmark evaluates models' ability to detect social bias (gender and other types) and hate speech in Korean online news comments. It probes whether models can distinguish between hate speech, offensive language, and neutral comments, and whether incorporating bias labels improves hate speech detection. Use when the user wants to benchmark on BEEP!, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Beerqa EvalA

Evaluates open-domain question answering systems on their ability to retrieve and synthesize information across varying numbers of reasoning steps (single-hop to three-hop) without relying on structured metadata or predefined hop counts. Use when the user wants to benchmark on SQuAD Open, HotpotQA, BeerQA, or asks about evaluating this task. Reports exact match (EM).

researchpythongo
0
3
Beexai EvalA

Evaluates post-hoc explainable AI (XAI) attribution methods on tabular data across binary classification, multi-class classification, and regression tasks. It measures how well feature importance scores align with core XAI desiderata—faithfulness, plausibility, robustness, and complexity—using ground-truth-aligned quantitative metrics. Use when the user wants to benchmark on inria-soda/tabular-benchmark, OpenML-CC18 Curated Classification, or asks about evaluating this task. Reports Infidelity.

researchpythongo
0
3
Behavior 1k EvalA

This evaluation probes a robot's ability to execute long-horizon, zero-shot rearrangement tasks in unexplored indoor-outdoor environments using grounded language reasoning. It measures how well the system interprets natural language instructions, reasons over 3D scene graphs, and coordinates sequential manipulation actions to satisfy multiple goal conditions. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Behavior1k EvalA

Evaluates embodied AI agents on long-horizon, human-centered manipulation tasks in a realistic physics-based simulation. It probes the agent's ability to plan and execute complex sequences of action primitives (pick, place, navigate, etc.) while handling rigid, articulated, and deformable objects. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports task success rate.

researchpythongo
0
3
Behavioral Fraud Patterns EvalA

This benchmark evaluates whether synthetic tabular data generators preserve complex behavioral fraud patterns beyond simple statistical fidelity. It specifically probes temporal inter-event timing, burst-like transaction sequences, shared-infrastructure graph motifs across accounts, and velocity-rule trigger rates. Use when the user wants to benchmark on IEEE-CIS Fraud Detection, Amazon Fraud Dataset, or asks about evaluating this task. Reports P1: IET W1.

researchpythongo
0
3
Behavioral Prediction EvalA

Predicts individual strategic decisions by conditioning on structured psychometric trait profiles (e.g., Big Five personality traits). It probes a model's ability to map high-dimensional psychological embeddings to discrete behavioral outcomes in unseen situational contexts. Use when the user wants to benchmark on Strategic Scenario Dataset, or asks about evaluating this task. Reports balanced accuracy, macro-F1.

researchpythonperformance
0
3
Being H05 Robot EvalA

Evaluates cross-embodiment generalization and manipulation capabilities of Vision-Language-Action models across heterogeneous real robots and simulation benchmarks. Probes spatial reasoning, long-horizon planning, bimanual coordination, and zero-shot transfer to unseen task-embodiment pairs. Use when the user wants to benchmark on Real-robot task suite, LIBERO, RoboCasa, or asks about evaluating this task. Reports success rate (%).

researchpythonperformance
0
3
Beir EvalA

Evaluates zero-shot information retrieval capabilities across 18 diverse domains and query types. It probes a model's ability to retrieve relevant documents without domain-specific fine-tuning, highlighting performance variations due to domain shifts, query length, and lexical versus semantic matching. Use when the user wants to benchmark on BEIR, or asks about evaluating this task. Reports nDCG@10.

researchpythongo
0
3
Beir Nl EvalA

Evaluates zero-shot information retrieval capabilities of lexical, dense, and reranking models on Dutch-language queries and documents. It probes how well models generalize to a machine-translated benchmark without fine-tuning, measuring both ranking quality and recall performance. Use when the user wants to benchmark on MSMARCO, TREC-COVID, NFCorpus, NQ, HotpotQA, FiQA-2018, ArguAna, Touche-2020, CQADupstack, Quora, DBPedia, SciDocs, SciFact, FEVER, Climate-FEVER, or asks about evaluating th...

researchpythongo
0
3
Bekhouche AccA

Compute Bekhouche/ACC via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Bekhouche/ACC.

developmentpython
0
3
Bekhouche NedA

Compute Bekhouche/NED via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Bekhouche/NED.

developmentpython
0
3
Belief Tracking Policy EvalA

Evaluates neural belief tracking models on their accuracy, calibration, and runtime efficiency, and measures how different uncertainty estimates (confidence, total uncertainty, knowledge uncertainty) affect downstream dialogue policy performance in both simulated and human-user environments. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).

researchpythongo
0
3
Bells EvalA

Evaluates LLM supervision systems and frontier models on their ability to detect harmful content across varying harm severities (benign, borderline, harmful) and adversarial sophistication levels (direct prompts vs. jailbreaks). It measures detection capability, robustness to adversarial transformations, and metacognitive coherence between harm classification and response behavior. Use when the user wants to benchmark on BELLS benchmark, or asks about evaluating this task. Reports BELLS Score.

researchpythongo
0
3
Bench Push EvalA

Evaluates the transferability and performance of reinforcement learning policies for mobile robot navigation and pushing-based manipulation tasks. It probes how well policies trained in simulation handle real-world sim-to-real gaps, clutter, and sparse rewards across varying obstacle densities. Use when the user wants to benchmark on Bench-Push (Maze & Box-Delivery), or asks about evaluating this task. Reports $S_{\text{manip}}$.

researchpythonperformance
0
3
Bench2drive Personalized Driving EvalA

This evaluation probes a vision-language-action model's ability to align autonomous driving behavior with both long-term individual driver habits and short-term natural language style instructions. It measures safety, efficiency, comfort, and stylistic fidelity in closed-loop simulation scenarios like merging, overtaking, and emergency braking. Use when the user wants to benchmark on Bench2Drive, or asks about evaluating this task. Reports Driving Score (DS).

ai-agentspythongo
0
3
Bench2drive Speed EvalA

Evaluates autonomous driving policies' ability to follow explicit user commands for target speed and overtake/follow behaviors in closed-loop simulations. It measures how well models track desired speeds and execute passing maneuvers while maintaining safety, comfort, and traffic compliance. Use when the user wants to benchmark on Bench2Drive-Speed, or asks about evaluating this task. Reports Speed-Adherence Score.

researchpythongo
0
3
Bench2drive Vl EvalA

Evaluates vision-language models in closed-loop autonomous driving by assessing their perception, prediction, planning, and behavioral reasoning capabilities within a CARLA simulator. It probes the model's ability to process raw sensor inputs, generate causally consistent natural language reasoning, and produce valid control actions across diverse and out-of-distribution driving scenarios. Use when the user wants to benchmark on Bench2Drive-VL, or asks about evaluating this task. Reports LLM-...

researchpythongo
0
3
Bench360 EvalA

Evaluates local LLM inference across multiple dimensions, including task-specific quality (e.g., accuracy, F1, ROUGE) and system-level performance (latency, throughput, energy, memory, cold-start) under simulated workloads (single-stream, batch, server). Use when the user wants to benchmark on mmlu, squad_v2, cnn_dailymail, or asks about evaluating this task. Reports accuracy.

devopspythondocker
0
3
Benchecg EvalA

Evaluates ECG foundation models on diverse clinical tasks including classification, regression, detection, and survival analysis across multiple populations and signal lengths. It probes the model's ability to generalize across datasets, modalities (ECG vs PPG), and long-context temporal dependencies. Use when the user wants to benchmark on CODE-15%, Sleep-Apnea-ECG, MIT-BIH Arrhythmia, PTB-XL, CPSC2018, MIMIC-IV-ECG, Exercise-ECG, or asks about evaluating this task. Reports BenchECG score.

researchpythongo
0
3
Benchie Fl EvalA

Evaluates Open Information Extraction (OIE) systems on their ability to extract fact-based triples from text. It uses a conservative exact-matching function with synset-based clustering to penalize non-informative copies and reward precise fact extraction, while also measuring correlation with downstream QA and knowledge base tasks. Use when the user wants to benchmark on BenchIE^FL, or asks about evaluating this task. Reports exact-match.

researchpythongo
0
3
Benchmark Accuracy EvalA

Evaluates zero-shot language model performance across a suite of 10 standard NLP benchmarks covering commonsense reasoning, science QA, and language modeling. It measures task accuracy and correlates it with word-level statistical overlap metrics to assess distributional alignment between pre-training data and evaluation sets. Use when the user wants to benchmark on ARC Easy, ARC Challenge, Hellaswag, MMLU, SciQ, OpenBookQA, PIQA, lambada, SocialIQA, SWAG, or asks about evaluating this task. ...

researchpythongo
0
3
Benchmark Diversity Stability EvalA

Evaluates the inherent trade-off between diversity (agreement of model rankings across tasks) and stability (sensitivity of final rankings to label noise) in multi-task machine learning benchmarks. It quantifies how much a benchmark's leaderboard ranking changes when trivial label noise is injected, and how diverse the rankings are across its constituent tasks. Use when the user wants to benchmark on GLUE, SuperGLUE, MTEB, BigBenchHard, MMLU, OpenLLM, VTAB, ImageNet, or asks about evaluating ...

datapythongo
0
3
Benchmax EvalA

BenchMAX evaluates the language-agnostic capabilities of large language models across 17 languages, including non-Latin scripts. It probes instruction following, reasoning, code generation, long-context modeling, tool use, and translation through a rigorously translated and human-post-edited pipeline. Use when the user wants to benchmark on BenchMAX, or asks about evaluating this task. Reports evaluation metrics.

researchpythongo
0
3
Benchmd EvalA

Evaluates modality-agnostic models across 19 real-world medical datasets spanning 1D, 2D, and 3D modalities. Probes performance under data scarcity (few-shot linear evaluation and finetuning) and out-of-distribution generalization across different hospitals and data distributions. Use when the user wants to benchmark on BenchMD, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Benchread EvalA

Evaluates retinal anomaly detection models across fundus photography and OCT modalities, testing their ability to distinguish normal from abnormal images and generalize to unseen anomalies under varying supervision levels. Use when the user wants to benchmark on Fundus Benchmark, OCT Benchmark, or asks about evaluating this task. Reports AUC-ROC.

researchpythongo
0
3
Benchtemp EvalA

Evaluates the effectiveness and efficiency of Temporal Graph Neural Networks (TGNNs) on link prediction and node classification tasks. It probes model performance across transductive and inductive settings (New-Old/New-New) to ensure fair cross-model comparisons. Use when the user wants to benchmark on BenchTemp (15 datasets), or asks about evaluating this task. Reports AUC.

researchpythonnode
0
3
Benchx EvalA

Evaluates Medical Vision-Language Pretraining (MedVLP) models on chest X-ray tasks including multi-label/binary classification, segmentation, report generation, and image-text retrieval. It specifically probes how standardized preprocessing and finetuning strategies affect model performance across heterogeneous architectures. Use when the user wants to benchmark on NIH, VinDr, COVIDx, SIIM, RSNA, Object-CXR, TBX11K, IUXray, MIMIC 5x200, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Bengal Ner El EvalA

Evaluates the performance of Named Entity Recognition (NER) and Entity Linking (EL) systems on automatically generated corpora. It probes a model's ability to accurately detect entity spans in text and correctly link them to a reference knowledge base (DBpedia) across varying document lengths, entity densities, and languages. Use when the user wants to benchmark on BENGAL (B1-B13, P1-P4, S1-S4), or asks about evaluating this task. Reports micro F1-score.

researchpythongo
0
3
Bengali Asr Diarization EvalA

Evaluates automatic speech recognition (ASR) accuracy and speaker diarization performance on long-form Bengali speech. It measures phonetic robustness and computational efficiency using public and private test splits under strict hardware constraints. Use when the user wants to benchmark on Lipi-Ghor-882, or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Bengalimoralbench EvalA

Evaluates large language models' ability to perform moral reasoning and align with human ethical judgments within Bengali language and South Asian socio-cultural contexts. It probes cultural grounding, commonsense reasoning, and fairness across five everyday moral domains using native-speaker consensus annotations. Use when the user wants to benchmark on BengaliMoralBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ber PerformanceA

Evaluates the bit error rate (BER) performance of a Reconfigurable Intelligent Surface (RIS) aided spatial media-based modulation system compared to baseline schemes (SM, MBM, QSM) under uncorrelated Rayleigh fading channels. Use when the user has predictions and gold and needs to compute Bit Error Rate (BER).

researchpythongo
0
3
Berkatil MapA

Compute berkatil/map via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of berkatil/map.

developmentpython
0
3
Berkatil MrrA

Compute berkatil/mrr via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of berkatil/mrr.

developmentpython
0
3
Berst EvalA

Probes the robustness of Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER) models under challenging real-world conditions, including varying distances, physical obstructions, and high-intensity vocalizations. It specifically tests whether models can maintain accuracy when linguistic context is removed via nonsense phrases and when acoustic features are degraded by far-field recording and shouting. Use when the user wants to benchmark on BERSt, or asks about evaluating th...

researchpythongo
0
3
Bert Noise Robustness EvalA

Evaluates BERT's robustness to synthetic character-level noise across sentiment classification and textual similarity tasks. It probes how spelling mistakes and typos disrupt subword tokenization and degrade contextual embeddings under varying noise intensities. Use when the user wants to benchmark on IMDB, SST-2, STS-B, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3