All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,229 views
Biwi Head Pose EvalA

Evaluates monocular head pose estimation accuracy by predicting 6DoF rotation (yaw, pitch, roll) from RGB images, comparing absolute single-image regression against relative two-view transformation prediction. Use when the user wants to benchmark on BIWI Kinect Head Pose Database, or asks about evaluating this task. Reports MAE.

researchpythondatabase
0
3
Bizbench EvalA

Evaluates large language models' ability to perform quantitative reasoning in business and finance, specifically focusing on synthesizing executable code from financial questions, extracting numeric values from structured and unstructured data, and applying domain-specific financial knowledge. Use when the user wants to benchmark on BizBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Bizfinbench EvalA

Evaluates LLMs on real-world financial reasoning tasks, including numerical calculation, temporal reasoning, information extraction, prediction recognition, and knowledge-based QA in Chinese. It probes the models' ability to handle noisy, context-dependent financial data and produce structured, reasoned outputs. Use when the user wants to benchmark on BizFinBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Bj Benchmark EvalA

This benchmark evaluates vision-language and large language models on clinical reasoning for musculoskeletal disorders. It probes capabilities ranging from medical knowledge recall and unimodal interpretation to open-ended multimodal diagnosis, treatment planning, and text-image inconsistency detection. The protocol highlights the performance gap between structured multiple-choice questions and complex, free-form clinical reasoning tasks. Use when the user wants to benchmark on B&J benchmark,...

researchpythongo
0
3
Bkdfedgnn EvalA

Evaluates the vulnerability of Federated Graph Neural Networks (FedGNN) to classification backdoor attacks across node-level and graph-level tasks. It systematically measures how global factors (data distribution, attacker count, attack timing, overlap) and local factors (trigger size, type, position, poisoning rate) influence attack success and transferability to clean clients. Use when the user wants to benchmark on Unspecified (13 datasets across 6 domains), or asks about evaluating this t...

researchpythonnode
0
3
Black Box Attribution EvalA

Evaluates the faithfulness of black-box attribution methods by measuring how well identified input regions align with the model's decision-making process. It tests the ability of explanation algorithms to pinpoint critical features that drive correct predictions or cause errors. Use when the user wants to benchmark on ImageNet, CUB-200-2011, CelebA, VGG-Face2, LC25000 (Lung), VGG-Sound, or asks about evaluating this task. Reports Deletion AUC, Insertion AUC.

researchpythongo
0
3
Black Box Llm Granularity EvalA

This evaluation probes a black-box LLM's ability to generate high-cardinality, continuous probability scores for binary classification tasks. It measures how well different prompting and post-processing methods improve operational granularity (control over precision-recall operating points) while maintaining predictive performance. Use when the user wants to benchmark on 11 binary classification datasets (combined into a joint dataset for one experiment), or asks about evaluating this task. R...

researchpythonapi
0
3
Blackswan EvalA

Evaluates vision-language models on abductive and defeasible reasoning in videos depicting unpredictable events. It probes the ability to infer hidden causes from limited visual cues and revise hypotheses when new evidence emerges, testing reasoning beyond simple statistical recall. Use when the user wants to benchmark on BlackSwanSuite, or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Blair Retrieval Recommendation EvalA

Evaluates a model's ability to align natural language reviews with item metadata for downstream recommendation and search tasks. It probes sequential next-item prediction, conventional keyword-based product retrieval, and complex long-context product search. Use when the user wants to benchmark on Amazon REVIEWS 2023 (Beauty, Games, Baby), ESCI, Amazon-C4, or asks about evaluating this task. Reports NDCG@10.

researchpythongit
0
3
Blas Performance BenchmarkA

Evaluates the runtime performance and throughput of a C++ expression template library (SALT) against optimized BLAS implementations (Intel MKL) and other template libraries (Eigen) for standard vector operations. Use when the user has predictions and gold and needs to compute performance (GFLOPS).

researchpythongo
0
3
Blast Forecasting EvalA

Evaluates the zero-shot forecasting capability of universal time series models across multiple domains and prediction horizons. It probes how well pre-trained models generalize to unseen datasets and measures predictive accuracy using standard error metrics. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, GlobalTemp, GIFT-Eval, or asks about evaluating this task. Reports MSE, MAE.

researchpython
0
3
BleurtA

Evaluates the correlation between automatic text generation scores and human quality ratings. Probes a model's ability to accurately rank or score generated translations and data-to-text outputs against human judgments, including robustness to domain/quality drift and few-shot adaptation. Use when the user has predictions and gold and needs to compute Kendall's Tau ($\tau$).

researchpythongo
0
3
BleuscoreA

Compute the BLEUScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BLEUScore, or asks how to score with BLEUScore.

documentationpython
0
3
Blimp EvalA

This benchmark probes language models' sensitivity to grammatical acceptability contrasts across 12 linguistic phenomena. It evaluates whether models can reliably distinguish acceptable sentences from minimally ungrammatical ones, revealing strengths in morphological agreement and weaknesses in complex syntactic and semantic constraints. Use when the user wants to benchmark on BLiMP, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Blimp Glue Superglue EvalA

Evaluates language understanding, linguistic acceptability, sentiment analysis, natural language inference, and factual reasoning under low-resource fine-tuning conditions. Use when the user wants to benchmark on BLiMP, GLUE (subset), SuperGLUE (subset), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Blink Vision Centric EvalA

This evaluation probes a model's ability to leverage raw visual representations for vision-centric tasks without relying on language priors or domain expertise. It tests pixel-level matching, depth perception, 3D object awareness, and art style recognition across multiple-choice and regression-style tasks. Use when the user wants to benchmark on CV-Bench (Depth Order), SPair-71k, FunKPoint, HPatches, MOCHI, WikiArt (BLINK Art Style), or asks about evaluating this task. Reports multiple-choice...

researchpythongo
0
3
Blinkflow EvalA

Evaluates the accuracy and robustness of event-based optical flow estimation models on complex scenes with dynamic objects, occlusions, and high-frequency motion. It measures how well models generalize to unseen scenarios and handle fine-grained flow details compared to traditional rigid/static scene benchmarks. Use when the user wants to benchmark on BlinkFlow, DSEC, MVSEC, or asks about evaluating this task. Reports AEE.

researchpythonangular
0
3
Blip3 Multimodal Benchmarks EvalA

Evaluates large multimodal models on single- and multi-image understanding, covering general VQA, domain knowledge, OCR, hallucination, and interleaved image-text reasoning. Use when the user wants to benchmark on SEED-IMG, MMB(dev), MMStar, MME(norm), RWQA, MMVet, MMMU(val), MathVista, TextVQA, OCRBench, POPE, HalBench, BLINK, QBench, MuirBench, Mantis-Eval, or asks about evaluating this task. Reports benchmark score, average score.

researchpythongo
0
3
Blizzard Indian G2p Dur EvalA

Evaluates DNN-based grapheme-to-phoneme and duration prediction models for three Indian languages (Hindi, Tamil, Telugu) using crowdsourced ASCII transliterated text. It measures the accuracy of predicted phoneme durations against forced-aligned ground truth to assess component quality for speech synthesis. Use when the user wants to benchmark on Blizzard Challenge 2015 (Hindi, Tamil, Telugu), or asks about evaluating this task. Reports RMSE (frames per phone).

researchpythongo
0
3
Blmore Challenge EvalA

Evaluates multimodal emotion recognition systems on detecting blended emotions (presence and salience) across unseen actors. It probes the model's ability to generalize actor-invariant emotional semantics from audio, visual, and combined modalities under a strict threshold-based discretization protocol. Use when the user wants to benchmark on BLEMORE, or asks about evaluating this task. Reports Score.

researchpythonexpress
0
3
Bloom Empirical EvalA

Evaluates BLOOM model variants against BERT-style and GPT-style baselines across diverse NLP tasks including text classification, question answering, zero/few-shot learning, multilingual transfer, and text generation. Use when the user wants to benchmark on GLUE, SQuAD, XNLI, MARC, Zero/FSL Benchmarks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Blue EvalA

Evaluates the cross-domain generalization and transfer learning capabilities of pre-trained language models across ten diverse biomedical and clinical NLP tasks. It probes sentence similarity, named entity recognition, relation extraction, document classification, and natural language inference to measure how well domain-specific pre-training captures clinical and biomedical semantics. Use when the user wants to benchmark on MedSTS, BIOSSES, BC5CDR-disease, BC5CDR-chemical, ShARe/CLEFE, DDI, ...

researchpythongo
0
3
Blurb EvalA

Evaluates biomedical language models on a comprehensive suite of downstream NLP tasks, including named entity recognition, relation extraction, sentence similarity, document classification, and question answering. It measures how well domain-specific pretraining transfers to specialized clinical and biomedical text understanding. Use when the user wants to benchmark on BLURB, or asks about evaluating this task. Reports BLURB score.

researchpythongo
0
3
Bmdal Regression EvalA

Evaluates the sample efficiency and predictive accuracy of batch-mode deep active learning methods for tabular regression tasks. It probes how well different kernel-based selection strategies reduce prediction error over sequential labeling rounds compared to random sampling. Use when the user wants to benchmark on UCI & OpenML Tabular Regression Benchmark, or asks about evaluating this task. Reports RMSE.

researchpythongit
0
3
Bnn Nvm Benchmark EvalA

This benchmark evaluates the inference accuracy and hardware performance of binary neural networks (BNNs) deployed on non-volatile memory crossbar architectures. It probes how hardware constraints like ADC resolution and first-layer input precision affect model accuracy, latency, energy efficiency, and chip area. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports accuracy.

businesspythongo
0
3
Bntt Snn EvalA

Evaluates the classification accuracy, inference latency, and energy efficiency of Spiking Neural Networks (SNNs) trained with Batch Normalization Through Time (BNTT) on standard image and neuromorphic datasets. It probes the model's ability to maintain high accuracy while drastically reducing time-steps and computational cost compared to ANN-SNN conversion and standard surrogate gradient methods. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, DVS-CIFAR10, or asks...

researchpython
0
3
Bomjin Code Eval OctopackA

Compute bomjin/code_eval_octopack via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bomjin/code_eval_octopack.

developmentpython
0
3
Bones EvalA

Evaluates the accuracy and computational efficiency of neural and traditional Shapley value estimators against ground truth attributions across tabular and image datasets. It measures how well different explainers approximate feature importance and how fast they run. Use when the user wants to benchmark on Monks, WBC, Census, Credit, Magic, ImageNette, Pet, or asks about evaluating this task. Reports L1 distance.

researchpythongo
0
3
Bongard Rwr Plus EvalA

Evaluates vision-language models' ability to perform abstract visual reasoning (AVR) by recognizing fine-grained, abstract visual concepts in Bongard-style matrix problems. It probes capabilities in concept selection, image-to-side classification, and free-form concept description generation. Use when the user wants to benchmark on Bongard-RWR+, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Booksum EvalA

Evaluates extractive and abstractive summarization models on long-form narrative texts across paragraph, chapter, and book granularities. It probes lexical overlap, semantic similarity, content coverage via question answering, and human-rated fluency, coherence, relevance, and factuality. Use when the user wants to benchmark on BookSum, or asks about evaluating this task. Reports ROUGE-1.

researchpythongit
0
3
Boolq EvalA

This benchmark evaluates a model's ability to answer naturally occurring yes/no questions based on a provided passage. It probes complex inferential reasoning and non-factoid inference, requiring the model to go beyond simple keyword matching or shallow statistical features to determine entailment or contradiction between the question and the passage. Use when the user wants to benchmark on BoolQ, or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Boom Ood EvalA

Evaluates the out-of-distribution (OOD) generalization of machine learning models for molecular property prediction. It probes whether models trained on in-distribution (ID) molecules can accurately extrapolate to novel chemical spaces, highlighting the disconnect between ID accuracy and OOD robustness. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports RMSE.

datapythongo
0
3
Booster EvalA

Evaluates stereo and monocular depth/disparity estimation models on images containing specular and transparent surfaces, which violate standard non-Lambertian assumptions and cause significant performance degradation in existing networks. Use when the user wants to benchmark on Booster, or asks about evaluating this task. Reports bad-2.

researchpythonperformance
0
3
Booststep EvalA

This evaluation probes the mathematical reasoning capability of large language models, specifically focusing on single-step reasoning and the effectiveness of step-aligned in-context learning. It measures how well models can solve challenging math problems across text and multi-modal domains when provided with fine-grained, step-level examples. Use when the user wants to benchmark on MATH500, AQuA, OlympiadBench-TO, MATHBench, AMC-10, AMC-12, MathVision, MathVerse, AIME, or asks about evaluat...

researchpythongo
0
3
Bootstrap3d EvalA

Evaluates text-to-multi-view diffusion models on their ability to generate prompt-aligned, high-quality 4-view images and reconstruct consistent 3D objects. It measures image-text alignment and visual fidelity against a synthetic ground-truth distribution. Use when the user wants to benchmark on GPTeval3D, Synthetic GT Distribution, or asks about evaluating this task. Reports FID.

researchpython
0
3
BootstrapperA

Compute the BootStrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BootStrapper, or asks how to score with BootStrapper.

documentationpython
0
3
Bop 6d Pose Refinement EvalA

Evaluates the ability to refine 6D object poses in cluttered real-world scenes using RGB or RGB-D inputs. It probes generalization to novel objects by measuring pose accuracy against ground truth under symmetry-aware error metrics. Use when the user wants to benchmark on LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HomebrewdDB, YCB-V, or asks about evaluating this task. Reports Average Recall (AR).

researchpythonperformance
0
3
Borehole Segmentation EvalA

This evaluation probes a model's ability to perform weakly supervised multimodal segmentation of acoustic borehole images by refining threshold-guided pseudo-labels using depth-aligned well logs. It measures how well the predicted segmentation aligns with a provisional target map, testing spatial coherence and multimodal feature fusion rather than absolute geological accuracy. Use when the user wants to benchmark on Antilope25 & Botorosa47 borehole intervals, or asks about evaluating this tas...

researchpythontesting
0
3
Botfails EvalA

This evaluation probes a model's ability to detect and classify robotic failures in real-world manipulation tasks. It specifically tests whether a system can distinguish between genuine task-disrupting failures and benign environmental deviations using multimodal observations and nominal demonstrations. Use when the user wants to benchmark on BotFails, Real-π dataset, or asks about evaluating this task. Reports AUROC.

researchpython
0
3
Bots Dvfs EvalA

Evaluates a hierarchical multi-agent reinforcement learning scheduler's ability to optimize task allocation, frequency scaling, and core selection for OpenMP DAG workloads on embedded systems. It probes the trade-off between makespan, energy consumption, and thermal constraints under real-time profiling feedback. Use when the user wants to benchmark on Barcelona OpenMP Tasks Suite (BOTS), or asks about evaluating this task. Reports makespan.

researchpythonperformance
0
3
Bounded Nbeddyn EvalA

This evaluation probes a model's ability to forecast and reconstruct the dynamics of partially observed, chaotic geophysical systems. It specifically tests short-term prediction accuracy and long-term topological stability (boundedness) under both in-distribution and out-of-attractor initial conditions. Use when the user wants to benchmark on Lorenz-63, Lorenz-96, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Bouquet Mt EvalA

Evaluates machine translation systems on a contamination-free, multilingual dataset covering diverse domains and registers. It measures translation quality at both sentence and paragraph levels to assess how well models handle linguistic diversity and cultural authenticity across 8 major languages. Use when the user wants to benchmark on BOUQuET, or asks about evaluating this task. Reports CometKiwi.

researchpython
0
3
Bowdbeg DocredA

Compute bowdbeg/docred via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bowdbeg/docred.

developmentpython
0
3
Bowdbeg Matching SeriesA

Compute bowdbeg/matching_series via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bowdbeg/matching_series.

developmentpython
0
3
Bowdbeg Patch SeriesA

Compute bowdbeg/patch_series via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bowdbeg/patch_series.

developmentpython
0
3
Bpmn Structured Extraction EvalA

Evaluates vision-language models' ability to extract structured information (names, types, and connectivity) from Business Process Model and Notation (BPMN) diagrams provided as images. It tests both raw visual understanding and the utility of OCR-enriched inputs for schema-constrained diagram parsing. Use when the user wants to benchmark on BPMN Diagrams (Custom), or asks about evaluating this task. Reports F1 Score.

researchpythongo
0
3
BraTS PLGG Classification EvalA

Evaluates the effectiveness of synthetic 3D MRI tumor ROI generation for data augmentation by measuring downstream binary classification performance on imbalanced brain tumor subtypes. Use when the user wants to benchmark on BraTS 2019, SickKids pLGG, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Brace Hallucination EvalA

Probes models' robustness in detecting subtle hallucinations in audio captions, specifically those introduced via LLM-driven noun substitution. It measures the ability to identify semantically flawed or factually incorrect descriptions against audio ground truth. Use when the user wants to benchmark on BRACE-Hallucination, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Brace Main EvalA

Evaluates the ability of audio-language models to align audio with captions and distinguish caption quality across different generation sources (human-human, human-machine, machine-machine). It probes fine-grained semantic and syntactic alignment capabilities under realistic captioning conditions. Use when the user wants to benchmark on BRACE-Main, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Bradley TerryA

Evaluates the stability and reliability of global pointwise scores (accuracy, AUC, F1) versus pairwise Bradley-Terry rankings for ordering NLP models across classification and text generation tasks. Use when the user has predictions and gold and needs to compute Bradley-Terry.

researchpythongo
0
3