All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,049 views
Bridgmanite Thermoelastic EvalA

Evaluates the accuracy of deep-learning molecular dynamics potentials in predicting the structural, thermodynamic, and elastic properties of bridgmanite (MgSiO3-perovskite) under high-pressure and high-temperature conditions relevant to Earth's lower mantle. Use when the user wants to benchmark on DFT reference dataset, Experimental benchmarks, PREM seismic model, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Brier Score LossA

Compute the brier_score_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute brier_score_loss, or asks how to score with brier_score_loss.

documentationpython
0
3
Bright EvalA

Evaluates a model's ability to perform reasoning-intensive text retrieval by matching complex, domain-diverse queries to relevant documents. It probes deep logical and conceptual alignment between queries and documents, going beyond simple keyword or semantic matching. Use when the user wants to benchmark on BRIGHT, or asks about evaluating this task. Reports nDCG@10.

researchpythongo
0
3
Brisc Segmentation EvalA

Evaluates deep learning models on multi-planar brain tumor segmentation from contrast-enhanced T1-weighted MRI scans. It probes multi-scale feature integration, cross-view generalization, and robustness to class imbalance across glioma, meningioma, and pituitary tumor types. Use when the user wants to benchmark on BRISC, or asks about evaluating this task. Reports mIoU.

researchpythongit
0
3
Browsecomp Plus EvalA

Evaluates the end-to-end effectiveness of deep-research agents in retrieving evidence and answering complex queries, as well as the standalone effectiveness of various retrievers. It probes the interplay between retrieval quality, reasoning capability, and search efficiency in agentic workflows. Use when the user wants to benchmark on BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.

ai-agentspythongo
0
3
Browser Inference EvalA

Evaluates the performance overhead and latency characteristics of running deep learning inference directly in web browsers compared to native environments. It probes the impact of WebAssembly runtime inefficiencies, SIMD limitations, and WebGL GPU abstraction on prediction, warmup, and setup phases across various models and hardware configurations. Use when the user wants to benchmark on Inference Benchmark (ResNet50, VGG16, MobileNetV2), or asks about evaluating this task. Reports prediction...

researchpythonbackend
0
3
Browsesafe Bench EvalA

Evaluates AI browser agents' ability to detect prompt injection attacks embedded in complex, realistic HTML environments. It probes whether models can distinguish malicious intent from benign distractors across diverse attack types, injection strategies, and linguistic styles. Use when the user wants to benchmark on BrowseSafe-Bench, or asks about evaluating this task. Reports balanced accuracy.

researchpythongo
0
3
Bs Gat Nf Iot EvalA

Evaluates network intrusion detection capabilities in IoT environments by classifying network flows into benign or multiple attack categories using graph-based and traditional machine learning models. Use when the user wants to benchmark on NF-BoT-IoT-v2, NF-ToN-IoT-v2, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Bsard EvalA

Evaluates the ability of information retrieval models to rank relevant Belgian statutory legal articles in response to natural language citizen questions. It probes domain-specific legal retrieval, handling the challenge of mapping unstructured queries to structured hierarchical legal texts. Use when the user wants to benchmark on BSARD, or asks about evaluating this task. Reports Recall@100.

researchpythongit
0
3
Bsc Nav EvalA

Evaluates embodied agents' spatial cognition and navigation capabilities across category-level, instance-level, and long-horizon instruction-following tasks, as well as active embodied question answering and real-world mobile manipulation. Use when the user wants to benchmark on MP3D & HM3D (Habitat Simulator), VLN-CE R2R, OpenEQA (A-EQA subset), Real-world Indoor Environment, or asks about evaluating this task. Reports Success Rate (SR).

developmentpythongo
0
3
Bscd Fsl EvalA

Evaluates cross-domain few-shot learning generalization from a source domain (ImageNet) to specialized imaging domains (agriculture, satellite, dermatology, radiology) that vary in perspective distortion, semantic content, and color depth. Use when the user wants to benchmark on CropDiseases, EuroSAT, ISIC2018, ChestX, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Bss EvalA

Evaluates speech language models' ability to process beyond-semantic speech attributes, including dialect comprehension, multi-turn context memory, emotion perception, age-aware response generation, and non-verbal cue handling. It probes whether models can integrate paralinguistic, affective, and contextual signals beyond literal semantic understanding to achieve human-like social interaction. Use when the user wants to benchmark on BoSS Evaluation Datasets, or asks about evaluating this task...

researchpythongo
0
3
Bstc EvalA

Evaluates Chinese-to-English speech translation accuracy and real-time simultaneous interpretation latency. It probes a model's ability to handle noisy ASR inputs, segment speech into meaningful units, and produce fluent translations under strict delay constraints. Use when the user wants to benchmark on BSTC, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Bstrai Classification ReportA

Compute bstrai/classification_report via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bstrai/classification_report.

developmentpython
0
3
Btzsc EvalA

Evaluates zero-shot text classification capabilities across diverse datasets using four model families: NLI cross-encoders, embedding models, rerankers, and instruction-tuned LLMs. It probes how well models can assign text to predefined categories without task-specific fine-tuning by verbalizing labels as descriptions. Use when the user wants to benchmark on BTZSC Benchmark, or asks about evaluating this task. Reports macro F1.

ai-agentspythongo
0
3
Bubbleml EvalA

This benchmark evaluates machine learning models on two distinct scientific tasks using the BubbleML dataset: predicting optical flow for bubble dynamics and solving multiphysics PDEs for temperature and velocity field propagation. It probes a model's ability to capture non-rigid object motion, sharp physical interfaces, and long-horizon temporal dynamics in phase-change simulations. Use when the user wants to benchmark on BubbleML, or asks about evaluating this task. Reports end-point error ...

datapythongit
0
3
Bucc Bitext Retrieval EvalA

Evaluates the capability of sentence embeddings to retrieve exact translation pairs (bitexts) from large monolingual corpora across different languages. It tests the model's ability to distinguish parallel sentences from semantically similar but non-parallel ones. Use when the user wants to benchmark on BUCC bitext mining task, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Bucketheadp65 Confusion MatrixA

Compute BucketHeadP65/confusion_matrix via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BucketHeadP65/confusion_matrix.

developmentpython
0
3
Bucketheadp65 Roc CurveA

Compute BucketHeadP65/roc_curve via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BucketHeadP65/roc_curve.

developmentpython
0
3
Budget Ai Researcher EvalA

Evaluates a retrieval-augmented generation framework's ability to synthesize novel, feasible, and interesting research abstracts by combining distant topics from AI conference literature. It probes long-range concept recombination and grounded ideation capabilities. Use when the user wants to benchmark on AI Conference Papers (ICLR, NeurIPS, ICML, ACL, ECCV), or asks about evaluating this task. Reports Novelty.

researchpythongo
0
3
Buelfhood Fbeta Score 2A

Compute buelfhood/fbeta_score_2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of buelfhood/fbeta_score_2.

developmentpython
0
3
Buelfhood Fbeta ScoreA

Compute buelfhood/fbeta_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of buelfhood/fbeta_score.

developmentpython
0
3
Builderbench EvalA

Evaluates an agent's ability to learn embodied reasoning, long-horizon planning, and physical/geometric intuition through self-supervised exploration, and generalizes these skills to construct unseen block structures. Use when the user wants to benchmark on BuilderBench, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Built Bench EvalA

Evaluates the ability of pre-trained text embedding models to capture domain-specific semantic alignment for built asset information. The benchmark probes clustering, information retrieval, and document reranking capabilities using technical terminology from architectural, structural, mechanical, and electrical systems. Use when the user wants to benchmark on BuiltBench, or asks about evaluating this task. Reports task-specific metrics.

ai-agentspythongo
0
3
Bull EvalA

Evaluates the ability of LLMs to generate correct SQL queries from natural language questions in financial domains. It tests schema linking, cross-database transfer, and output calibration capabilities specific to fund, stock, and macroeconomic data. Use when the user wants to benchmark on BULL, or asks about evaluating this task. Reports execution accuracy (EX).

researchpythongo
0
3
Burgers Robustness EvalA

Evaluates the adversarial robustness and baseline accuracy of neural operators trained to solve the viscous Burgers’ equation. It probes whether targeted active learning and architectural denoising can mitigate sensitivity to input perturbations compared to uniform or random sampling strategies. Use when the user wants to benchmark on Viscous Burgers’ Equation (Spectral Solver), or asks about evaluating this task. Reports Combined (%).

researchpythongo
0
3
Bwgnn Anomaly Detection EvalA

Evaluates graph neural networks and baselines for node anomaly detection in heterogeneous and homogeneous graphs. It probes the model's ability to identify fraudulent or anomalous accounts/users based on node features and graph structure under both supervised and semi-supervised settings. Use when the user wants to benchmark on YelpChi, Amazon, T-Finance, T-Social, or asks about evaluating this task. Reports F1-macro.

researchpythongo
0
3
Bwor EvalA

Evaluates LLMs' ability to automate operations research problem solving through mathematical modeling, code generation, and solver-based optimization. It probes whether reasoning agents can correctly translate natural language OR problems into executable models and compute optimal solutions. Use when the user wants to benchmark on BWOR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Bxaic EvalA

Evaluates the faithfulness and localization accuracy of explainable AI (XAI) methods for Graph Neural Networks on molecular graphs. It measures how well explainers identify ground-truth chemical motifs (nodes/edges) versus correctly identifying when the entire graph is important, using threshold-free metrics. Use when the user wants to benchmark on B-XAIC, or asks about evaluating this task. Reports NE.

researchpythonnode
0
3
Byol Lrl EvalA

Evaluates LLM capabilities in low- and extreme-low-resource languages (Chichewa, Māori) across reasoning, reading comprehension, factual knowledge, and machine translation. It measures both language-specific adaptation gains and preservation of multilingual/English capabilities. Use when the user wants to benchmark on BYOL Evaluation Benchmarks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
C Mteb EvalA

Evaluates the generalization and effectiveness of Chinese text embedding models across six core NLP tasks: retrieval, semantic textual similarity (STS), pair classification, single-label classification, re-ranking, and clustering. It measures how well dense vector representations capture semantic relationships for diverse downstream applications. Use when the user wants to benchmark on C-MTEB, or asks about evaluating this task. Reports Average Performance.

ai-agentspythongo
0
3
C Mteb Sentence Embed EvalA

Evaluates the zero-shot quality of sentence and word embeddings from decoder-only LLMs on downstream tasks without additional training. Probes capabilities in text classification, clustering, retrieval, semantic similarity, and polysemous word disambiguation. Use when the user wants to benchmark on C-MTEB, SLPWC (C-SEM), WSD, or asks about evaluating this task. Reports Average C-MTEB normalized score.

ai-agentspythongo
0
3
C Srrg EvalA

Evaluates multimodal large language models on automated structured radiology report generation, specifically testing their ability to produce clinically accurate findings and impressions while integrating rich clinical context (multi-view X-rays, indications, techniques, prior studies) to mitigate temporal hallucinations. Use when the user wants to benchmark on C-SRRG, or asks about evaluating this task. Reports F1-SRRG-BERT.

researchpythongo
0
3
C Sts EvalA

Evaluates how well sentence embedding models capture conditional semantic similarity by measuring how accurately they rank similarity scores under different conditions compared to human ratings. It probes condition-aware representation learning and the ability to modulate embeddings based on specific contextual constraints. Use when the user wants to benchmark on C-STS, or asks about evaluating this task. Reports Spearman Rank correlation.

researchpythongo
0
3
C2c EvalA

Evaluates the strategic coordination and negotiation capabilities of language model agents in a long-horizon, mixed-motive multi-agent environment with asymmetric objectives and private information. Use when the user wants to benchmark on C2C, or asks about evaluating this task. Reports win rate.

ai-agentspythongo
0
3
C2f Chart EvalA

Evaluates a model's ability to classify chart types from images using a coarse-to-fine curriculum learning approach. It measures performance on broad and fine-grained chart categories to assess hierarchical classification capabilities. Use when the user wants to benchmark on ICPR 2022 UB Unitec PMC Dataset, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
C2prompt EvalA

Evaluates federated continual learning methods on image classification benchmarks. It probes long-term knowledge accumulation, progressive performance across sequential tasks, and the model's ability to retain old knowledge while learning new ones under non-IID client distributions. Use when the user wants to benchmark on ImageNet-R, DomainNet, CIFAR-100, or asks about evaluating this task. Reports Avg.

researchpythonperformance
0
3
C2rope 3d Vqa EvalA

Evaluates a 3D large multimodal model's ability to perform spatial reasoning and visual question answering on multi-view 3D scene data. It probes the model's capacity to retain early visual context, understand spatial relationships, and generate accurate text responses to complex 3D scene queries. Use when the user wants to benchmark on ScanQA, SQA3D, or asks about evaluating this task. Reports EM@1.

researchpythongo
0
3
C3 Bench EvalA

Evaluates the robustness and multi-tasking capabilities of LLM-based agents by probing their ability to handle complex tool dependencies, propagate hidden information across tasks, and maintain stable decision policies under dynamic, multi-round interactions. Use when the user wants to benchmark on C^3-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
C4 Loss Scaling EvalA

Evaluates how optimal batch size and learning rate scale with model size and target loss during pre-training. It probes the stability of hyperparameters across different model scales and quantifies the relationship between batch size and validation loss on a standard text corpus. Use when the user wants to benchmark on C4, or asks about evaluating this task. Reports C4 Loss.

researchpythongit
0
3
Ca Afp EvalA

Evaluates federated learning models for human activity recognition under non-IID data distributions. It measures classification accuracy, fairness across heterogeneous clients, and communication efficiency during model pruning and clustering. Use when the user wants to benchmark on WISDM, UCI-HAR, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
Caa EvalA

Evaluates the robustness of Large Audio-Language Models (LALMs) against adversarial audio attacks in conversational settings. It probes response consistency, semantic preservation, and linguistic quality when audio inputs are perturbed with content, emotional, explicit, or implicit noise. Use when the user wants to benchmark on CAA, or asks about evaluating this task. Reports WER.

researchpythongit
0
3
Caad 3k EvalA

Evaluates a model's ability to detect contextual anomalies where normality depends on subject-context alignment rather than intrinsic appearance. It probes cross-context generalization by testing on unseen subject-context combinations and zero-shot transfer to real-world out-of-context benchmarks. Use when the user wants to benchmark on CAAD-3K, MVTec-AD, VisA, MIT-OOC, COCO-OOC, or asks about evaluating this task. Reports I-AUROC.

researchpythontesting
0
3
Cab EvalA

Evaluates whether LLMs exhibit biased responses to automatically generated, realistic open-ended questions. It probes for hidden biases across sensitive attributes (e.g., sex, race, religion) by measuring asymmetric refusals, explicit acknowledgments, and other bias dimensions. Use when the user wants to benchmark on CAB, or asks about evaluating this task. Reports fitness score.

researchpythongo
0
3
Cabuar Burned Area Delineation EvalA

This benchmark evaluates a model's ability to accurately delineate wildfire-affected regions from satellite imagery. It probes pixel-level binary segmentation and change detection capabilities using pre- and post-fire Sentinel-2 multispectral data. Use when the user wants to benchmark on CaBuAr, or asks about evaluating this task. Reports pixel-level accuracy.

researchpythongit
0
3
Cache Gen EvalA

Evaluates the efficiency and generation quality of a KV cache compression and streaming system for LLM serving. It measures how effectively the system reduces network bandwidth and time-to-first-token across varying context lengths, network conditions, and concurrent requests while maintaining task-specific accuracy, F1, or perplexity. Use when the user wants to benchmark on LongChat, TriviaQA, NarrativeQA, Wikitext, or asks about evaluating this task. Reports TTFT.

researchpythongo
0
3
Cads Whole Body Ct EvalA

Evaluates AI models on voxel-level segmentation of 167 whole-body anatomical structures from CT scans. It probes generalization across diverse imaging protocols, patient demographics, and pathological conditions, with a focus on clinical utility in radiation oncology. Use when the user wants to benchmark on CADS-dataset, 18 Public Benchmark Datasets, or asks about evaluating this task. Reports Dice coefficient.

researchpythongo
0
3
Caf 7m EvalA

Evaluates whether context-aided forecasting models can effectively leverage supplementary descriptive context to improve probabilistic time series forecasts, and tests generalization to out-of-domain real-world datasets. Use when the user wants to benchmark on CAF-7M, CGTSF, GIFT-Eval, or asks about evaluating this task. Reports CRPS.

researchpythongit
0
3
Cafp Fairness EvalA

Evaluates whether a post-processing framework can reduce group-level disparities in predictions while maintaining predictive accuracy. It probes a model's ability to balance fairness constraints (demographic parity and equalized odds) against standard classification performance across multiple benchmark datasets. Use when the user wants to benchmark on Adult Income (UCI), COMPAS Recidivism, German Credit, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
Cage Korset EvalA

Probes large language models' vulnerability to culturally-adapted adversarial prompts across Korean and Khmer contexts. It measures how effectively models resist harmful intent when framed within local socio-technical norms, safety policies, and cultural specifics, rather than generic or translated attacks. Use when the user wants to benchmark on KorSET, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythongit
0
3