
Claude Skills by qhjqhj00
github.com/qhjqhj00Compute the d2_absolute_error_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_absolute_error_score, or asks how to score with d2_absolute_error_score.
Compute the d2_brier_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_brier_score, or asks how to score with d2_brier_score.
Compute the d2_log_loss_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_log_loss_score, or asks how to score with d2_log_loss_score.
Compute the d2_pinball_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_pinball_score, or asks how to score with d2_pinball_score.
Compute the d2_tweedie_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute d2_tweedie_score, or asks how to score with d2_tweedie_score.
Evaluates enterprise network vulnerability to lateral attacks by simulating adversarial movement across authentication graphs. It also measures the effectiveness of defense strategies in predicting attacker movement based on graph topology and credential hygiene levels. Use when the user wants to benchmark on G_s, G_l, G_lanl, or asks about evaluating this task. Reports Network Vulnerability.
Evaluates a diffusion-based reinforcement learning agent's capability to dynamically assign AI-generated content tasks to edge service providers under stochastic workloads, optimizing for user utility while preventing system crashes and minimizing training time. Use when the user wants to benchmark on Custom AIGC Edge Simulation Environment, Gym Benchmark Tasks, or asks about evaluating this task. Reports Cumulative Reward.
Evaluates a model's ability to predict driver maneuvers and generate hierarchical, human-understandable explanations from ego-centric driving videos. It probes spatio-temporal feature disentanglement and the impact of gaze modality on interpretability. Use when the user wants to benchmark on DAAD-X, or asks about evaluating this task. Reports Acc.
This benchmark evaluates AI-based data assimilation models for weather forecasting. It tests their ability to integrate sparse or noisy observations with background atmospheric fields to produce accurate analysis fields, which are then used to initialize skillful medium-range weather predictions. Use when the user wants to benchmark on ERA5, GDAS, or asks about evaluating this task. Reports RMSE.
Evaluates the ability of computer vision models to perform multi-label semantic segmentation for identifying and localizing various reinforced concrete defects on bridge infrastructure. It probes pixel-level classification accuracy across diverse, real-world damage types and structural components. Use when the user wants to benchmark on dacl10k, or asks about evaluating this task. Reports mean IoU.
Compute daiyizheng/valid via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of daiyizheng/valid.
Compute DaliaCaRo/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DaliaCaRo/accents_unplugged_eval.
Compute danasone/ru_errant via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danasone/ru_errant.
Compute danieldux/hierarchical_softmax_loss via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/hierarchical_softmax_loss.
Compute danieldux/isco_hierachical_accuracy_v2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/isco_hierachical_accuracy_v2.
Compute danieldux/isco_hierarchical_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of danieldux/isco_hierarchical_accuracy.
Compute dannashao/span_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of dannashao/span_metric.
Evaluates the quality of Chinese vision-language pre-training datasets by measuring downstream performance on cross-modal retrieval and multimodal reasoning benchmarks after continual pre-training. Use when the user wants to benchmark on Flickr30K-CN, MSCOCO-CN, MUGE, DCI-CN, DOCCI-CN, or asks about evaluating this task. Reports R@1/5/10.
Evaluates cross-domain patent retrieval systems by measuring how well they rank relevant patent documents or passages when queries and targets share or lack overlapping IPC3 classifications. Use when the user wants to benchmark on DAPFAM, or asks about evaluating this task. Reports NDCG@100.
Evaluates speech enhancement models by measuring their ability to restore clean speech from noisy or degraded inputs. It probes perceptual quality, acoustic fidelity, and semantic preservation using human listening tests and objective feature-space distances. Use when the user wants to benchmark on DAPS, Noisy VCTK, or asks about evaluating this task. Reports MUSHRA.
Evaluates the accuracy and computational efficiency of the DARF simulation framework for predicting speech recognition thresholds (SRT) in normal-hearing and hearing-impaired listeners across various acoustic maskers and hearing aid conditions. Use when the user wants to benchmark on Empirical SRT datasets (Hochmuth et al. 2015, Hülsmeier et al., Schädler et al. 2020a), or asks about evaluating this task. Reports SRT.
Evaluates the ability of unsupervised machine learning models to detect deviations from Standard Model physics in high-energy collider data without assuming specific new physics signatures. It probes model-agnostic anomaly detection by measuring how well density estimation and reconstruction-based methods separate background events from potential signal events. Use when the user wants to benchmark on Dark Machines Anomaly Score Challenge Dataset, or asks about evaluating this task. Reports re...
Compute DarrenChensformer/action_generation via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DarrenChensformer/action_generation.
Compute DarrenChensformer/eval_keyphrase via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DarrenChensformer/eval_keyphrase.
Compute DarrenChensformer/relation_extraction via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DarrenChensformer/relation_extraction.
This evaluation probes the robustness, privacy compliance, bias/fairness, and hallucination resistance of medical large language models under dynamic, adversarial stress. It measures how well models maintain safety and accuracy when prompts are iteratively mutated by autonomous agents to exploit vulnerabilities, mimicking real-world clinical interactions rather than static benchmark conditions. Use when the user wants to benchmark on MedQA, Privacy-trap scenarios, Medical bias dataset, or ask...
This evaluation probes the robustness of machine learning models against data poisoning attacks by measuring classification accuracy degradation and recovery under label flipping and image replacement attacks. It assesses how well statistical anomaly detection, adversarial training, and ensemble learning defenses mitigate performance drops and false prediction rates. Use when the user wants to benchmark on CIFAR-10, Insurance Claims, or asks about evaluating this task. Reports classification ...
Evaluates a model's ability to retrieve relevant tables and text passages from a hybrid corpus to satisfy complex, multi-part analytical user requests (Data Product Requests). It probes multi-modal data integration and semantic clustering capabilities by requiring complete alignment between a request and its underlying data assets. Use when the user wants to benchmark on HybridQA, TAT-QA, ConvFinQA, or asks about evaluating this task. Reports Full Recall@100.
Evaluates whether distributional or embedding similarity between a model's pretraining data and downstream tasks predicts few-shot or finetuned performance. It probes the 'similarity hypothesis' by measuring correlations between aggregate and example-level text similarities and model accuracy. Use when the user wants to benchmark on BIG-bench Lite, GLUE, or asks about evaluating this task. Reports correlation coefficient.
Evaluates semi-supervised anomaly detection models for supply chain fraud under severe class imbalance and limited label availability. Probes the ability to leverage unsupervised pre-filtering and self-training to improve precision, recall, and F1-score while maintaining low false positive rates. Use when the user wants to benchmark on DataCo Smart Supply Chain Dataset, or asks about evaluating this task. Reports F1-Score.
Evaluates the zero-shot generalization capability of vision-language models across a diverse suite of image classification benchmarks. It measures how well pre-trained image-text alignment transfers to unseen downstream tasks without fine-tuning. Use when the user wants to benchmark on ImageNet, DataComp evaluation datasets, or asks about evaluating this task. Reports ImageNet.
Evaluates zero-shot generalization of vision-language models across a broad suite of classification and retrieval tasks, with a focus on image-text alignment and retrieval accuracy. Use when the user wants to benchmark on DataComp Zero-Shot Suite, or asks about evaluating this task. Reports ImageNet accuracy.
Evaluates the quality, accountability, and documentation standards of dataset and benchmark papers across major AI conferences. It probes whether papers provide transparent data collection guidelines, quality assurance practices, and clear provenance using a structured rubric. Use when the user wants to benchmark on Conference Dataset & Benchmark Papers (2021-2024), or asks about evaluating this task. Reports datarubrics_compliance_rate.
Evaluates large language models on their ability to parse, translate, and perform arithmetic reasoning with datetime information. It probes structured string formatting (ISO-8601) and multi-step temporal calculations across diverse linguistic contexts. Use when the user wants to benchmark on DATETIME, or asks about evaluating this task. Reports accuracy.
Compute davebulaval/meaningbert via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of davebulaval/meaningbert.
Compute the davies_bouldin_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute davies_bouldin_score, or asks how to score with davies_bouldin_score.
Compute the DaviesBouldinScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DaviesBouldinScore, or asks how to score with DaviesBouldinScore.
Evaluates protein-ligand binding affinity prediction models on a modification-aware dataset, testing their ability to generalize across different train-test splits (new ligands, new proteins, modifications) and assessing robustness to wild-type overfitting and few-shot fine-tuning. Use when the user wants to benchmark on DAVIS-complete, or asks about evaluating this task. Reports Rp.
Evaluates whether large language models can discover genuinely new biological knowledge by generating correct scientific hypotheses or mechanisms from post-release literature, enforcing strict temporal separation to prevent data leakage. Use when the user wants to benchmark on DBench-Bio, or asks about evaluating this task. Reports Score.
Evaluates the reconstruction fidelity of high-spatial-compression autoencoders and the generation quality and efficiency of latent diffusion models that utilize them. It benchmarks performance across multiple datasets and resolutions to assess trade-offs between compression ratio, image quality, and computational throughput. Use when the user wants to benchmark on ImageNet, FFHQ, MapillaryVistas, MJHQ, or asks about evaluating this task. Reports rFID, FID.
Evaluates multimodal large language models' ability to generate executable HTML/JavaScript code for dynamic chart animations from text or video prompts. It probes instruction following, code executability, and fine-grained semantic alignment between generated visualizations and input specifications. Use when the user wants to benchmark on DCG-8K, or asks about evaluating this task. Reports Execution Pass Rate.
Compute the dcg_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute dcg_score, or asks how to score with dcg_score.
Evaluates the effectiveness of data curation strategies for language models by training base models on curated corpora and measuring performance on 53 downstream tasks. It isolates data quality effects from architectural and computational variables using fixed training recipes across multiple compute scales. Use when the user wants to benchmark on DCLM downstream tasks, or asks about evaluating this task. Reports MMLU 5-shot accuracy.
Evaluates the ability of multi-task learning models to predict post-click conversion rate (CVR) and click-through conversion rate (CTCVR) while mitigating selection bias in recommendation and search systems. It probes whether causal debiasing mechanisms improve ranking quality on both clicked and unclicked items across diverse e-commerce and industrial datasets. Use when the user wants to benchmark on Ali-CCP, Ali-Express (AE-ES), Ali-Express (AE-FR), Ali-Express (AE-NL), Ali-Express (AE-US),...
Evaluates the performance of machine learning and traditional algorithms for reconstructing particle tracks in drift chamber detectors. It probes hit-level matching accuracy, track-level reconstruction efficiency, charge identification correctness, and momentum resolution under realistic detector conditions. Use when the user wants to benchmark on DCTracks, or asks about evaluating this task. Reports track efficiency.
This benchmark evaluates the ability of deep learning models to detect AI-generated art (deeparts) versus conventional art (conarts) and identify their generative origins. It probes detector generalization across different state-of-the-art diffusion models and tests continual learning capabilities under evolving data streams with strict memory constraints. Use when the user wants to benchmark on DDDB, or asks about evaluating this task. Reports AA.
Probes large language models' ability to perform iterative medical diagnostic reasoning through a simulated patient-doctor dialogue. It evaluates whether models can gather clinical evidence, generate differential diagnoses, and converge on the correct ground-truth pathology within a limited number of interaction turns. Use when the user wants to benchmark on DDXPlus, or asks about evaluating this task. Reports diagnostic accuracy.
Evaluates an agent's ability to iteratively collect clinical evidence and generate a ranked list of differential diagnoses for a simulated patient. It probes the system's diagnostic reasoning, evidence-gathering efficiency, and alignment with ground-truth pathologies. Use when the user wants to benchmark on DDXPlus, or asks about evaluating this task. Reports DDF1.
This benchmark probes the acoustic faithfulness of Audio Multimodal Large Language Models (Audio MLLMs) by measuring how reliably they attend to acoustic cues (emotional prosody, background sounds, speaker identity) when faced with conflicting textual semantics or misleading prompts. It specifically diagnoses the tendency of models to prioritize text over audio (text dominance) under progressive levels of interference. Use when the user wants to benchmark on DEAF, or asks about evaluating thi...
Evaluates the capability of neural architectures to perform binary emotion recognition (valence, arousal, dominance) directly from raw, multi-channel EEG time-series data without hand-crafted features. It measures how well a model generalizes across subjects using a standard cross-validation protocol. Use when the user wants to benchmark on DEAP, or asks about evaluating this task. Reports Accuracy.