
Claude Skills by qhjqhj00
github.com/qhjqhj00This evaluation probes the effectiveness of explainable AI (XAI)-driven feature selection methods on network intrusion detection systems. It measures how well various black-box machine learning models classify network traffic flows into normal or specific attack categories when trained on different subsets of extracted features. Use when the user wants to benchmark on CICIDS-2017, RoEduNet-SIMARGL2021, or asks about evaluating this task. Reports Accuracy (Acc).
Evaluates the reliability and operational suitability of white-box explainable AI methods (DeepLift, Integrated Gradients, LRP) when applied to deep neural network-based intrusion detection systems. It probes how well these methods preserve model accuracy, maintain consistency under repeated runs, resist adversarial noise, and compute efficiently across real-world network traffic datasets. Use when the user wants to benchmark on NSL-KDD, RoEduNet-SIMARGL2021, CICIDS-2017, or asks about evalua...
Evaluates a machine learning model's ability to predict local chemical descriptors (oxidation state and coordination number) from experimental X-ray absorption spectra, specifically testing how well spectral domain mapping bridges the gap between simulated training data and real experimental measurements. Use when the user wants to benchmark on Combinatorial Zinc Titanate Thin Film XANES, or asks about evaluating this task. Reports OS/CN prediction accuracy.
Evaluates a model's ability to perform cross-cultural machine translation, specifically probing its capacity to accurately transcreate culturally nuanced entity names across multiple language pairs rather than merely transliterating or omitting them. Use when the user wants to benchmark on XC-Translate, WMT (17-21), or asks about evaluating this task. Reports M-ETA.
Evaluates large language models on multilingual code understanding, generation, translation, and retrieval across 11 programming languages. It probes the model's ability to produce executable, correct code by validating outputs against unit tests rather than relying on lexical overlap. Use when the user wants to benchmark on xCodeEval, or asks about evaluating this task. Reports pass@5.
Evaluates the ability of multi-agent LLM frameworks to automatically fix buggy Ruby code using test-driven feedback loops. It probes iterative code repair, self-reflection, and test generation capabilities under varying difficulty levels and error types. Use when the user wants to benchmark on xCodeEval (Ruby subset), or asks about evaluating this task. Reports pass@1.
Evaluates a model's ability to detect violent events in untrimmed audio-visual videos under weak supervision. It probes the model's capacity to fuse complementary audio and visual cues to localize violence in class-imbalanced, long-range video sequences. Use when the user wants to benchmark on XD-Violence, or asks about evaluating this task. Reports AP.
Evaluates the ability of graph neural networks to detect fraudulent transactions in large-scale, highly imbalanced e-commerce transaction graphs. It probes model performance under extreme class imbalance and measures the trade-off between detection accuracy, inference speed, and scalability across different graph sizes and distributed settings. Use when the user wants to benchmark on eBay-xlarge, eBay-large, eBay-small, or asks about evaluating this task. Reports AUC.
Evaluates cross-lingual transfer capabilities of pre-trained language models across 11 diverse natural language understanding and generation tasks spanning over 100 languages. It measures how well models fine-tuned on English can generalize to zero-shot testing in other languages. Use when the user wants to benchmark on XGLUE, or asks about evaluating this task. Reports accuracy.
Evaluates a Text-to-SQL framework's ability to generate correct and efficient SQL queries from natural language questions across complex, cross-domain databases. It probes schema filtering, multi-generator candidate creation, and selection robustness. Use when the user wants to benchmark on BIRD, Spider, or asks about evaluating this task. Reports Execution Accuracy (EX).
Evaluates the ability of text-to-SQL models to generate correct SQL or GQL queries for natural language questions across relational and graph databases. It measures execution accuracy by comparing the runtime results of generated queries against reference queries on specific database instances. Use when the user wants to benchmark on Spider, Bird, SQL-Eval, NL2GQL, or asks about evaluating this task. Reports Execution Accuracy (EX).
Evaluates multimodal large language models on ultra-high-resolution remote sensing imagery using vision-language question answering. It probes both perception (e.g., object classification, counting, spatial relations) and reasoning capabilities across various sub-tasks. Use when the user wants to benchmark on XLRS-Bench, LRS-VQA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the cross-lingual transfer capability of neural news recommenders. It probes how well models trained monolingually on English news can generate accurate recommendations in 14 other languages under zero-shot and few-shot settings, with and without bilingual user consumption patterns. Use when the user wants to benchmark on xMIND, or asks about evaluating this task. Reports AUC.
This benchmark probes the cross-modal consistency and reasoning capabilities of omni-language models by evaluating semantic equivalence across all six possible modality combinations (text, vision, audio) for both context and candidate inputs. It measures how well models maintain performance when modalities are swapped or combined, highlighting modality-specific biases and directional asymmetries. Use when the user wants to benchmark on XModBench, or asks about evaluating this task. Reports ac...
Evaluates multilingual and cross-lingual capabilities of LLMs on Natural Language Inference (XNLI) and topic classification (SIB-200) across English and low-resource South Asian languages (Bangla, Hindi, Urdu) using zero-shot prompting. Use when the user wants to benchmark on XNLI, SIB-200, or asks about evaluating this task. Reports F1macro.
Evaluates automatic machine translation metrics by measuring their correlation with human quality judgments across multiple language pairs. It probes whether metrics exhibit cross-lingual scoring bias and how reliably they rank translation systems or quality triplets relative to human assessments. Use when the user wants to benchmark on XQ-MEval, or asks about evaluating this task. Reports Kendall-τ.
Probes large-scale 3D scene understanding by evaluating a model's ability to reason across multiple rooms, locate specific objects, generate embodied task plans, and produce detailed captions in complex, high-density environments. It specifically tests spatial awareness, contextual inference, and fine-grained detail retention beyond single-room benchmarks. Use when the user wants to benchmark on XR-Scene, or asks about evaluating this task. Reports CIDEr.
Evaluates an LLM's ability to perform cross-lingual retrieval-augmented generation by answering questions in a target language using supporting documents in English or mixed languages, while ignoring topically related distractors. It specifically probes cross-document reasoning capabilities and response language consistency. Use when the user wants to benchmark on XRAG, or asks about evaluating this task. Reports response language consistency.
Evaluates the capability of vision-language models to generate accurate and clinically relevant radiology reports from chest X-ray images. It probes both linguistic fluency and coverage against reference reports, as well as clinical accuracy in identifying pathological findings. Use when the user wants to benchmark on IU X-ray, MIMIC-CXR, CheXpert Plus, or asks about evaluating this task. Reports ROUGE-L.
Evaluates ML inference accelerators under realistic Extended Reality (XR) workloads. It probes the system's ability to handle real-time, multi-task, multi-model (MTMM) pipelines with dynamic dependencies while meeting strict latency, energy, and quality-of-experience (QoE) constraints. Use when the user wants to benchmark on XRBench Scenarios, or asks about evaluating this task. Reports overall score.
Evaluates the fidelity and stability of state-explaining methods in Reinforcement Learning across tabular and image-based environments. It measures how accurately explanations identify critical states and how consistently they perform under perturbations. Use when the user wants to benchmark on XRL-Bench Environments (DunkCityDynasty-v1, LunarLander-v2, CartPole-v0, FlappyBird-v0, Breakout-v0, Pong-v0), or asks about evaluating this task. Reports AIM.
Probes whether large language models exhibit exaggerated safety behaviors by refusing safe prompts due to lexical overfitting or system prompt effects. It measures the model's ability to distinguish between genuinely unsafe requests and safe prompts that merely resemble unsafe content. Use when the user wants to benchmark on XSTest, or asks about evaluating this task. Reports response_classification.
Evaluates a model's ability to perform extreme abstractive summarization by generating a single-sentence summary from a full news article, requiring synthesis, paraphrasing, and inference across document sections. Use when the user wants to benchmark on XSum, or asks about evaluating this task. Reports automatic metrics.
Evaluates cross-task semantic consistency in unified multimodal models by measuring how well generation and understanding tasks align on shared scene-graph facts. It probes whether architectural unification leads to representation-level coherence or merely independent task accuracy, specifically highlighting failures like consistent hallucination. Use when the user wants to benchmark on XTC-Bench, or asks about evaluating this task. Reports CCTA, AW-CCTA.
Evaluates an LLM's ability to perform context-aware, instruction-guided revisions of academic paper sections. It probes controllable editing capabilities, specifically measuring adherence to revision instructions, clarity, conciseness, and alignment with scientific writing standards through automated pairwise comparisons and human scoring. Use when the user wants to benchmark on XtraQA, or asks about evaluating this task. Reports Length-controlled (LC) win rate.
Evaluates zero-shot cross-lingual transfer of multilingual language models. Models are trained exclusively on English-labeled data and then tested on 40 typologically diverse languages across nine tasks spanning sentence classification, structured prediction, question answering, and sentence retrieval. Use when the user wants to benchmark on XTREME, or asks about evaluating this task. Reports accuracy.
Evaluates zero-shot cross-lingual transfer by training models on English data and testing them on 50 typologically diverse languages across classification, QA, and retrieval tasks. It probes fine-grained diagnostic capabilities and cross-lingual alignment using structured performance breakdowns. Use when the user wants to benchmark on XQuAD, XCOPA, Mewsli-X, LAReQA, CheckList, or asks about evaluating this task. Reports Exact Match.
Evaluates cross-lingual speech representations across 102 languages and four task families: automatic speech recognition, speech translation, speech classification, and speech-text retrieval. It probes the ability of self-supervised speech models to generalize across languages, domains, and varying data regimes from high-resource to low-resource settings. Use when the user wants to benchmark on Fleurs, MLS, VoxPopuli, CoVoST-2, Minds-14, or asks about evaluating this task. Reports WER.
This benchmark evaluates the quality and stability of explanation methods for time series classification models. It probes four key properties: robustness to input perturbations, faithfulness to model predictions, computational complexity of the explanations, and reliability against known ground-truth informative features. Use when the user wants to benchmark on XTSC-Bench Synthetic Datasets, or asks about evaluating this task. Reports Faithfulness Correlation.
Compute xu1998hz/sescore_english_coco via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_coco.
Compute xu1998hz/sescore_english_mt via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_mt.
Compute xu1998hz/sescore_english_webnlg via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_english_webnlg.
Compute xu1998hz/sescore_german_mt via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_german_mt.
Compute xu1998hz/sescore via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore.
This benchmark evaluates the ability of multilingual and cross-lingual models to generate accurate English summaries from source documents in German, French, and Czech. It probes supervised, zero-shot, and few-shot cross-lingual transfer capabilities, as well as model robustness on out-of-domain news text. Use when the user wants to benchmark on XWikis, D_en→en, Voxeurop, or asks about evaluating this task. Reports ROUGE-L recall.
This evaluation probes a model's ability to optimize exploration-exploitation trade-offs in personalized click-through rate (CTR) prediction. It measures how effectively an exploration strategy improves cumulative user engagement and advertiser retention compared to standard ranking baselines. Use when the user wants to benchmark on Yahoo! R6B, or asks about evaluating this task. Reports CTR.
Compute ybelkada/cocoevaluate via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ybelkada/cocoevaluate.
Predicts a 1-to-5 star rating for a restaurant review based on its text content. It probes a model's ability to capture sentiment, domain-specific linguistic patterns, and fine-grained textual features for multi-class classification. Use when the user wants to benchmark on Yelp Dataset, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to predict sentiment on long, complex documents by leveraging discourse structure. It probes whether incorporating hierarchical discourse trees improves sentiment classification and regression over standard sequential baselines, particularly for longer texts where sentiment is more subtle and diverse. Use when the user wants to benchmark on Yelp'13, or asks about evaluating this task. Reports accuracy.
This evaluation probes a graph neural network's ability to detect fraudulent or spam reviews within a multi-relational graph structure. It measures classification robustness against structural and semantic inconsistencies by training on varying fractions of labeled data and testing on the remainder. Use when the user wants to benchmark on YelpChi, or asks about evaluating this task. Reports F1-score.
Compute Yeshwant123/mcc via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Yeshwant123/mcc.
Evaluates a multimodal model's ability to personalize to a few-shot visual concept for both understanding (recognition and QA) and pixel-level image generation. It probes whether learnable soft prompts can capture subject-specific details without catastrophic forgetting, while measuring token efficiency compared to standard prompting. Use when the user wants to benchmark on Yo’LLaVA dataset, or asks about evaluating this task. Reports Recognition Accuracy.
Evaluates monolingual automatic speech recognition performance across multiple languages using a large-scale, YouTube-collected audio dataset. It probes the model's ability to accurately transcribe spoken language from noisy, real-world video subtitles after alignment filtering. Use when the user wants to benchmark on YODAS, or asks about evaluating this task. Reports CER.
Assesses large language models' knowledge of Japanese yokai (folklore creatures) through multiple-choice questions. It probes cultural and linguistic familiarity with Japanese folklore, revealing how training data exposure and language affect cross-cultural knowledge retention. Use when the user wants to benchmark on YokaiEval, or asks about evaluating this task. Reports correctness.
Evaluates open-vocabulary object detection and segmentation capabilities using text, visual, and prompt-free inputs on zero-shot and fine-tuned settings. Use when the user wants to benchmark on LVIS, COCO, or asks about evaluating this task. Reports Fixed AP.
Compute yonting/average_precision_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yonting/average_precision_score.
This benchmark evaluates graph neural networks and generative models on AI-generated neural network architectures represented as directed acyclic graphs. It probes two capabilities: local component-level refinement (predicting data flows and operator types) and global end-to-end architecture generation. Use when the user wants to benchmark on Younger, or asks about evaluating this task. Reports F1.
Evaluates recommender systems on implicit feedback datasets, testing their ability to rank relevant items for users. It probes model versatility across cold-start, offline, and instant recommendation scenarios using side information and sequential context features. Use when the user wants to benchmark on YouTube Implicit Feedback Subset, or asks about evaluating this task. Reports NDCG@100.
Compute yqsong/execution_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yqsong/execution_accuracy.
Evaluates a model's ability to detect topic boundaries in unstructured spoken transcriptions (text segmentation) and generate coherent chapter titles (smart chaptering). It probes hierarchical structuring, real-time/online processing constraints, and cross-domain generalization to meeting transcripts. Use when the user wants to benchmark on WIKI-727K, YTSEG, QMSUM, YTSEG[TITLES], or asks about evaluating this task. Reports F1.