
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates a Retrieval-Augmented Generation system's ability to ground responses in retrieved legal documents and directly address jurisdiction-specific AI regulation queries. It tests the system's performance on both single-entity and multi-jurisdictional comparison tasks using automatic LLM-based scoring. Use when the user wants to benchmark on Global AI Regulation Test Queries, or asks about evaluating this task. Reports faithfulness.
This evaluation probes a model's ability to recover and predict sparse latent fitness functions from observed sequence-fitness data distorted by global epistasis. It compares data efficiency and predictive accuracy under complete versus incomplete (subsampled) data regimes, highlighting the robustness of contrastive losses over mean-squared error. Use when the user wants to benchmark on NK model synthetic data, FLIP benchmark, or asks about evaluating this task. Reports Spearman correlation.
This benchmark probes physical commonsense reasoning by testing whether models can distinguish correct from incorrect solutions to everyday physical tasks. It specifically evaluates cultural and linguistic grounding by using items constructed natively in 116 language varieties, avoiding translation artifacts that often skew multilingual evaluations. Use when the user wants to benchmark on Global PIQA, or asks about evaluating this task. Reports accuracy.
Evaluates AI music generation models for global and cultural bias by measuring how well generated tracks match reference tracks across different world regions and genres. It probes the models' out-of-distribution capabilities and tendency to default to mainstream styles over authentic regional ones. Use when the user wants to benchmark on GlobalDISCO, or asks about evaluating this task. Reports FAD.
Evaluates the audio fidelity, speaker diversity, and transcript alignment of English multi-speaker TTS datasets. It also benchmarks how well zero-shot speaker-adaptive TTS models trained on these corpora generalize to unseen speakers and global accents. Use when the user wants to benchmark on GLOBE, VCTK, Common Voice, LibriTTS, LibriTTS-R, or asks about evaluating this task. Reports NMOS.
This benchmark evaluates session-based news recommendation systems by predicting the next article a user will click based on their recent interaction history. It probes a model's ability to capture short-term user intent, handle temporal dynamics, and leverage both content and contextual features in a streaming environment. Use when the user wants to benchmark on Globo.com, or asks about evaluating this task. Reports HR@5.
Evaluates open-world knowledge graph question answering by testing a model's ability to answer single-hop and multi-hop questions over incomplete graphs. It probes the integration of structural graph signals with textual semantics to handle missing answer paths and domain-specific reasoning without relying on fine-tuning or retrieval-only pipelines. Use when the user wants to benchmark on GLOW-Bench, Arxiv2023, ogbn-arxiv, ogbn-products, or asks about evaluating this task. Reports Exact Match...
Evaluates the ability of efficient fine-tuning and structured sparsity methods to maintain performance across a diverse suite of natural language understanding tasks. It probes task-specific classification accuracy and correlation metrics under both full-data and limited-data regimes, as well as pre-training perplexity on a large-scale corpus. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports GLUE average score.
Evaluates the transfer learning capability of pre-trained language models across a diverse suite of natural language understanding tasks, including sentence classification, semantic textual similarity, and natural language inference. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports GLUE Average.
Evaluates few-shot text classification performance across multiple natural language understanding tasks. It probes a model's ability to generalize from extremely limited labeled examples (16 per class) by generating synthetic training data and fine-tuning a classifier. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports Average performance.
Evaluates the quantization performance of pre-trained language models (both discriminative BERT-style and generative GPT-style) across standard NLP classification, regression, and language modeling tasks under data-free zero-shot quantization settings. Use when the user wants to benchmark on GLUE, WikiText2, Penn Treebank (PTB), WikiText103, or asks about evaluating this task. Reports GLUE Avg..
Evaluates the generalization and stability of finetuned pretrained language models on low-resource NLP tasks. It probes how well models adapt to sentiment classification, natural language inference, paraphrasing, similarity assessment, and linguistic acceptability when trained on severely limited data (300–1000 examples). Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports averaged evaluation metrics.
This protocol evaluates whether language models memorize benchmark surface features or demonstrate true semantic robustness. It measures performance degradation when inputs are paraphrased, lexically/syntactically perturbed, or adversarially rewritten, contrasting deterministic greedy decoding with stochastic distributional evaluation. Use when the user wants to benchmark on GLUE (MNLI, QQP, QNLI, SST-2), or asks about evaluating this task. Reports GLUE robustness ratio.
Evaluates the downstream generalization and data efficiency of pre-trained language models by fine-tuning them on standard natural language understanding and instruction-following benchmarks. Use when the user wants to benchmark on GLUE, SuperNatural-Instructions (SNI), or asks about evaluating this task. Reports GLUE.
Evaluates the ability of pre-trained language models to perform a diverse set of natural language understanding tasks, including sentence classification, paraphrase detection, semantic similarity, natural language inference, and extractive question answering. Use when the user wants to benchmark on GLUE, SQuAD 1.1, SQuAD 2.0, or asks about evaluating this task. Reports Task-specific metrics (Accuracy, F1, Spearman, Matthews).
Evaluates the generalization and downstream performance of pretrained language models on a suite of natural language understanding tasks (GLUE) and reading comprehension (SQuAD 2.0). Use when the user wants to benchmark on GLUE, SQuAD 2.0, or asks about evaluating this task. Reports GLUE.
Evaluates general language understanding across classification, regression, and entailment tasks, alongside language modeling capability. It also measures computational efficiency and internal attention allocation strategies under resource constraints. Use when the user wants to benchmark on GLUE Benchmark, WikiText-103, or asks about evaluating this task. Reports MNLI-m Accuracy.
Evaluates machine learning models on glycan analysis tasks, including taxonomic classification, immunogenicity prediction, glycosylation type prediction, and protein-glycan binding affinity estimation. It probes the ability of sequence-based and graph-based encoders to capture multi-relational glycan structures and benefit from multi-task learning. Use when the user wants to benchmark on GlycanML, or asks about evaluating this task. Reports Macro-F1.
Tests the ability of language models to automatically classify documents into predefined group membership categories (e.g., gender, geographic location) for group fairness evaluation in information retrieval. Use when the user wants to benchmark on TREC fair ranking track 2021, TREC fair ranking track 2022, NTCIR fairweb1 (Chuweb-21D), or asks about evaluating this task. Reports accuracy.
Evaluates a vision-language model's ability to understand and reason over diverse medical imaging modalities (CT, MRI, X-ray, pathology slides) to answer clinical questions, diagnose diseases, and perform anatomical or lesion recognition tasks. Use when the user wants to benchmark on PMCVQA, PathVQA, VQA-RAD, SLAKE, OmniMedVQA, GMAI-MMBench, MMMU Health & Medicine track, or asks about evaluating this task. Reports accuracy.
Compute GMFTBY/dailydialog_evaluate via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of GMFTBY/dailydialog_evaluate.
Compute GMFTBY/dailydialogevaluate via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of GMFTBY/dailydialogevaluate.
Evaluates a model's ability to recognize named entities in multimodal social media content and ground them to visual regions, while dynamically deciding when to use internal knowledge versus external search tools. Use when the user wants to benchmark on Twitter-GMNER*, Twitter-FMNERG*, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to perform pixel-level semantic segmentation for drivable areas and road anomalies using multi-modal visual inputs. It specifically probes how effectively networks can fuse RGB imagery with depth-related features (e.g., transformed disparity) to improve detection accuracy for ground mobile robots. Use when the user wants to benchmark on GMRP, KITTI road, KITTI semantic segmentation, or asks about evaluating this task. Reports IoU.
Evaluates the accuracy of density functionals in predicting broad molecular properties, including atomization energies, barrier heights, and noncovalent interactions, relative to high-level quantum chemistry reference data. Use when the user wants to benchmark on GMTKN55 (specifically GMTKN53 subset), or asks about evaluating this task. Reports WTMAD1.
Evaluates the adversarial robustness of Graph Neural Networks against node/edge injection and modification attacks, comparing Hamiltonian-based models against standard GNNs and defense baselines. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, Coauthor, Computers, Ogbn-Arxiv, Polblogs, or asks about evaluating this task. Reports accuracy.
Evaluates the expressiveness and practical performance of GNN-AK, a framework that replaces star-shaped neighbor aggregation with subgraph-based encoding in Message Passing Neural Networks. It probes the model's ability to distinguish complex graph structures (e.g., strongly regular graphs, substructures) and predict graph-level properties on standard benchmarks. Use when the user wants to benchmark on ZINC-12K, CIFAR10, PATTERN, MolHIV, MolPCBA, EXP, SR25, or asks about evaluating this task....
This evaluation probes how graph neural network (GNN) performance depends on the ratio of observed training nodes to feature dimensionality (N_obs/D). It tests whether standard benchmarks are biased toward low-dimensional regimes and evaluates the effectiveness of decoupling feature extraction from graph propagation. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, or asks about evaluating this task. Reports Accuracy.
Evaluates an LLM-guided framework's ability to automatically propose and refine Graph Neural Network architectures for node classification across diverse graph datasets, including out-of-distribution and heterophilic graphs, without requiring extensive training or search. Use when the user wants to benchmark on NAS-Bench-Graph & OOD Graphs (Cora, Citeseer, PubMed, CS, Physics, Photo, Computer, ogbn-arXiv, DBLP, Flickr, Actor), or asks about evaluating this task. Reports accuracy.
Evaluates the quality and interpretability of explanations generated by various Graph Neural Network (GNN) explainers across different architectures and graph datasets. It probes how well explanations align with human expectations (plausibility) and model decision logic (fidelity). Use when the user wants to benchmark on Grid, Grid-House, Stars, House-Color, or asks about evaluating this task. Reports F1-Fidelity.
Evaluates Graph Neural Networks and traditional ML models on a compiled Internet routing dataset for link prediction and node classification. It probes the ability to infer missing AS-AS connections and predict voluntary PeeringDB attributes under conditions of graph sparsity and label imbalance. Use when the user wants to benchmark on Internet Routing AS Graph Dataset, or asks about evaluating this task. Reports binary cross entropy.
Evaluates the runtime and memory efficiency of low-level Graph Neural Network and sparse tensor operations on NVIDIA A100 GPUs. It probes how input sparsity, tensor dimensions, and reduce factors affect computational overhead when operations are pushed to near-full GPU memory capacity. Use when the user has predictions and gold and needs to compute median_runtime.
Compares graph neural networks against deep fully-connected feedforward networks for binary classification of b-quark pairs in top-quark-antiquark collisions at the LHC. It probes whether explicit relational inductive biases in GNNs outperform permutation-invariant DNNs when provided with equivalent kinematic and relational features. Use when the user wants to benchmark on LHC t\bar{t} b\bar{b} event classification, or asks about evaluating this task. Reports mean ROC-AUC ($\mu_{\mathrm{AUC}}$).
Evaluates Graph Neural Networks for classifying emotional states from 32-channel EEG signals. It probes spatial-temporal feature extraction, cross-subject generalization, and robustness to hyperparameter settings in neuroscience signal processing. Use when the user wants to benchmark on FACED, or asks about evaluating this task. Reports Accuracy.
Evaluates the quality, robustness, and feasibility of perturbation-based GNN explainers. It probes whether generated subgraphs (factual or counterfactual) reliably preserve or flip model predictions, remain stable under topological or architectural perturbations, and satisfy domain-specific structural constraints. Use when the user wants to benchmark on Mutagenicity, Proteins, IMDB-B, AIDS, MUTAG, NCI1, Graph-SST2, DD, REDDIT-B, ogbg-molhiv, Tree-Cycles, Tree-Grid, BA-Shapes, or asks about ev...
Evaluates a self-supervised, classification-based anomaly detection framework (GOAD) on its ability to distinguish normal data from anomalies without labeled anomalies during training. It probes the model's robustness to data contamination and adversarial attacks across image and tabular domains. Use when the user wants to benchmark on CIFAR-10, FashionMNIST, Arrhythmia, Thyroid, KDD, KDDRev, or asks about evaluating this task. Reports ROC-AUC.
Evaluates large multimodal models' ability to perform goal-oriented embodied navigation in complex urban 3D airspace. It probes geometric perception, cross-view understanding, spatial imagination, and long-term memory by requiring models to navigate from a start point to a semantic goal using visual observations and historical context. Use when the user wants to benchmark on Embodied Navigation Benchmark, or asks about evaluating this task. Reports SR.
Evaluates cross-domain generalization of emotion classification models by measuring how well a model fine-tuned on GoEmotions transfers to external benchmarks with limited labeled data, compared to training from scratch. Use when the user wants to benchmark on ISEAR, EmoInt, Emotion-Stimulus, or asks about evaluating this task. Reports average F1-score.
This benchmark evaluates the capability of large language models to perform a wide range of financial natural language processing tasks in both English and Chinese. It probes domain-specific understanding, information extraction, reasoning, and generation across sentiment analysis, classification, entity/relation extraction, summarization, question answering, and stock movement prediction. Use when the user wants to benchmark on FPB, Fiqa-SA, Headlines, FOMC, lendingclub, NER, FinRE, CFA, EDT...
Evaluates reinforcement learning agents' ability to learn multi-agent coordination, strategic decision-making, and long-horizon planning in a physics-based 3D football simulation. It probes sample efficiency, reward shaping robustness, and performance under varying opponent difficulty levels. Use when the user wants to benchmark on Google Research Football (Football Benchmarks), or asks about evaluating this task. Reports average goal difference.
Evaluates an LLM's ability to generate correct API invocation code from natural language prompts, with or without retrieved documentation. It measures how well the model selects the appropriate API, avoids hallucinating non-existent APIs, and respects functional constraints like accuracy thresholds. Use when the user wants to benchmark on APIBench, or asks about evaluating this task. Reports AST accuracy.
Compute gorkaartola/metric_for_tp_fp_samples via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of gorkaartola/metric_for_tp_fp_samples.
Evaluates whether dependency-aware structural retrieval of agent skills improves task completion performance and computational efficiency compared to flat library access and semantic-only retrieval. It probes an agent's ability to assemble functionally complete, prerequisite-aware skill bundles for long-horizon technical and embodied sequential decision-making tasks. Use when the user wants to benchmark on SkillsBench, ALFWorld, or asks about evaluating this task. Reports average reward.
Evaluates tracking performance on a large-scale dataset of 10,000 videos with diverse object categories, testing generalization to real-world scenarios with language guidance. Use when the user wants to benchmark on GOT-10k, or asks about evaluating this task. Reports AO.
Evaluates the effectiveness of LLM-based prompt optimizers across complex reasoning, knowledge-intensive, and common NLP tasks. It measures how much optimized prompts improve model performance compared to baseline prompts and other optimization methods. Use when the user wants to benchmark on Big-Bench Hard (BBH), GSM8K, MMLU, WSC, WebNLG, or asks about evaluating this task. Reports average accuracy.
Evaluates large vision-language models on complex mathematical and visual reasoning tasks, measuring both correctness and computational efficiency. It specifically probes the model's ability to avoid excessive chain-of-thought generation (overthinking) by dynamically routing computation between fast perception, slow perception, and slow reasoning paths. Use when the user wants to benchmark on MathVision, MathVerse, MathVista, DynaMath, MM-Vet, or asks about evaluating this task. Reports accur...
Evaluates the few-shot, one-shot, and zero-shot learning capabilities of large autoregressive language models across diverse NLP tasks including language modeling, cloze completion, question answering, translation, and commonsense reasoning. Use when the user wants to benchmark on Penn Tree Bank (PTB), LAMBADA, HellaSwag, StoryCloze 2016, Natural Questions, WebQuestions, TriviaQA, WMT14/WMT16 Translation, or asks about evaluating this task. Reports accuracy.
Evaluates whether a language model's final numerical answer to a grade-school math word problem matches the ground truth. It uses an external LLM to extract the final answer from the model's generated solution and compares it against the gold answer. Use when the user has predictions and gold and needs to compute GPT4-based-Exact-Match.
Evaluates the optical character recognition and document understanding capabilities of GPT-4V across multiple tasks including scene text recognition, handwritten text recognition, mathematical expression recognition, and table structure recognition. It probes the model's robustness to different languages, handwriting styles, complex layouts, and input resolutions. Use when the user wants to benchmark on CUTE80, SCUT-CTW1500, Total-Text, WordArt, ReCTS, MLT19, IAM, CASIA-HWDB, CROHME2014, HME1...
Evaluates large language models on Arabic natural language understanding and generation across 44 tasks and over 60 datasets, covering both Modern Standard Arabic and dialectal varieties. Use when the user wants to benchmark on GPTAraEval Benchmark Suite, or asks about evaluating this task. Reports macro-F1.