
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the trade-off between classification performance (AUC) and model resilience (robustness to Monte Carlo simulation variations) in quark/gluon and top-quark jet tagging. It probes whether complex neural architectures generalize better to different physics simulators compared to simpler, physics-informed models. Use when the user wants to benchmark on Pythia 8 / Herwig 7 Jet Samples, or asks about evaluating this task. Reports AUC.
Evaluates Japanese financial text embedding models across classification, retrieval, and clustering tasks. It probes the models' ability to capture domain-specific semantics, handle regulatory and economic terminology, and generalize to zero-shot financial scenarios. Use when the user wants to benchmark on JFinTEB, or asks about evaluating this task. Reports macro-F1.
Compute jialinsong/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jialinsong/apps_metric.
Evaluates object detection, recognition, and precise tiling assembly capabilities for modular robotic manipulation. It tests the pipeline's ability to segment oddly shaped pieces, recognize their identity, and assemble them with high spatial accuracy. Use when the user wants to benchmark on Jigsaw puzzle set, or asks about evaluating this task. Reports score.
Evaluates spatial-temporal reasoning and hardware-agnostic robotic manipulation skills through a structured jigsaw puzzle assembly protocol. It measures vision-based segmentation, object recognition, pick planning success, and motion planning efficiency across three progressively complex physical tasks. Use when the user wants to benchmark on Jigsaw Manipulation Benchmark, or asks about evaluating this task. Reports Task Score.
Compute jijihuny/ecqa via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jijihuny/ecqa.
Evaluates a multimodal embedding model's ability to retrieve relevant documents, images, and code from large corpora, and to measure semantic similarity between text pairs across multiple languages and modalities. Use when the user wants to benchmark on J-VDR, ViDoRe, CLIPB, MMTEB, MTEB-en, COIR, LEMB, STS-m, STS-en, or asks about evaluating this task. Reports nDCG@10.
Compute jjkim0807/code_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jjkim0807/code_eval.
Evaluates Japanese biomedical large language models across five tasks: multiple-choice question answering, named entity recognition, machine translation, document classification, and semantic text similarity. It probes domain-specific knowledge, multilingual comprehension, and in-context learning capabilities. Use when the user wants to benchmark on JMedBench, or asks about evaluating this task. Reports F1-entity, Accuracy.
This benchmark evaluates large multimodal models' ability to understand Japanese-language visual content and answer questions across multiple disciplines. It specifically probes the gap between general language translation capabilities (culture-agnostic subset) and deep cultural knowledge (culture-specific subset), revealing how models handle language variation bias and culturally grounded reasoning. Use when the user wants to benchmark on JMMMU, or asks about evaluating this task. Reports ac...
Evaluates multimodal language models' ability to perform integrated visual-textual reasoning on Japanese-language tasks where questions and reference images are combined into a single composite image. It specifically probes OCR capabilities, visual perception, and cross-modal alignment in a multilingual context. Use when the user wants to benchmark on JMMMU-Pro, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to normalize job titles by mapping them to standardized ESCO occupation labels using semantic similarity. It probes the model's capacity to handle hierarchical occupational taxonomies and filter out irrelevant contextual tokens like locations. Use when the user wants to benchmark on JobBERT Vacancy Titles, or asks about evaluating this task. Reports MRR.
Evaluates behavioral cloning and extraction fidelity in bias-mitigation LLM agents processing job descriptions. It measures how well mutated agent outputs align with expert golden references using cross-item similarity and a multi-facet diagnostic rubric. Use when the user wants to benchmark on JobFair Corpus, or asks about evaluating this task. Reports BERTScore (F1) against 5-NN average.
Evaluates the ability of a genetic algorithm to evolve a navigation program for an artificial ant to traverse a complex, toroidal grid trail with gaps and high-difficulty sections. The benchmark measures how well the evolved program generalizes beyond the standard Santa Fe trail into a chaotic extended sector. Use when the user wants to benchmark on John Muir Ant Problem, or asks about evaluating this task. Reports score.
Evaluates models' ability to detect physical face spoofing attacks and digital face forgeries using visual appearance and physiological rPPG cues. It measures cross-domain generalization and compares separate versus joint multi-task learning protocols. Use when the user wants to benchmark on SiW, 3DMAD, HKBU-MarsV2, MSU-MFSD, 3DMask, ROSE-Youtu, FaceForensics++, DFDC, CelebDFv2, or asks about evaluating this task. Reports AUC, EER.
Evaluates a layer-granular DNN offloading framework by measuring inference latency and energy consumption across standard discriminative, generative, and autoencoder neural network architectures on mobile-cloud setups. Use when the user wants to benchmark on JointDNN Deep Architecture Benchmarks (AlexNet, OverFeat, VGG16, Deep Speech, ResNet, NiN, Chair, Pix2Pix), or asks about evaluating this task. Reports latency.
Compute Josh98/nl2bash_m via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Josh98/nl2bash_m.
Evaluates the quality of computed game-theoretic strategies in the word-guessing game Jotto by measuring expected guesses against a benchmark opponent and in self-play, as well as the equilibrium approximation error (epsilon). Use when the user wants to benchmark on Jotto (2-5 letter variants), or asks about evaluating this task. Reports epsilon.
Compute JP-SystemsX/nDCG via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of JP-SystemsX/nDCG.
Compute jpxkqx/peak_signal_to_noise_ratio via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jpxkqx/peak_signal_to_noise_ratio.
Compute jpxkqx/signal_to_reconstruction_error via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jpxkqx/signal_to_reconstruction_error.
Evaluates a model's ability to jointly recognize individual actions, social group activities, and global crowd-level activities in crowded panoramic scenes. It probes multi-granular activity recognition and hierarchical graph-based scene understanding. Use when the user wants to benchmark on JRDB-PAR, or asks about evaluating this task. Reports Overall F1 ($\mathcal{F}_a$).
Evaluates whether language models can generate JSON objects that strictly comply with a given JSON Schema under constrained decoding. It measures the empirical coverage of schema features supported by the model and decoding framework. Use when the user wants to benchmark on JSONSchemaBench, or asks about evaluating this task. Reports Top 1 Empirical Coverage.
Probes the factual and logical reliability of LLM-based judges and reward models by testing their ability to distinguish between objectively correct responses and subtly flawed ones across knowledge, reasoning, math, and coding domains. Use when the user wants to benchmark on JudgeBench, or asks about evaluating this task. Reports accuracy.
Compute juliakaczor/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of juliakaczor/accents_unplugged_eval.
Compute jzm-mailchimp/joshs_second_test_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jzm-mailchimp/joshs_second_test_metric.
Evaluates a model's ability to predict the number of k-barriers for intrusion detection in wireless sensor networks deployed over circular regions, based on geometric and deployment parameters. Use when the user wants to benchmark on WSN k-barrier simulation dataset, or asks about evaluating this task. Reports RMSE.
Evaluates automated exoplanet vetting pipelines on K2 transit lightcurves by comparing their planet candidate versus false positive dispositions against established ground truth from the NASA Exoplanet Archive. It probes the ability of tools to correctly identify true transiting planets while filtering out astrophysical false positives like eclipsing binaries and blended stars. Use when the user wants to benchmark on K2 Planet Candidate Catalog (DAVE Benchmark), or asks about evaluating this ...
Compute k4black/codebleu via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of k4black/codebleu.
Evaluates the ability of security tools to detect network misconfigurations in Kubernetes Helm charts, focusing on discrepancies between declared configurations and actual runtime behavior, including port exposure, label collisions, and network policy effectiveness. Use when the user wants to benchmark on Kubernetes Helm Charts (287 apps), or asks about evaluating this task. Reports misconfiguration detection (found/partially found/missed).
Evaluates an automated video classification pipeline for multi-species wildlife behavior monitoring by comparing machine learning predictions against expert manual annotations and traditional ground-based sampling methods. Use when the user wants to benchmark on KABR Drone Behavioral Dataset (custom), or asks about evaluating this task. Reports accuracy.
Evaluates models' capabilities in knowledge abstraction, concretization, and completion within entity-concept knowledge graphs. It specifically probes multi-hop reasoning through hierarchical relations and cross-view knowledge transfer between entities and concepts. Use when the user wants to benchmark on KACC, or asks about evaluating this task. Reports Hits@10.
Evaluates the zero-shot and few-shot generalization capability of text-to-SQL parsers on realistic, industrial-style database schemas with obscure column names and unrestricted natural language questions. It probes the model's ability to perform schema linking, constraint parsing, and SQL generation without extensive domain-specific training data. Use when the user wants to benchmark on KaggleDBQA, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates a provenance-based intrusion detection system's ability to identify anomalous system behavior and reconstruct attack footprints from whole-system kernel-level logs. It probes the model's capacity to distinguish between benign and malicious activity in temporal windows without relying on attack signatures. Use when the user wants to benchmark on Manzoor et al., DARPA-E3-THEIA, DARPA-E3-CADETS, DARPA-E3-ClearScope, DARPA-E5-THEIA, DARPA-E5-CADETS, DARPA-E5-ClearScope, DARPA-OpTC, or a...
This benchmark probes an LLM's ability to understand and generate culturally appropriate responses for Filipino contexts. It evaluates whether models can align with the lived experiences, values, and preferred strategies of action of average native Filipino speakers across nuanced socio-cultural scenarios. Use when the user wants to benchmark on Kalahi, or asks about evaluating this task. Reports MC1.
Evaluates large language models' ability to manipulate and apply stored knowledge across logical reasoning, reading comprehension, and natural language understanding tasks. It specifically probes the 'known & incorrect' phenomenon where models possess relevant facts but fail to apply them correctly during inference. Use when the user wants to benchmark on AbsR, Commonsense (Common), Big Bench Hard (BBH), RACE-H, RACE-M, MMLU, ARC-c, ARC-e, or asks about evaluating this task. Reports accuracy.
Evaluates multilingual vision-language reasoning by testing models on multiple-choice questions about images entirely in their native language. It probes cultural and linguistic authenticity, assessing how well models handle complex multimodal reasoning without relying on English translations. Use when the user wants to benchmark on Kaleidoscope, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to manipulate specific semantic attributes (e.g., technique, skill level) in human motion data while preserving untargeted attributes and anatomical accuracy. It probes latent space disentanglement and the capacity for controlled, attribute-level motion editing. Use when the user wants to benchmark on Kyokushin karate dataset, or asks about evaluating this task. Reports linear separability.
Compute kashif/mape via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kashif/mape.
Evaluates the acoustic quality, intelligibility, and diacritic sensitivity of a Kashmiri text-to-speech system. It probes the model's ability to accurately map Perso-Arabic script with explicit diacritics to natural-sounding speech and maintain spectral fidelity under low-resource conditions. Use when the user wants to benchmark on Curated Kashmiri Corpus, or asks about evaluating this task. Reports MCD, MOS.
Evaluates multilingual sentiment classification models on Kazakh customer reviews, probing their ability to handle code-switching, mixed scripts, and imbalanced class distributions across polarity and numerical score prediction tasks. Use when the user wants to benchmark on KazSAnDRA, or asks about evaluating this task. Reports macro-F1.
Evaluates a model's ability to learn first-order logic rules for knowledge base completion and object classification. It probes rule generation efficiency, scalability to longer rules, and few-shot generalization on relational data. Use when the user wants to benchmark on Even-and-Successor (ES), FB15K-237, WN18, Visual Genome (via GQA), or asks about evaluating this task. Reports MRR.
Compute kbmlcoding/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kbmlcoding/apps_metric.
Evaluates the ability of distantly supervised relation extraction and knowledge base validation systems to correctly predict and rank triples in web-scale knowledge graphs. It probes how well global graph structure and confidence scoring can refine noisy extractions and reduce logical inconsistencies. Use when the user wants to benchmark on NYT-FB, CC-DBP, NELL-165, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to answer complex questions over knowledge graphs by retrieving relevant subgraphs and generating correct answer entities. It probes multi-hop reasoning capabilities and robustness to missing edges in incomplete knowledge bases. Use when the user wants to benchmark on ComplexWebQuestions, WebQuestionsSP, WebQuestions, GrailQA, or asks about evaluating this task. Reports Hit@1.
Evaluates instruction-tuned LLMs on e-commerce tasks across two Amazon KDD Cup'24 tracks. It measures model performance on development and official test sets to assess retrieval, ranking, and generation capabilities in a commercial setting. Use when the user wants to benchmark on Amazon KDD Cup'24, or asks about evaluating this task. Reports scores.
Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on KDD CUP 99, or asks about evaluating this task. Reports accuracy.
Evaluates the capability of an unsupervised cellular automata-based framework to detect network intrusions and anomalous traffic patterns. It probes the model's ability to learn spatial-temporal dependencies from raw network connection records and distinguish between normal and malicious activities. Use when the user wants to benchmark on KDD Cup, or asks about evaluating this task. Reports Intrusion Detection Vs Positive Rate.
Evaluates a network intrusion detection system's ability to classify TCP/IP connections as either normal or one of several attack types based on 41 network features. It measures how well the model discriminates between benign traffic and specific intrusion categories such as DoS, Probe, R2L, and U2R. Use when the user wants to benchmark on KDD-Cup 99, or asks about evaluating this task. Reports accuracy.
Evaluates an intrusion detection system's ability to classify network traffic into normal and specific attack categories (probe, dos, u2r, r2l) using genetic algorithm-optimized feature selection and rule generation. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate (DR).