All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,406 views
Jet Tagging Resilience EvalA

Evaluates the trade-off between classification performance (AUC) and model resilience (robustness to Monte Carlo simulation variations) in quark/gluon and top-quark jet tagging. It probes whether complex neural architectures generalize better to different physics simulators compared to simpler, physics-informed models. Use when the user wants to benchmark on Pythia 8 / Herwig 7 Jet Samples, or asks about evaluating this task. Reports AUC.

researchpythonangular
0
3
Jfin Teb EvalA

Evaluates Japanese financial text embedding models across classification, retrieval, and clustering tasks. It probes the models' ability to capture domain-specific semantics, handle regulatory and economic terminology, and generalize to zero-shot financial scenarios. Use when the user wants to benchmark on JFinTEB, or asks about evaluating this task. Reports macro-F1.

ai-agentspythongo
0
3
Jialinsong Apps MetricA

Compute jialinsong/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jialinsong/apps_metric.

developmentpython
0
3
Jigsaw Puzzle Tiling EvalA

Evaluates object detection, recognition, and precise tiling assembly capabilities for modular robotic manipulation. It tests the pipeline's ability to segment oddly shaped pieces, recognize their identity, and assemble them with high spatial accuracy. Use when the user wants to benchmark on Jigsaw puzzle set, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Jigsaw Robot Manipulation EvalA

Evaluates spatial-temporal reasoning and hardware-agnostic robotic manipulation skills through a structured jigsaw puzzle assembly protocol. It measures vision-based segmentation, object recognition, pick planning success, and motion planning efficiency across three progressively complex physical tasks. Use when the user wants to benchmark on Jigsaw Manipulation Benchmark, or asks about evaluating this task. Reports Task Score.

researchpythongo
0
3
Jijihuny EcqaA

Compute jijihuny/ecqa via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jijihuny/ecqa.

developmentpython
0
3
Jina Embeddings V4 EvalA

Evaluates a multimodal embedding model's ability to retrieve relevant documents, images, and code from large corpora, and to measure semantic similarity between text pairs across multiple languages and modalities. Use when the user wants to benchmark on J-VDR, ViDoRe, CLIPB, MMTEB, MTEB-en, COIR, LEMB, STS-m, STS-en, or asks about evaluating this task. Reports nDCG@10.

ai-agentspythongo
0
3
Jjkim0807 Code EvalA

Compute jjkim0807/code_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jjkim0807/code_eval.

developmentpython
0
3
Jmedbench EvalA

Evaluates Japanese biomedical large language models across five tasks: multiple-choice question answering, named entity recognition, machine translation, document classification, and semantic text similarity. It probes domain-specific knowledge, multilingual comprehension, and in-context learning capabilities. Use when the user wants to benchmark on JMedBench, or asks about evaluating this task. Reports F1-entity, Accuracy.

researchpythongo
0
3
Jmmmu EvalA

This benchmark evaluates large multimodal models' ability to understand Japanese-language visual content and answer questions across multiple disciplines. It specifically probes the gap between general language translation capabilities (culture-agnostic subset) and deep cultural knowledge (culture-specific subset), revealing how models handle language variation bias and culturally grounded reasoning. Use when the user wants to benchmark on JMMMU, or asks about evaluating this task. Reports ac...

researchpythongo
0
3
Jmmmu Pro EvalA

Evaluates multimodal language models' ability to perform integrated visual-textual reasoning on Japanese-language tasks where questions and reference images are combined into a single composite image. It specifically probes OCR capabilities, visual perception, and cross-modal alignment in a multilingual context. Use when the user wants to benchmark on JMMMU-Pro, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Jobbert EvalA

Evaluates a model's ability to normalize job titles by mapping them to standardized ESCO occupation labels using semantic similarity. It probes the model's capacity to handle hierarchical occupational taxonomies and filter out irrelevant contextual tokens like locations. Use when the user wants to benchmark on JobBERT Vacancy Titles, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Jobfair Behavioral Transfer EvalA

Evaluates behavioral cloning and extraction fidelity in bias-mitigation LLM agents processing job descriptions. It measures how well mutated agent outputs align with expert golden references using cross-item similarity and a multi-facet diagnostic rubric. Use when the user wants to benchmark on JobFair Corpus, or asks about evaluating this task. Reports BERTScore (F1) against 5-NN average.

ai-agentspythongo
0
3
John Muir Ant EvalA

Evaluates the ability of a genetic algorithm to evolve a navigation program for an artificial ant to traverse a complex, toroidal grid trail with gaps and high-difficulty sections. The benchmark measures how well the evolved program generalizes beyond the standard Santa Fe trail into a chaotic extended sector. Use when the user wants to benchmark on John Muir Ant Problem, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Joint Face Spoofing Forgery EvalA

Evaluates models' ability to detect physical face spoofing attacks and digital face forgeries using visual appearance and physiological rPPG cues. It measures cross-domain generalization and compares separate versus joint multi-task learning protocols. Use when the user wants to benchmark on SiW, 3DMAD, HKBU-MarsV2, MSU-MFSD, 3DMask, ROSE-Youtu, FaceForensics++, DFDC, CelebDFv2, or asks about evaluating this task. Reports AUC, EER.

researchpythontesting
0
3
Jointdnn Benchmarks EvalA

Evaluates a layer-granular DNN offloading framework by measuring inference latency and energy consumption across standard discriminative, generative, and autoencoder neural network architectures on mobile-cloud setups. Use when the user wants to benchmark on JointDNN Deep Architecture Benchmarks (AlexNet, OverFeat, VGG16, Deep Speech, ResNet, NiN, Chair, Pix2Pix), or asks about evaluating this task. Reports latency.

researchpythonperformance
0
3
Josh98 Nl2bash MA

Compute Josh98/nl2bash_m via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Josh98/nl2bash_m.

developmentpythonbash
0
3
Jotto Game EvalA

Evaluates the quality of computed game-theoretic strategies in the word-guessing game Jotto by measuring expected guesses against a benchmark opponent and in self-play, as well as the equilibrium approximation error (epsilon). Use when the user wants to benchmark on Jotto (2-5 letter variants), or asks about evaluating this task. Reports epsilon.

researchpythongo
0
3
Jp Systemsx NdcgA

Compute JP-SystemsX/nDCG via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of JP-SystemsX/nDCG.

developmentpython
0
3
Jpxkqx Peak Signal To Noise RatioA

Compute jpxkqx/peak_signal_to_noise_ratio via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jpxkqx/peak_signal_to_noise_ratio.

developmentpython
0
3
Jpxkqx Signal To Reconstruction ErrorA

Compute jpxkqx/signal_to_reconstruction_error via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jpxkqx/signal_to_reconstruction_error.

developmentpython
0
3
Jrdb Par EvalA

Evaluates a model's ability to jointly recognize individual actions, social group activities, and global crowd-level activities in crowded panoramic scenes. It probes multi-granular activity recognition and hierarchical graph-based scene understanding. Use when the user wants to benchmark on JRDB-PAR, or asks about evaluating this task. Reports Overall F1 ($\mathcal{F}_a$).

researchpythongo
0
3
Jsonschemabench Coverage EvalA

Evaluates whether language models can generate JSON objects that strictly comply with a given JSON Schema under constrained decoding. It measures the empirical coverage of schema features supported by the model and decoding framework. Use when the user wants to benchmark on JSONSchemaBench, or asks about evaluating this task. Reports Top 1 Empirical Coverage.

researchpythongo
0
3
Judgebench EvalA

Probes the factual and logical reliability of LLM-based judges and reward models by testing their ability to distinguish between objectively correct responses and subtly flawed ones across knowledge, reasoning, math, and coding domains. Use when the user wants to benchmark on JudgeBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Juliakaczor Accents Unplugged EvalA

Compute juliakaczor/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of juliakaczor/accents_unplugged_eval.

developmentpython
0
3
Jzm Mailchimp Joshs Second Test MetricA

Compute jzm-mailchimp/joshs_second_test_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jzm-mailchimp/joshs_second_test_metric.

developmentpython
0
3
K Barrier Prediction EvalA

Evaluates a model's ability to predict the number of k-barriers for intrusion detection in wireless sensor networks deployed over circular regions, based on geometric and deployment parameters. Use when the user wants to benchmark on WSN k-barrier simulation dataset, or asks about evaluating this task. Reports RMSE.

researchpythonperformance
0
3
K2 Vetting EvalA

Evaluates automated exoplanet vetting pipelines on K2 transit lightcurves by comparing their planet candidate versus false positive dispositions against established ground truth from the NASA Exoplanet Archive. It probes the ability of tools to correctly identify true transiting planets while filtering out astrophysical false positives like eclipsing binaries and blended stars. Use when the user wants to benchmark on K2 Planet Candidate Catalog (DAVE Benchmark), or asks about evaluating this ...

researchpythongo
0
3
K4black CodebleuA

Compute k4black/codebleu via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of k4black/codebleu.

developmentpython
0
3
K8s Misconfig Detection EvalA

Evaluates the ability of security tools to detect network misconfigurations in Kubernetes Helm charts, focusing on discrepancies between declared configurations and actual runtime behavior, including port exposure, label collisions, and network policy effectiveness. Use when the user wants to benchmark on Kubernetes Helm Charts (287 apps), or asks about evaluating this task. Reports misconfiguration detection (found/partially found/missed).

devopspythongo
0
3
Kabr Drone Behavior EvalA

Evaluates an automated video classification pipeline for multi-species wildlife behavior monitoring by comparing machine learning predictions against expert manual annotations and traditional ground-based sampling methods. Use when the user wants to benchmark on KABR Drone Behavioral Dataset (custom), or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Kacc EvalA

Evaluates models' capabilities in knowledge abstraction, concretization, and completion within entity-concept knowledge graphs. It specifically probes multi-hop reasoning through hierarchical relations and cross-view knowledge transfer between entities and concepts. Use when the user wants to benchmark on KACC, or asks about evaluating this task. Reports Hits@10.

researchpythongo
0
3
Kaggledbqa EvalA

Evaluates the zero-shot and few-shot generalization capability of text-to-SQL parsers on realistic, industrial-style database schemas with obscure column names and unrestricted natural language questions. It probes the model's ability to perform schema linking, constraint parsing, and SQL generation without extensive domain-specific training data. Use when the user wants to benchmark on KaggleDBQA, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Kairos EvalA

Evaluates a provenance-based intrusion detection system's ability to identify anomalous system behavior and reconstruct attack footprints from whole-system kernel-level logs. It probes the model's capacity to distinguish between benign and malicious activity in temporal windows without relying on attack signatures. Use when the user wants to benchmark on Manzoor et al., DARPA-E3-THEIA, DARPA-E3-CADETS, DARPA-E3-ClearScope, DARPA-E5-THEIA, DARPA-E5-CADETS, DARPA-E5-ClearScope, DARPA-OpTC, or a...

researchpythongo
0
3
Kalahi EvalA

This benchmark probes an LLM's ability to understand and generate culturally appropriate responses for Filipino contexts. It evaluates whether models can align with the lived experiences, values, and preferred strategies of action of average native Filipino speakers across nuanced socio-cultural scenarios. Use when the user wants to benchmark on Kalahi, or asks about evaluating this task. Reports MC1.

researchpythongit
0
3
Kale EvalA

Evaluates large language models' ability to manipulate and apply stored knowledge across logical reasoning, reading comprehension, and natural language understanding tasks. It specifically probes the 'known & incorrect' phenomenon where models possess relevant facts but fail to apply them correctly during inference. Use when the user wants to benchmark on AbsR, Commonsense (Common), Big Bench Hard (BBH), RACE-H, RACE-M, MMLU, ARC-c, ARC-e, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Kaleidoscope EvalA

Evaluates multilingual vision-language reasoning by testing models on multiple-choice questions about images entirely in their native language. It probes cultural and linguistic authenticity, assessing how well models handle complex multimodal reasoning without relying on English translations. Use when the user wants to benchmark on Kaleidoscope, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Karate Attribute Manipulation EvalA

Evaluates a model's ability to manipulate specific semantic attributes (e.g., technique, skill level) in human motion data while preserving untargeted attributes and anatomical accuracy. It probes latent space disentanglement and the capacity for controlled, attribute-level motion editing. Use when the user wants to benchmark on Kyokushin karate dataset, or asks about evaluating this task. Reports linear separability.

researchpythonperformance
0
3
Kashif MapeA

Compute kashif/mape via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kashif/mape.

developmentpython
0
3
Kashmiri Tts EvalA

Evaluates the acoustic quality, intelligibility, and diacritic sensitivity of a Kashmiri text-to-speech system. It probes the model's ability to accurately map Perso-Arabic script with explicit diacritics to natural-sounding speech and maintain spectral fidelity under low-resource conditions. Use when the user wants to benchmark on Curated Kashmiri Corpus, or asks about evaluating this task. Reports MCD, MOS.

researchpythongit
0
3
Kazsandra EvalA

Evaluates multilingual sentiment classification models on Kazakh customer reviews, probing their ability to handle code-switching, mixed scripts, and imbalanced class distributions across polarity and numerical score prediction tasks. Use when the user wants to benchmark on KazSAnDRA, or asks about evaluating this task. Reports macro-F1.

researchpythongo
0
3
Kb Completion EvalA

Evaluates a model's ability to learn first-order logic rules for knowledge base completion and object classification. It probes rule generation efficiency, scalability to longer rules, and few-shot generalization on relational data. Use when the user wants to benchmark on Even-and-Successor (ES), FB15K-237, WN18, Visual Genome (via GQA), or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Kbmlcoding Apps MetricA

Compute kbmlcoding/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kbmlcoding/apps_metric.

developmentpython
0
3
Kbp Auc EvalA

Evaluates the ability of distantly supervised relation extraction and knowledge base validation systems to correctly predict and rank triples in web-scale knowledge graphs. It probes how well global graph structure and confidence scoring can refine noisy extractions and reduce logical inconsistencies. Use when the user wants to benchmark on NYT-FB, CC-DBP, NELL-165, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Kbqa Hit1 EvalA

Evaluates a model's ability to answer complex questions over knowledge graphs by retrieving relevant subgraphs and generating correct answer entities. It probes multi-hop reasoning capabilities and robustness to missing edges in incomplete knowledge bases. Use when the user wants to benchmark on ComplexWebQuestions, WebQuestionsSP, WebQuestions, GrailQA, or asks about evaluating this task. Reports Hit@1.

researchpythongo
0
3
Kdd Cup 24 EvalA

Evaluates instruction-tuned LLMs on e-commerce tasks across two Amazon KDD Cup'24 tracks. It measures model performance on development and official test sets to assess retrieval, ranking, and generation capabilities in a commercial setting. Use when the user wants to benchmark on Amazon KDD Cup'24, or asks about evaluating this task. Reports scores.

researchpythongo
0
3
Kdd Cup 99 EvalA

Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on KDD CUP 99, or asks about evaluating this task. Reports accuracy.

researchpythonnode
0
3
Kdd Cup EvalA

Evaluates the capability of an unsupervised cellular automata-based framework to detect network intrusions and anomalous traffic patterns. It probes the model's ability to learn spatial-temporal dependencies from raw network connection records and distinguish between normal and malicious activities. Use when the user wants to benchmark on KDD Cup, or asks about evaluating this task. Reports Intrusion Detection Vs Positive Rate.

researchpythongo
0
3
Kdd99 Accuracy EvalA

Evaluates a network intrusion detection system's ability to classify TCP/IP connections as either normal or one of several attack types based on 41 network features. It measures how well the model discriminates between benign traffic and specific intrusion categories such as DoS, Probe, R2L, and U2R. Use when the user wants to benchmark on KDD-Cup 99, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Kdd99 Ids EvalA

Evaluates an intrusion detection system's ability to classify network traffic into normal and specific attack categories (probe, dos, u2r, r2l) using genetic algorithm-optimized feature selection and rule generation. Use when the user wants to benchmark on KDD99, or asks about evaluating this task. Reports Detection Rate (DR).

researchpythongo
0
3