
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates vision-language models' ability to perform multimodal scientific reasoning in chemistry and materials research. It probes capabilities across data extraction, experimental understanding, and data interpretation, specifically testing spatial reasoning, cross-modal synthesis, and multi-step inference. Use when the user wants to benchmark on MaCBench, or asks about evaluating this task. Reports accuracy.
Evaluates a low-resource language model's capability on standard commonsense reasoning, reading comprehension, and factual knowledge tasks adapted to Macedonian. It measures how well continued pretraining and instruction tuning improve performance on these benchmarks compared to multilingual baselines. Use when the user wants to benchmark on Macedonian Benchmarks (ARC Easy, ARC Challenge, BoolQ, HellaSwag, OpenBookQA, PIQA, WinoGrande), or asks about evaluating this task. Reports accuracy.
Evaluates an LLM agent's susceptibility to unethical steering prompts in a text-based adventure game environment. It probes the capability of anomaly detection systems to classify agent trajectories as ethical or unethical based on their interaction traces. Use when the user wants to benchmark on MACHIAVELLI, or asks about evaluating this task. Reports AUPRC.
Evaluates a model's ability to detect out-of-distribution (OOD) malware variants and classify known malware families without using OOD samples during training. It probes both classification accuracy on in-distribution data and the statistical separation capability between known and novel threats using cluster-driven decision boundaries. Use when the user wants to benchmark on Unspecified malware dataset (25 families), or asks about evaluating this task. Reports AUROC.
Evaluates end-to-end autonomous materials discovery pipelines by measuring how effectively different policies (planners, generators, selectors, and agentic orchestrators) can find thermodynamically stable compounds under constrained oracle query budgets. It probes the trade-offs between discovery efficiency, structural diversity, and adaptivity as chemical complexity and stability thresholds increase. Use when the user wants to benchmark on MADE Benchmark Environments, or asks about evaluatin...
Evaluates the multilingual machine translation and zero-shot/few-shot translation capabilities of models trained on the MADLAD-400 dataset. It probes cross-lingual generalization, low-resource language handling, and the impact of data auditing on translation quality across multiple benchmarks and language pairs. Use when the user wants to benchmark on WMT, Flores-200, NTREX, GATONES, or asks about evaluating this task. Reports BLEU.
Evaluates the short-term traffic flow forecasting capability of various machine learning and deep learning models on real-world urban road networks. It specifically probes how well models generalize across different traffic profiles and prediction horizons while maintaining computational efficiency. Use when the user wants to benchmark on Madrid Traffic Dataset, or asks about evaluating this task. Reports R^2 (Coefficient of determination).
Evaluates a multimodal AI model's ability to predict clinical outcomes and adverse reactions for drug combinations from preclinical data. It probes robustness to missing modalities and generalization to novel drugs under strict hold-out splits. Use when the user wants to benchmark on TWOSIDES, DrugBank, or asks about evaluating this task. Reports AUROC.
Evaluates audio embedding models across 30 tasks spanning speech, music, environmental sounds, bioacoustics, emotion recognition, and cross-modal audio-text reasoning in over 100 languages. It probes the ability of models to generalize across acoustic domains, handle multilingual alignment, and perform both supervised and unsupervised audio understanding tasks. Use when the user wants to benchmark on MAEB, or asks about evaluating this task. Reports Average Score.
This evaluation protocol assesses the transfer learning capability of self-supervised and supervised vision models on multimodal, multitemporal, and multispectral Earth observation data. It probes downstream performance on tree species classification and agricultural/land cover segmentation tasks across varying dataset scales and fusion strategies. Use when the user wants to benchmark on TreeSatAI-TS, PASTIS-HD, FLAIR#2, FLAIR-HUB, or asks about evaluating this task. Reports weighted F1 score...
Evaluates the capability of models to generate high-fidelity alpha mattes for multiple human instances in images and videos, focusing on detail preservation, instance separation, and temporal consistency across frames. Use when the user wants to benchmark on HIM2K+M-HIM2K, V-HIM60, or asks about evaluating this task. Reports MAD.
Evaluates the predictive performance of mean-aggregation GNNs with non-negative weights on link prediction and node classification tasks, while assessing the soundness, monotonicity, and logical complexity of the extracted explanatory rules. Use when the user wants to benchmark on WN18RRv1, FB237v1, NELLv1, LUBM, LogInfer-WN-hier, LogInfer-WN-sym, LogInfer-WN-hier_nmhier, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of a self-supervised graph representation learning model to detect Advanced Persistent Threats (APTs) in system audit logs. It probes multi-granularity anomaly detection (batched log-level and system entity-level) under a strict unsupervised setting where only benign data is available for training. Use when the user wants to benchmark on StreamSpot, Unicorn Wget, DARPA Engagement 3, or asks about evaluating this task. Reports Precision.
Evaluates generative models on cartoon animation tasks including audio-driven facial animation, face reenactment, image-to-video generation, and frame interpolation. It probes a model's ability to produce stylized, temporally consistent video with accurate facial details and cross-modal alignment. Use when the user wants to benchmark on MagicAnime-Bench, or asks about evaluating this task. Reports VSR.
Evaluates instruction-based image editing models on their ability to modify source images according to text instructions while preserving original content and style. It tests both single-turn and multi-turn editing capabilities against ground truth edited images. Use when the user wants to benchmark on MagicBrush, or asks about evaluating this task. Reports CLIP image similarity.
Evaluates text-to-image generation models on their ability to produce images free of fine-grained artifacts, specifically probing subject anatomy, attributes, and interactions. It measures detection accuracy using a hierarchical taxonomy of artifact types to benchmark model robustness against visual inconsistencies. Use when the user wants to benchmark on MagicData340K, or asks about evaluating this task. Reports F1-Score.
Evaluates a lightweight vision-language model's performance on reasoning, OCR, and real-world understanding benchmarks, alongside deployment efficiency metrics like inference latency and throughput on mobile hardware. Use when the user wants to benchmark on HallusionBench, MMBench, RealworldQA, MMStar, OCRBench, AI2D, TextVQA, CRPE, MME Realworld, DocVQA, or asks about evaluating this task. Reports accuracy.
This protocol evaluates a model's capability to perform multimodal agentic tasks, specifically UI navigation and robotic manipulation. It probes spatial-temporal reasoning, action grounding, and zero-shot or few-shot transfer across digital interfaces and physical simulators. Use when the user wants to benchmark on ScreenSpot, VisualWebBench, SimplerEnv, Mind2Web, AITW, LIBERO, or asks about evaluating this task. Reports step_success_rate.
Evaluates a model-free runtime system for dynamically scaling uncore frequencies in heterogeneous CPU-GPU architectures. It probes the system's ability to balance energy efficiency and performance across diverse HPC, molecular dynamics, and deep learning workloads. Use when the user wants to benchmark on Altis, ECP proxy applications, AI-enabled applications, MLPerf benchmarks, Altis-SYCL, or asks about evaluating this task. Reports Energy Delay Product (EDP).
Evaluates the accuracy and inference speed of monocular depth estimation models on resource-constrained mobile devices. It measures depth prediction quality using invariant standard root mean squared error and records inference time on a Raspberry Pi 4 to assess real-time capability. Use when the user wants to benchmark on MAI&AIM2022 challenge dataset, or asks about evaluating this task. Reports si-RMSE.
Evaluates the instruction-following capability, output preference quality, and general reasoning performance of fine-tuned LLMs. It probes how well models adhere to explicit constraints, generate preferred responses relative to a baseline, and solve standard academic benchmarks. Use when the user wants to benchmark on AlpacaEval, IFEval, ARC, HellaSwag, Winogrande, MMLU, TruthfulQA, or asks about evaluating this task. Reports AlpacaEval.
Probes large language models' ability to reason about legacy mainframe systems, interpret COBOL code, and generate accurate technical summaries. It tests domain-specific code understanding through multiple-choice questions, open-ended QA, and text generation tasks. Use when the user wants to benchmark on MainframeBench, or asks about evaluating this task. Reports Accuracy.
Evaluates retrieval models' ability to follow complex, task-specific instructions across diverse domains and long-tail tasks. It measures how instruction tuning impacts generalization and performance on heterogeneous query-document relevance tasks compared to non-instruction-tuned baselines. Use when the user wants to benchmark on MAIR, or asks about evaluating this task. Reports nDCG@10.
Evaluates the robustness and consistency of a model's mathematical reasoning by aggregating predictions across multiple inference runs per problem. It measures whether the most frequent prediction among k samples matches the ground truth. Use when the user has predictions and gold and needs to compute Majority@k.
Compute maksymdolgikh/seqeval_with_fbeta via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maksymdolgikh/seqeval_with_fbeta.
Evaluates radiative transfer models' ability to simulate exoplanet transit and direct-imaging spectra under controlled atmospheric conditions. Probes how atmospheric discretization, opacity treatments, and spectroscopic databases impact spectral predictions. Use when the user wants to benchmark on MALBEC Test Suite, or asks about evaluating this task. Reports ppm.
Evaluates a model's ability to classify malware samples into one of seven predefined families based on their opcode sequences. It probes the model's capacity to capture structural and sequential patterns in low-level code representations. Use when the user wants to benchmark on Malicia, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the ability of machine learning models to classify malware binaries by converting them into grayscale images and predicting their specific family among 25 categories. It probes classification accuracy and computational efficiency across different neural network architectures. Use when the user wants to benchmark on Malimg, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of deep learning models to classify malware families from their visual representations (malware images). It probes feature extraction robustness and classification accuracy on an imbalanced dataset. Use when the user wants to benchmark on MalImg, or asks about evaluating this task. Reports Accuracy.
Evaluates memory-aware long sequence compression techniques in large-scale sequential recommendation. It probes how well methods balance memory overhead, computational cost, and ranking accuracy when processing long user interaction histories to predict future clicks. Use when the user wants to benchmark on Amazon-Electronic, MicroVideo1.7M, KuaiVideo, or asks about evaluating this task. Reports AUC.
Evaluates graph neural networks for Android malware family classification under intra-family and cross-family distribution shifts. It probes how semantic feature enrichment (function metadata and LLM embeddings) and test-time/domain adaptation methods mitigate performance degradation when models encounter unseen malware families. Use when the user wants to benchmark on MalNet-Tiny, MalNet-Tiny-Common, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness of deep neural network malware classifiers against adversarial attacks. It probes whether an attacker can successfully misclassify malicious Android applications by adding a limited number of valid features to their manifest files. Use when the user wants to benchmark on DREBIN, or asks about evaluating this task. Reports misclassification rate.
This benchmark evaluates an AI system's ability to analyze low-level process execution logs and identify malicious signals from malware detonations. It probes structured data parsing, security event correlation, and malware family classification capabilities. Use when the user wants to benchmark on CyberSOCEval Malware Analysis, or asks about evaluating this task. Reports accuracy.
Evaluates a CNN's ability to classify Windows PE files as malware or benign by learning spatial features from grayscale images generated from dynamic API call argument sequences. It probes the model's resilience to obfuscation by leveraging behavioral temporal patterns converted into visual representations. Use when the user wants to benchmark on Windows PE Malware Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates host-based intrusion detection models on their ability to classify multi-label malware behaviors from truncated Windows API call sequences. Probes how well different neural architectures handle sequential behavioral data and feature selection strategies for detecting overlapping malicious activities. Use when the user wants to benchmark on Behavioural Reports of Multi-Stage Malware, or asks about evaluating this task. Reports F1-score.
This evaluation probes a model's ability to classify long-sequence binary and executable files into malware families or benign/malicious categories. It tests robustness to varying sequence lengths, compression formats, and real-world malware distribution characteristics compared to synthetic long-range benchmarks. Use when the user wants to benchmark on Kaggle (BIG 2015), Drebin, EMBER, LRA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the effectiveness of unsupervised clustering algorithms on large-scale malware binary datasets. It probes how well different feature representations and clustering methods can group malware samples into coherent families while handling real-world noise and benign samples. Use when the user wants to benchmark on Bodmas, Ember, Security, or asks about evaluating this task. Reports Homogeneity.
Binary classification of software binaries as benign or malicious based on their control flow graphs. It probes the model's ability to learn graph-structured representations and route them through specialized experts for accurate detection. Use when the user wants to benchmark on BODMAS, DikeDataset, PMML, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of LLMs and their ensembles to correctly classify malware samples into one of ten canonical families based on their behavior or code semantics. It probes robustness to class imbalance and the effectiveness of hierarchical decision-making under obfuscation. Use when the user wants to benchmark on Gold-standard malware family dataset, or asks about evaluating this task. Reports Macro F1-score.
Evaluates the capability of anti-malware detection engines and behavior profiling methods to correctly group malware variants into their respective families based on runtime Windows API call sequences and parameters. Use when the user wants to benchmark on 40Bot, 419Mal, or asks about evaluating this task. Reports Pairwise Classification Score (PCS).
Evaluates static, dynamic, and hybrid analysis pipelines for malware detection. Models are trained on opcode and API call sequences to distinguish malware families from benign Windows executables. Use when the user wants to benchmark on Malware Detection Dataset, or asks about evaluating this task. Reports Area under the ROC curve.
Evaluates a deep learning model's ability to classify malware images across multiple tasks, including binary classification, malware family classification, and detection of obfuscation techniques across Windows, Android, macOS, and Linux platforms. Use when the user wants to benchmark on Maling benchmark dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of machine learning and deep learning models to classify Windows PE binaries as benign or malicious, and further categorize malicious samples into specific families or ransomware types using static analysis features and grayscale image representations. Use when the user wants to benchmark on Ember, Bodmas, PEMachineLearning, or asks about evaluating this task. Reports accuracy.
Evaluates a static analysis tool's ability to detect malware by extracting system call data flow trees and matching them against a learned tree automaton. It probes the model's capability to generalize semantic malware signatures from a small training set to a larger, unseen test set while avoiding false positives on benign software. Use when the user wants to benchmark on VX Heavens & Windows XP Benign, or asks about evaluating this task. Reports detection_rate.
Evaluates how well vision-language models and CNNs capture clinically relevant mammography concepts at the neuron level. It quantifies concept coverage, alignment strength, and how domain-specific pretraining or task-specific fine-tuning shifts learned representations. Use when the user wants to benchmark on VinDR-Mammo, EMBED, or asks about evaluating this task. Reports unique_concepts_captured.
Evaluates a breast-specific foundational model's ability to generalize across in-distribution and out-of-distribution mammographic datasets for zero-shot diagnosis, linear probing, full fine-tuning, and pathology localization. It probes the model's robustness, data efficiency, and representation quality for clinical tasks like cancer detection and risk prediction. Use when the user wants to benchmark on EMBED, VinDr, RSNA, or asks about evaluating this task. Reports AUROC.
Evaluates a federated learning framework for quantitative breast density estimation from mammographic images. It probes the model's ability to segment breast and dense tissue, predict percent density, and generalize across different medical institutions while preserving patient privacy. Use when the user wants to benchmark on MC, UPHS, or asks about evaluating this task. Reports PD MAE.
Evaluates lightweight CNN architectures for pixel-wise lesion segmentation in mammograms. It measures segmentation accuracy and computational efficiency, while also probing cross-dataset generalization under domain shift and the sensitivity of performance metrics to post-processing thresholds. Use when the user wants to benchmark on INbreast, DMID, or asks about evaluating this task. Reports Dice Score.
Evaluates deep CNN architectures for classifying mammographic abnormalities (calcifications and masses) and localizing them using class activation maps. It probes the model's ability to learn patch-based features and generalize to full-image localization without explicit spatial supervision. Use when the user wants to benchmark on Mammography dataset (unspecified), or asks about evaluating this task. Reports accuracy.
Probes a deep learning model's ability to classify mammograms into BI-RADS categories (normal, benign, malignant) and localize suspicious lesions using weakly and semi-supervised learning. It evaluates both image-level diagnostic accuracy and region-level detection performance under clinically relevant operating points. Use when the user wants to benchmark on IMG, INbreast, or asks about evaluating this task. Reports AUROC.