
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates native full-duplex audio-language models on speech understanding, speech generation, and real-time conversational capabilities. It measures how well models handle asynchronous text-audio streams, responsiveness to interruptions, and overall dialogue quality compared to specialized ASR/TTS systems and other full-duplex chatbots. Use when the user wants to benchmark on Fleurs-zh, LibriSpeech-clean, LlamaQuestions, Seed-TTS-en, Seed-TTS-zh, Custom Chinese Speech Instruction-Following S...
Evaluates a model's ability to predict river water levels and forecast floods using spatiotemporal radar precipitation data. It probes the model's accuracy across multiple forecasting lead times (2h to 12h) and its robustness in capturing extreme hydrological events compared to baseline and deep learning models. Use when the user wants to benchmark on Goslar, Göttingen, or asks about evaluating this task. Reports NSE.
This protocol evaluates the computational throughput and real-time efficiency of embedded CPU and GPU platforms. It measures peak floating-point operations per second (FLOPS) using a controlled matrix rotation kernel, and assesses practical system performance via end-to-end inference latency and power consumption on a robotic vision pipeline. Use when the user has predictions and gold and needs to compute FLOPS.
Evaluates a unified vision foundation model's zero-shot and fine-tuned capabilities across diverse computer vision tasks. It probes the model's ability to perform image captioning, visual question answering, object detection, referring expression comprehension, and semantic segmentation using a single sequence-to-sequence architecture. Use when the user wants to benchmark on COCO, Flickr30k, RefCOCO/+/g, VQAv2, ADE20K, or asks about evaluating this task. Reports CIDEr.
Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports XSTS.
Evaluates the translation quality of Neural Machine Translation (NMT) models trained on filtered pseudo-parallel corpora. It measures how well few-shot Quality Estimation (QE) based corpus filtering improves MT performance across low-resource and mid-resource language pairs compared to baselines and other filtering methods. Use when the user wants to benchmark on FLORES101, or asks about evaluating this task. Reports BLEU.
Evaluates multilingual machine translation quality across 60 languages and 234 translation directions. It specifically probes a model's ability to handle high-, medium-, and low-resource languages while mitigating directional degeneration in symmetric multi-way translation. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports COMET-22.
Evaluates the accuracy of predicted optical flow fields on spherical 360° video frames, measuring both endpoint displacement and angular deviation. It also assesses egocentric activity recognition performance using these flow features to test rotation-invariant representation learning. Use when the user wants to benchmark on FLOW360, EGOK360, or asks about evaluating this task. Reports EPE.
Evaluates computational workflow anomaly detection by benchmarking models on detecting injected CPU and HDD performance anomalies in distributed workflow execution logs. It tests the ability of tabular, graph, and text-based methods to identify anomalous nodes within Directed Acyclic Graph (DAG) workflow executions. Use when the user wants to benchmark on Flow-Bench, or asks about evaluating this task. Reports ROC-AUC.
Evaluates the scalability, heterogeneity handling, realism, and privacy overhead of the Flower federated learning framework across various datasets and device configurations. Use when the user wants to benchmark on Amazon Book Reviews, FEMNIST, RealWorld, CIFAR-10, FashionMNIST, or asks about evaluating this task. Reports training time.
This evaluation protocol assesses the effectiveness of various transformer-based architectures for flow-based network intrusion detection. It systematically tests different input encodings, transformer blocks, and classification heads across three standard NIDS datasets to determine optimal configurations for accuracy, model size, and inference speed. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, CSE-CIC-IDS2018, or asks about evaluating this task. Reports F1 score.
Evaluates a network intrusion detection model's ability to classify benign versus malicious traffic flows in real-world IoT environments. It specifically probes robustness to severe class imbalance, feature sparsity mitigation via context-aware embeddings, and temporal generalization across different time periods. Use when the user wants to benchmark on MAWI, or asks about evaluating this task. Reports F1-Score.
Probes French sequence classification capabilities across sentiment analysis, paraphrase identification, and natural language inference. Use when the user wants to benchmark on CLS, PAWSX, XNLI, or asks about evaluating this task. Reports Accuracy.
Evaluates reinforcement learning algorithms for active flow control tasks, measuring their ability to stabilize fluid dynamics and reduce drag or enhance heat transfer. It probes algorithmic robustness, sample efficiency, and the capacity to transfer policies across dimensionalities and domain sizes. Use when the user wants to benchmark on FluidGym, or asks about evaluating this task. Reports mean reward per step.
Evaluates the ability of reinforcement learning and trajectory optimization algorithms to control complex, multi-phase fluid systems interacting with rigid bodies. It probes sample efficiency, gradient-based optimization stability, and sim-to-real transfer in high-dimensional, non-smooth fluid dynamics. Use when the user wants to benchmark on FluidLab, or asks about evaluating this task. Reports accumulated reward.
Evaluates how well NLP models maintain performance when subjected to minimal, linguistically-grounded perturbations (e.g., syntactic voice changes, negation, style shifts, geographical/temporal biases) across classification and generation tasks. It probes model brittleness to covariate shifts introduced by natural language modifications rather than adversarial noise. Use when the user wants to benchmark on KnowRef, Few-NERD, GSM8K, IFEval, or asks about evaluating this task. Reports Unrobustn...
Evaluates the runtime performance of an LLM serving engine under bursty, heterogeneous, and long-context workloads. It probes the system's ability to dynamically switch between data and tensor parallelism to optimize latency and throughput while maintaining memory efficiency compared to static and alternative dynamic baselines. Use when the user wants to benchmark on ShareGPT, CodeActInstruct, HumanEval, Synthetic Workloads, or asks about evaluating this task. Reports TTFT.
Evaluates music information retrieval models on genre classification tasks using a large-scale, open music dataset. It probes the model's ability to map audio tracks to hierarchical genre labels (single-label or multi-label) using raw audio or precomputed features. Use when the user wants to benchmark on FMA (Free Music Archive), or asks about evaluating this task. Reports accuracy.
Compute fnvls/bleu_1234 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of fnvls/bleu_1234.
Compute fnvls/bleu1234 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of fnvls/bleu1234.
This evaluation probes a model's ability to classify Android malware by fusing audio and visual representations derived from raw APK binaries. It measures supervised classification performance across multiple malware families and benign samples using standard accuracy and macro-F1 metrics. Use when the user wants to benchmark on CICMalDroid-2020, Mal-Net, or asks about evaluating this task. Reports Accuracy.
Evaluates end-to-end reasoning, error propagation, and user experience in multi-modal cascading agents. It assesses technical performance (ASR/TTS fidelity, tool calling) and behavioral quality (reasoning, semantic similarity, contextual consistency) across simulated customer service journeys. Use when the user wants to benchmark on FOCAL Customer Journeys, or asks about evaluating this task. Reports Accuracy Similarity.
Evaluates visual object tracking performance specifically for anti-UAV scenarios, probing a model's ability to maintain target localization under abrupt camera motion, extreme scale variations, and small target sizes in thermal infrared imagery. Use when the user wants to benchmark on AntiUAV, AntiUAV410, or asks about evaluating this task. Reports AUC.
Evaluates a discrete-event simulation framework that fuses dynamic scene graphs with urban environments to model hierarchical, interconnected spaces under partial observability. It probes the simulator's capacity to reproduce emergent temporal behaviors, the accuracy of state reconstruction when agent views are sparse, and the computational efficiency of the underlying simulation engine. Use when the user wants to benchmark on FOGMACHINE Scenarios (Bruchsal, Wenningstedt, Trier), or asks abou...
Evaluates the ability of state-of-the-art audio deepfake detectors to distinguish real speech from face-to-voice (FOICE) synthesized speech, and assesses how fine-tuning on FOICE data affects robustness against unseen synthesis pipelines like SpeechT5. Use when the user wants to benchmark on FOICE, SpeechT5, or asks about evaluating this task. Reports EER.
Evaluates how fairness interventions affect predictive accuracy and fairness violations across different geographic regions and time periods. It probes the stability of fairness metrics under distribution shift and the efficacy of pre-processing, in-processing, and post-processing interventions on tabular demographic data. Use when the user wants to benchmark on Folktables (ACS PUMS), or asks about evaluating this task. Reports accuracy.
Evaluates the calibration and predictive accuracy of language models when used as risk scorers for tabular prediction tasks. It probes whether models can accurately quantify outcome uncertainty (calibration) while maintaining discriminative power (AUC) on natural-language versions of tabular datasets. Use when the user wants to benchmark on folktexts, or asks about evaluating this task. Reports ECE.
Evaluates whether information retrieval models can follow complex, long-form instructions derived from TREC narratives to determine document relevance. It probes the model's ability to interpret conditional, negated, and composite relevance criteria rather than relying solely on keyword matching. Use when the user wants to benchmark on Robust04, News21, Core17, or asks about evaluating this task. Reports p-MRR.
Evaluates fine-grained semantic segmentation and ingredient localization in food images. It probes a model's ability to handle pixel-wise mask prediction under high appearance variability, long-tailed class distributions, and cross-domain generalization to unseen cuisines. Use when the user wants to benchmark on FoodSeg103, or asks about evaluating this task. Reports mIoU.
This benchmark evaluates the robustness and accuracy of Text-to-SQL systems when translating natural language questions into SQL queries across different database schema designs. It probes how data model complexity, training data size, and language model scale impact execution accuracy on real-world user queries. Use when the user wants to benchmark on FootballDB, or asks about evaluating this task. Reports exact execution matching (EX).
Evaluates the quality of universal image embeddings for flat object retrieval across diverse 2D domains (e.g., logos, paintings, currency) under varying visual distortions. It probes both candidate rank accuracy and the matching score margin to assess out-of-distribution generalization. Use when the user wants to benchmark on FORB, or asks about evaluating this task. Reports mAP@5.
Evaluates hierarchical multi-label classification of academic papers into a taxonomy of 170 research fields. It tests zero-shot/few-shot prompting and weakly-labeled data integration for field-of-research prediction. Use when the user wants to benchmark on FoRC4CL 2025, or asks about evaluating this task. Reports Micro-F1.
Evaluates a video generation model's ability to adhere to specified physics-based force signals (local point forces or global wind fields) and produce visually realistic, physically plausible dynamics. It probes generalization across diverse objects, materials, and motion categories using human preference judgments. Use when the user wants to benchmark on Local Point Force Benchmark, Global Force Benchmark, or asks about evaluating this task. Reports 2AFC win rate.
This benchmark evaluates Large Vision Language Models (LVLMs) on their ability to detect and attribute image or video forgeries. It probes generalization and reasoning capabilities across five dimensions: semantics, modalities, tasks, forgery types, and generation models, using multi-choice visual questions. Use when the user wants to benchmark on Forensics-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform temporally grounded, multimodal video understanding in surveillance settings. It probes precise temporal localization, identity-based search, and complex reasoning over long videos using text-only or image+text queries. Use when the user wants to benchmark on ForeSeaQA, or asks about evaluating this task. Reports accuracy.
Evaluates a fine-tuned LLM's ability to predict future biomedical concepts and clinical disorders from patient clinical timelines. It measures concept prediction accuracy via precision and recall across different temporal windows and candidate counts, and assesses clinical risk forecasting by checking how many of the top-5 predicted disorders match the ground truth for the next month. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports Precision.
Evaluates a vision-language agent's ability to perform joint change detection and natural language captioning on bi-temporal remote sensing imagery. It probes the model's capacity to segment deforestation and built-environment changes at the pixel level while generating accurate semantic descriptions of those changes. Use when the user wants to benchmark on Forest-Change, LEVIR-MCI-Trees, or asks about evaluating this task. Reports MIoU.
Evaluates multimodal large language models on fine-grained manufacturing tasks, including workpiece verification, surface defect inspection, and assembly verification. It probes the models' ability to combine visual grounding with domain-specific knowledge to identify anomalies or classify conditions. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates the quality of generated Semantic Identifiers (SIDs) for generative retrieval in industrial recommendation and search systems. It measures how well SIDs capture item relationships and distribute usage fairly, and assesses their impact on downstream retrieval hitrate and online transaction metrics. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports HR@K.
Evaluates pixel-level localization accuracy for detecting AI-generated and traditionally tampered image forgeries. It probes a model's ability to distinguish manipulated regions from authentic content by measuring spatial overlap and detection trade-offs against ground-truth masks. Use when the user wants to benchmark on OpenSDID, GIT10K, CocoGlide, Inpaint32K, IMD2020, NIST16, CASIA, or asks about evaluating this task. Reports F1-score.
This benchmark probes the ability of multimodal large language models to recognize regular and irregular geometric shapes from images and accurately count their sides. It further evaluates multi-step visual-mathematical reasoning by requiring models to identify multiple shapes, map them to side counts, and compute their sum. Use when the user wants to benchmark on Forgotten Polygons, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of auxiliary-task learning methods to mitigate negative transfer and improve target task performance across multi-task, multi-domain, and semi-supervised learning settings. It probes how well a model can dynamically combine or select auxiliary tasks without degrading the primary task. Use when the user wants to benchmark on NYUv2, DomainNet, AliExpress, CIFAR-10, SVHN, or asks about evaluating this task. Reports Δm.
Evaluates large language models' domain knowledge in petroleum geoscience using a 505-question multiple-choice benchmark. It probes understanding across seven specialized subdomains, including petrophysics, reservoir engineering, and drilling, while measuring performance variance by model size, cost, and question difficulty. Use when the user wants to benchmark on FormationEval, or asks about evaluating this task. Reports Accuracy.
Evaluates large language models and speech systems on three endangered Formosan Austronesian languages (Atayal, Amis, Paiwan) across machine translation, automatic speech recognition, and text summarization. It probes zero-shot, few-shot (10-shot), and fine-tuning adaptation capabilities in typologically complex, low-resource settings. Use when the user wants to benchmark on FormosanBench, or asks about evaluating this task. Reports BLEU.
Evaluates the ability of PDF document parsers to accurately extract mathematical formulas and preserve their semantic meaning. It probes format variability handling, representational non-uniqueness, and semantic equivalence recognition beyond simple character matching. Use when the user wants to benchmark on PDF Formula Extraction Benchmark, or asks about evaluating this task. Reports LLM-as-a-Judge.
Evaluates LLM safeguard robustness against national security and public safety (NSPS) risks by measuring both the model's tendency to generate harmful content in response to adversarial prompts and its tendency to incorrectly refuse legitimate benign requests. Use when the user wants to benchmark on FORTRESS, or asks about evaluating this task. Reports Average Risk Score (ARS).
Evaluates cross-dataset generalization and geographic domain shift in building change detection models. It measures how well models trained on one geographic region or dataset perform when tested on entirely different datasets, highlighting the impact of training data diversity on remote sensing model transferability. Use when the user wants to benchmark on FOTBCD-Binary, LEVIR-CD+, WHU-CD, or asks about evaluating this task. Reports IoU.
Evaluates a lightweight foundational model for ECG-based cardiac analysis, specifically testing its ability to classify signals as Normal/Abnormal and perform fine-grained disease classification across multiple cardiac conditions. Use when the user wants to benchmark on PTB-XL, CinC 2017, MedalCare-XL, PTB, or asks about evaluating this task. Reports F1-score.
Evaluates the ability of a machine learning closure model to stabilize reduced-order models for turbulent geophysical fluid dynamics. Specifically, it probes whether an extreme learning machine can predict mode-dependent eddy viscosities to maintain long-time integration accuracy and statistical steady-state behavior in coarse-grained ocean circulation simulations. Use when the user wants to benchmark on Four-gyre barotropic circulation problem, or asks about evaluating this task. Reports L2-...
Evaluates the capability of a Fourier-domain low-rank adapter to generate high-quality, diverse images for style transfer and concept editing. It also assesses the adapter's performance on standard language understanding benchmarks compared to baseline adapters like LoRA. Use when the user wants to benchmark on Paintings, Blue-Fire, 3D, Origami, GLUE, or asks about evaluating this task. Reports HPSv2.1.