
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the fidelity, utility, and privacy of synthetic tabular data generated by language models compared to real data and other generative baselines. It measures how well the synthetic distribution matches the original across statistical, downstream utility, and privacy dimensions. Use when the user wants to benchmark on Adult, Default, Shoppers, Magic, Beijing, or asks about evaluating this task. Reports C2ST.
This evaluation probes a model's ability to perform few-shot in-context learning and standard classification on high-dimensional, heterogeneous tabular data. It specifically measures how well biaxial attention and meta-learning improve performance across medical, financial, and energy domains, and how robust the model is to varying support set sizes and selection strategies. Use when the user wants to benchmark on TALENT, OpenML-CC18, or asks about evaluating this task. Reports accuracy (ACC).
Evaluates large language models on predictive tabular tasks, including classification, regression, and missing value imputation. It probes the model's ability to reason over structured data, handle mixed numerical and textual features, and perform few-shot or long-context learning on tables. Use when the user wants to benchmark on Kaggle (Classification & Regression), Tabular Benchmark (Grinsztajn et al., 2022), or asks about evaluating this task. Reports ROC-AUC.
Evaluates the calibration and reliability of confidence scores produced by LLMs when answering questions over tabular data. It probes how well predicted confidence aligns with actual accuracy across different elicitation methods and dataset complexities. Use when the user wants to benchmark on WikiTableQuestions, TableBench, or asks about evaluating this task. Reports smooth ECE.
Evaluates zero-shot and few-shot transfer learning capabilities of a language model on diverse tabular prediction tasks. It probes the model's ability to generalize across unseen datasets without fine-tuning, leveraging serialized row data and column headers to predict categorical or regression targets. Use when the user wants to benchmark on UniPredict Benchmark, Grinsztajn Benchmark, AutoML Multimodal Benchmark (AMLB), OpenML CC-18 Benchmark, OpenML CTR-23 Benchmark, or asks about evaluatin...
Evaluates the cross-dataset transferability of pretrained tabular generative models by measuring how well synthesized data preserves column distributions and pairwise correlations compared to ground truth tables. Use when the user wants to benchmark on Kaggle, GitTables, or asks about evaluating this task. Reports overall average.
Evaluates computational extrapolation and algorithmic generalization in tabular learning models by testing their ability to predict target values outside the training distribution. It probes whether models learn statistical interpolation versus deterministic computation on program-verified synthetic math problems. Use when the user wants to benchmark on TabularMath, or asks about evaluating this task. Reports rounded consistency.
This protocol re-evaluates tabular benchmarks to measure how validation strategy (holdout vs. 5-fold cross-validation) and hyperparameter optimization budgets affect model selection and reported performance. It probes the robustness of empirical conclusions in tabular machine learning when standard holdout validation is replaced with cross-validation ensembles. Use when the user wants to benchmark on TabZilla-hard, Grinsztajn et al. (2022) benchmark, or asks about evaluating this task. Report...
Evaluates code generation models on competition-level algorithmic programming problems. It probes fine-grained capabilities across different programming skills and difficulty levels by measuring whether generated Python programs correctly solve given problems under strict constraints. Use when the user wants to benchmark on TACO, or asks about evaluating this task. Reports pass@k.
Evaluates the perceptual naturalness and quality of text-to-speech synthesis. It probes the model's ability to generate high-fidelity audio waveforms that are indistinguishable from human speech. Use when the user wants to benchmark on Internal US English Test Set, Custom 100-Sentence Test Set, News Headlines Test Set, or asks about evaluating this task. Reports MOS.
Evaluates an agent's ability to perform logical deduction and multi-hop reasoning over long, noise-rich unstructured text without pre-defined schemas or tables. It probes the model's capacity to actively synthesize scattered evidence and filter out irrelevant distractors to arrive at a correct decision. Use when the user wants to benchmark on TACT, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates the effectiveness of various text embedding models combined with different anomaly detection algorithms for identifying text anomalies. It probes how well embedding-based anomaly detection generalizes across diverse domains (spam, fake news, hate speech) and distinguishes between patterned versus context-dependent anomalies. Use when the user wants to benchmark on Email-Spam, SMS-Spam, COVID-Fake, LIAR2, Hate-Speech, OLID, or asks about evaluating this task. Reports AUROC.
Evaluates computer vision models on detecting traffic accidents from surveillance footage across image classification, video classification, and object detection tasks. It probes the model's ability to distinguish accident scenarios from normal traffic and localize accident events in real-world highway scenes. Use when the user wants to benchmark on TAD, or asks about evaluating this task. Reports F1-score.
This benchmark evaluates an LLM's ability to generate and execute hybrid relational queries that combine traditional SQL operations with semantic reasoning over textual data. It probes capabilities like semantic joins, information extraction, and multi-hop reasoning by measuring execution accuracy against expert-verified ground truth. Use when the user wants to benchmark on TAG+, or asks about evaluating this task. Reports execution accuracy.
Evaluates the effectiveness of an adversarial agent in jailbreaking safety-aligned operator agents through conversational interaction. It measures how well a small attacker model can trigger prohibited tool usage on unseen malicious tasks using reinforcement learning. Use when the user wants to benchmark on TagAlong-Dojo, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates the ability of automatic speech recognition (ASR) systems to transcribe Taiwanese Hokkien (Taigi) audio into text. It specifically measures how well models map Taigi phonetic patterns to Mandarin character sequences using a standardized set of public service announcement recordings. Use when the user wants to benchmark on Taigi ASR Benchmark, or asks about evaluating this task. Reports Character Error Rate (CER).
Evaluates speech intent recognition in a low-resource, real-world setting for Taiwanese Taigi, specifically testing domain adaptation and robustness to domain mismatch between mined training data and real-world elderly speech. Use when the user wants to benchmark on TaigiSpeech, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to ground free-form natural language commands to specific objects in autonomous driving scenes. It probes spatial and relational language understanding, disambiguation of same-category objects, and handling of long-range referents and complex sentences. Use when the user wants to benchmark on Talk2Car, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness of state-of-the-art automatic speech recognition (ASR) models on real-world, unstructured conversational speech compared to controlled, read-speech benchmarks. It specifically probes how conversational disfluencies, interruptions, and variable audio durations impact transcription accuracy. Use when the user wants to benchmark on TalkBank, or asks about evaluating this task. Reports Word Error Rate.
Evaluates the robustness and generalization of deepfake detectors on talking-head videos under distribution shifts in identity and generator type. It probes whether detectors rely on spurious background cues or robust facial artifacts, and measures performance degradation when facing unseen generators and identities. Use when the user wants to benchmark on TalkingHeadBench, or asks about evaluating this task. Reports TPR@FPR=1% (T1).
Evaluates the performance isolation and resource sharing capabilities of GPU scheduling systems for concurrent deep learning workloads. It probes how well a system maintains tail latency for high-priority inference tasks while maximizing throughput for best-effort training tasks under varying traffic loads and workload combinations. Use when the user wants to benchmark on Tally Benchmark Suite, or asks about evaluating this task. Reports 99th-percentile latency.
Evaluates LLMs on automated unit test maintenance tasks, including test creation, repair, and updating across Python, Java, and Go. Probes context-aware reasoning, code generation, and the ability to dynamically adapt tests to production code changes. Use when the user wants to benchmark on TAM-Eval, or asks about evaluating this task. Reports mutation_score.
Probes a robot policy's ability to execute contact-rich bimanual manipulation tasks using visuo-tactile feedback. It evaluates robustness to visual disturbances, generalization to unseen object appearances, and the effectiveness of tactile pretraining and recovery data in imitation learning. Use when the user wants to benchmark on TAMEn Contact-Rich Manipulation Tasks, or asks about evaluating this task. Reports success rate (%).
Evaluates automatic speech recognition (ASR) systems on agglutinative languages (Tamil and Kannada) to measure how subword dictionary learning and segmentation techniques (Morfessor, BPE, extended-BPE) reduce out-of-vocabulary rates and improve word error rates compared to baseline word-level models. Use when the user wants to benchmark on Tamil and Kannada ASR dataset, or asks about evaluating this task. Reports WER.
This benchmark probes a model's ability to detect visual tampering on parcel logistics items by comparing a single RGB image to a reference database. It evaluates the pipeline's robustness in detecting corner keypoints, performing perspective transformation to generate viewpoint-invariant views, and accurately identifying appearance changes across varying angles, lighting, and lens distortions. Use when the user wants to benchmark on TAMPAR, or asks about evaluating this task. Reports F1-Score.
This benchmark evaluates the robustness of LLM safety safeguards against fine-tuning-based tampering attacks. It measures whether a model can maintain low accuracy on weaponized knowledge (forget) and high benign capabilities (retain) after undergoing various supervised fine-tuning attacks. Use when the user wants to benchmark on WMDP, MMLU, HarmBench, MT-Bench, or asks about evaluating this task. Reports Post-Attack Forget accuracy, Attack Success Rate (ASR).
Evaluates a model's ability to rank candidate answer sentences for a given question. It specifically probes the stability and robustness of pre-trained transformer models when transferred to a target domain using a two-stage fine-tuning protocol. Use when the user wants to benchmark on WikiQA, TREC-QA, or asks about evaluating this task. Reports MAP.
Evaluates whether automatically composing domain-specific data augmentation transformations improves end-task classification performance. It probes the ability of a learned sequence model to generate effective augmentation pipelines compared to heuristic or random baselines across image and text domains. Use when the user wants to benchmark on MNIST, CIFAR-10, ACE (Employment relation extraction), DDSM (Mammography), or asks about evaluating this task. Reports test set accuracy.
Evaluates novel view synthesis quality and 3D geometric reconstruction accuracy of NeRF variants on a single outdoor scene. Use when the user wants to benchmark on Tanks & Temples (Truck scene), or asks about evaluating this task. Reports CD.
Evaluates open-domain, multi-hop question answering by requiring models to aggregate data from multiple sources and generate structured answer tables. It probes capabilities in multi-hop reasoning, data normalization, unit conversion, and table construction. Use when the user wants to benchmark on TANQ, or asks about evaluating this task. Reports F1.
This benchmark evaluates the ability of machine learning models to predict click-through rates (CTR) for advertisements on a large-scale e-commerce platform. It specifically probes how well models capture static user-ad interactions versus dynamic, temporal user behavior sequences to forecast future clicks. Use when the user wants to benchmark on Alibaba's Taobao Advertising Dataset, or asks about evaluating this task. Reports AUC.
Evaluates few-shot and zero-shot Russian language understanding across six tasks probing logical reasoning, multi-hop inference, commonsense knowledge, and ethical judgment. It also measures model robustness against linguistic adversarial perturbations like typos, deletions, and modality changes. Use when the user wants to benchmark on TAPE, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to track arbitrary points in 3D space over time from monocular video, assessing spatio-temporal consistency, depth accuracy, and occlusion handling. It probes whether models can reconstruct scene geometry up to scale or only maintain relative depth consistency for local interactions. Use when the user wants to benchmark on TAPVid-3D, or asks about evaluating this task. Reports 3D Average Jaccard (3D-AJ).
Evaluates the performance, network traffic, and hardware overhead of the Tardis 2.0 cache coherence protocol under Total Store Order (TSO) consistency. It measures speedup and traffic reduction against a full-map MSI directory baseline and a baseline Tardis implementation across 20 multicore system benchmarks. Use when the user wants to benchmark on Splash2, PARSEC, HPCG, YCSB, TPCC, or asks about evaluating this task. Reports speedup.
Evaluates a generative world model's ability to produce controllable road images, predict geographic coordinates from street-view imagery, and generate self-consistent navigation actions on held-out spatiotemporal data. Use when the user wants to benchmark on STRIDE, or asks about evaluating this task. Reports georeferencing_error_m.
Evaluates the ability of conditional generative models to produce chemically valid and structurally similar molecular graphs conditioned on specific target protein sequences. It probes both the generative quality (validity, uniqueness, novelty) and the chemical fidelity (similarity to known drugs) of the generated molecules. Use when the user wants to benchmark on Combined Drug-Target Dataset (BIOSNAP, BindingDB, DAVIS, DrugBank), or asks about evaluating this task. Reports Tanimoto similarity.
Evaluates whether world models can perform mapless path planning toward semantic targets in real-world environments. It probes spatio-temporal consistency, trajectory accuracy, and the ability to reason about explicit versus implicit goals without prior map information. Use when the user wants to benchmark on Target-Bench, or asks about evaluating this task. Reports WO.
Evaluates an algorithm's ability to identify minimal control input sets for steering complex networks to target states. It probes scalability, solution optimality, and the capacity to prioritize biologically relevant nodes (e.g., drug targets) across synthetic and biological interaction networks. Use when the user wants to benchmark on Breast DEF, Breast HCC1428, Ovarian DEF, Pancreatic AsPC-1, Social Interaction 1, Erdos-Renyi 1000, Scale Free 1000, Small World 1000, or asks about evaluating...
This protocol evaluates a model's ability to maintain high accuracy on clean, unperturbed data while resisting adversarial attacks. It specifically probes the trade-off between clean accuracy and robustness under l_infinity perturbation constraints using both a custom simulated manifold dataset and the standard CIFAR-10 benchmark. Use when the user wants to benchmark on Transformed Hemisphere, CIFAR-10, or asks about evaluating this task. Reports Clean test accuracy.
This benchmark evaluates the parallel runtime performance and scalability of different programming systems by executing parameterized task graphs. It probes how efficiently systems handle varying degrees of parallelism, communication patterns, and computational versus memory-bound workloads. Use when the user wants to benchmark on Task Bench, or asks about evaluating this task. Reports minimum effective task granularity (METG).
Evaluates task graph scheduling algorithms by measuring their makespan on standard and adversarially modified network topologies. It probes how well algorithms minimize execution time across heterogeneous datasets and reveals performance reversals under minor structural changes. Use when the user wants to benchmark on Parallel Chains (and 15 other SAGA framework datasets), or asks about evaluating this task. Reports makespan.
Evaluates the visual perceptual capabilities of large multimodal language models across object recognition, attribute recognition, spatial and temporal reasoning, and action recognition using programmatically generated image and video question-answering tasks. Use when the user wants to benchmark on Task-Me-Anything, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of neural dialogue systems to track user intent and belief states, generate contextually appropriate responses, and successfully complete task-oriented conversations (specifically restaurant search) by jointly modeling intent, belief, and database interaction. Use when the user wants to benchmark on Wizard-of-Oz restaurant dialogue corpus, or asks about evaluating this task. Reports Objective task success rate.
Evaluates a model's ability to perform 3D semantic segmentation on multi-parametric MRI scans of brain tumours. It probes the algorithm's capacity to distinguish and delineate multiple tumour sub-regions (edema, enhancing, non-enhancing) under high clinical variability in scanner hardware and acquisition protocols. Use when the user wants to benchmark on Task01_BrainTumour, or asks about evaluating this task. Reports semantic segmentation accuracy.
Evaluates the quality of synthetically generated task automation instructions and tool invocation graphs. It probes whether the generated data is natural, appropriately complex, and correctly aligned with the underlying tool dependencies. Use when the user wants to benchmark on Hugging Face Tools, Multimedia Tools, Daily Life APIs, or asks about evaluating this task. Reports Alignment.
Evaluates discrete reasoning capabilities over hybrid tabular and textual data, specifically focusing on financial question answering tasks that require arithmetic operations, counting, and span extraction. Use when the user wants to benchmark on FinQA, TAT-QA, TAT-DQA, or asks about evaluating this task. Reports EM.
Evaluates cross-lingual sentence retrieval accuracy for low-resource languages by finding the most similar sentence in a target language corpus for each source sentence. It probes the alignment quality of vector spaces for languages with limited parallel data. Use when the user wants to benchmark on Tatoeba test setup (LASER), or asks about evaluating this task. Reports Accuracy.
Evaluates the task completion success rate of LLM agents performing long-horizon, tool-using agentic workflows across customer service and software engineering domains. Probes the agent's ability to execute environment-altering (mutating) actions safely and maintain goal alignment over extended trajectories. Use when the user wants to benchmark on $ au$-Bench Airline, $ au$-Bench Retail, $ au$-Bench-V Air, $ au$-Bench-V Ret, SWE-Bench Verified, or asks about evaluating this task. Reports score.
This benchmark evaluates full-duplex voice agents on real-world conversational tasks across retail, airline, and telecom domains. It jointly measures task completion success and real-time interaction quality, including responsiveness, latency, interruption handling, and selectivity under varying acoustic conditions like noise, diverse accents, and natural turn-taking. Use when the user wants to benchmark on $\tau^2$-bench, or asks about evaluating this task. Reports pass@1.
Evaluates conversational agents' ability to collaborate with a user simulator in a dual-control environment where both parties share tool access to a dynamic world. It probes coordination, communication under decentralized control, and adherence to domain-specific policies while resolving multi-step tasks. Use when the user wants to benchmark on $\tau^2$-Bench, or asks about evaluating this task. Reports pass^1.