
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates a dual-stream text-to-speech model's ability to generate high-quality, speaker-similar speech with low latency and high efficiency under streaming and offline conditions. It probes the model's robustness to complex text, alignment accuracy, and real-time generation speed compared to autoregressive and interleaved baselines. Use when the user wants to benchmark on LibriSpeech test-clean, SeedTTS test-zh, SeedTTS test-hard, or asks about evaluating this task. Reports RTF.
Evaluates the quality and alignment of a synthetic passage retrieval test collection (SynDL) by measuring how well system rankings on synthetic relevance judgments match those from official human-annotated TREC Deep Learning Track collections. It probes whether LLM-generated judgments can reliably substitute human assessors for deep relevance evaluation and system ranking. Use when the user wants to benchmark on SynDL, or asks about evaluating this task. Reports Kendall rank correlation coeff...
Evaluates the impact of curriculum learning and synthetic data generation on legal LLM fine-tuning. It measures performance across legal summarization, classification, and question-answering benchmarks to determine if synthetic data improves model capabilities over real-data-only baselines. Use when the user wants to benchmark on EurLex-Sum, EurLex, LexGLUE, BigLaw-Bench, CUAD, or asks about evaluating this task. Reports training loss.
Evaluates the logical reasoning and cross-domain generalization capabilities of models trained with verifiable synthetic data. It measures accuracy on mathematical, coding, and logical reasoning benchmarks using a multi-sample generation and verification protocol. Use when the user wants to benchmark on MATH 500, AIME 2024, AMC 2023, LiveCodeBench, SynLogic coding validation split, or asks about evaluating this task. Reports avg@8.
Evaluates a hybrid neural forecasting framework's ability to predict regional particulate matter (PM1, PM2.5, PM10) concentrations. It probes both average forecasting accuracy across spatial grids and the model's capacity to capture rare, high-impact pollution spikes and extreme events. Use when the user wants to benchmark on ERA5 & CAMS, or asks about evaluating this task. Reports Latitude-Weighted RMSE.
Evaluates unsupervised anomaly detection in medical imaging by training a generative model on healthy images and reconstructing anomalous inputs. It measures how well the model localizes and segments pathological regions by comparing the reconstruction residuals against ground-truth anomaly masks. Use when the user wants to benchmark on BraTS 2023 (Brain MRI), LiTS (Liver CT), Carotid US, or asks about evaluating this task. Reports Dice.
Evaluates vision-language models' ability to generate physically grounded, spatially accurate weather forecast discussions from Numerical Weather Prediction (NWP) images. It probes the model's capacity to identify and correctly locate synoptic-scale phenomena (e.g., pressure systems) in generated text, revealing limitations of traditional lexical metrics in domain-specific evaluation. Use when the user wants to benchmark on SynopticBench, or asks about evaluating this task. Reports Space-local.
Evaluates the effectiveness of an automated framework for synthesizing paralinguistic speech datasets on downstream paralinguistic text-to-speech generation and event detection tasks. It measures how well the generated data improves model performance in producing and recognizing paralinguistic features like laughter, sighs, and gasps compared to real-world annotated datasets. Use when the user wants to benchmark on SynParaSpeech, or asks about evaluating this task. Reports PMOS.
Probes whether music foundation models encode discrete and continuous Western music theory concepts by training linear or MLP classifiers on their internal audio embeddings. Use when the user wants to benchmark on SynTheory, or asks about evaluating this task. Reports accuracy.
Evaluates whether synthetic data generated by LLMs can effectively substitute real data for benchmarking NLP models, measuring both absolute performance alignment and relative ranking preservation across tasks. Additionally quantifies the self-bias of LLMs when they generate data and subsequently solve the same tasks. Use when the user wants to benchmark on Headlines, Tweet-News, CrossNER-Literature, CrossNER-Politics, SNIPS, ATIS, or asks about evaluating this task. Reports MSPD.
Evaluates the distributional fidelity of generated pharmacokinetic and drug-target interaction properties against real data, and measures the utility of the synthetic data for downstream regression tasks. Use when the user wants to benchmark on TDCommons/BindingDB PK & DTI Collection, or asks about evaluating this task. Reports Hellinger Distance (HD).
Evaluates a flow matching generative model's ability to produce realistic 3D subsurface geological models, both unconditionally and conditioned on sparse borehole data. It probes the model's capacity for geological interpolation, structural feature reconstruction, and probabilistic uncertainty estimation. Use when the user wants to benchmark on Synthetic Geology / StructuralGeo Dataset, or asks about evaluating this task. Reports probabilistic confidence intervals.
Evaluates the generalization capability of synthetic image detectors across different generative models, image resolutions, and real-world sources. It probes whether detectors rely on dataset-specific artifacts or scale-dependent biases rather than learning robust forgery signatures. Use when the user wants to benchmark on SuSy Benchmarking Datasets, or asks about evaluating this task. Reports recall.
Evaluates the quality and downstream utility of synthetic medical images generated by GANs by measuring how well classifiers trained on synthetic data perform compared to those trained on real data. It probes the trade-offs between image resolution, label complexity, and sample size on both visual fidelity and predictive performance. Use when the user wants to benchmark on Chest radiographs, Brain CT scans, or asks about evaluating this task. Reports AUC_real - AUC_syn.
Evaluates how synthetic or human-generated column descriptions impact LLM performance on text-to-SQL tasks, and assesses the quality of LLM-generated descriptions across varying semantic difficulty levels. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports Mean quality scores.
Evaluates the downstream utility and statistical fidelity of synthetic tabular data generated by various models. It probes whether synthetic data preserves classification accuracy, model selection rankings, feature importance rankings, and distributional similarity compared to real data. Use when the user wants to benchmark on Tabular Classification from Numerical features benchmark suite (filtered), or asks about evaluating this task. Reports AUROC.
Evaluates a 3D scene representation's capability for novel view synthesis, relighting, and inverse rendering (estimating diffuse albedo and roughness) from posed RGB images. Use when the user wants to benchmark on Synthetic4Relight, or asks about evaluating this task. Reports PSNR.
Probes the ability of an agent to select the optimal tabular data synthesizer for a given dataset and objective (privacy, fidelity, or utility) based on dataset stress profiles and a capability registry. It evaluates whether stress-aware, intent-conditioned matching outperforms heuristics, zero-shot LLMs, and meta-learning baselines in ranking generative models. Use when the user wants to benchmark on OpenML Tabular Benchmark (Abalone, Bean, IndianLiverPatient, Obesity, faults, insurance, wil...
Evaluates the ability of LLMs to infer private personal attributes (e.g., occupation, age, location, income) from concatenated user comments. It also assesses the fidelity of synthetic comments compared to real human text via human studies. Use when the user wants to benchmark on SynthPAI, or asks about evaluating this task. Reports 0-1 accuracy.
Evaluates the effectiveness of a synthetic referring expression dataset for training language-guided video object segmentation models. It measures segmentation accuracy when models are trained on synthetic versus human annotations and evaluated on standard referring video segmentation benchmarks. Use when the user wants to benchmark on DAVIS-2017, Refer-YouTube-VOS, or asks about evaluating this task. Reports J&F.
This benchmark evaluates 2D and 3D point tracking capabilities across diverse synthetic domains, including rapid camera motion, articulated objects, and occlusions. It probes a model's ability to maintain spatio-temporal correspondence, handle depth-adaptive spatial errors, and correctly classify occlusion or out-of-frame status under significant distribution shifts. Use when the user wants to benchmark on SynthVerse, or asks about evaluating this task. Reports AJ_3D.
Evaluates the expected cost incurred by a crowdsourcing platform when inferring service states from biased, strategic user reviews under different baseline mechanisms. Use when the user has predictions and gold and needs to compute system loss.
This evaluation probes the trade-off between total system throughput and user fairness in UAV-enabled wireless networks. It measures how effectively a resource allocation and trajectory design scheme balances maximizing aggregate data rates against ensuring equitable service across users with varying channel conditions. Use when the user has predictions and gold and needs to compute system throughput.
Evaluates the ability of representation learning models to factorize audio into independent semantic factors (timbre, amplitude, frequency). It measures how well the learned latent space aligns with these ground-truth factors using standard disentanglement metrics. Use when the user wants to benchmark on SynTone, or asks about evaluating this task. Reports MIG.
Evaluates the quality and effectiveness of a synthetic multilingual voice command dataset for on-device keyword spotting. It probes whether TTS-synthesized audio can support high-accuracy classification across different model complexities and languages (English and Chinese). Use when the user wants to benchmark on SYNTTS-COMMANDS, or asks about evaluating this task. Reports classification accuracy.
Evaluates Retrieval-Augmented Generation (RAG) systems on their ability to retrieve relevant text-and-table contexts from financial reports and perform numerical reasoning to answer questions. It measures both retrieval effectiveness and the accuracy of the generated numerical answers. Use when the user wants to benchmark on T2-RAGBench, or asks about evaluating this task. Reports Number Match (NM), MRR@3.
Evaluates the fine-grained quality of text-to-3D generated meshes across multiple dimensions including textual alignment, visual quality, and authenticity. It measures how well generative models adhere to complex compositional prompts and produce structurally sound, aesthetically pleasing 3D assets. Use when the user wants to benchmark on T23D-CompBench, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
Evaluates text-to-image models' ability to handle high compositional density and multi-step visual reasoning. It probes instance, attribute, and relation binding, text rendering, and deductive/inductive/abductive inference capabilities. Use when the user wants to benchmark on T2I-CoReBench, or asks about evaluating this task. Reports Overall Score.
Evaluates the ability to deanonymize text-to-image models by identifying which model generated a given image, exploiting model-specific visual signatures in embedding space. Use when the user wants to benchmark on T2I Leaderboard Prompts, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates text-to-image generation models on their ability to align with complex, compositional prompts through iterative fine-grained reasoning and self-refinement. It probes capabilities in object counting, attribute binding, spatial relationships, and handling long, dense prompts. Use when the user wants to benchmark on GenEval, T2I-CompBench, DPGBench, or asks about evaluating this task. Reports GenEval, T2I-CompBench, and DPGBench alignment scores.
Evaluates the safety and alignment of text-to-image (T2I) models by measuring their susceptibility to generating harmful content across a hierarchical taxonomy of risks. It probes whether models can be prompted to produce NSFW, copyright-infringing, or politically sensitive images, and tests the effectiveness of various defense mechanisms and safety filters. Use when the user wants to benchmark on T2I-RiskyPrompt, or asks about evaluating this task. Reports risk ratio.
Evaluates the visual quality and text-3D alignment of generated 3D scenes across varying prompt complexities (single object, object with surroundings, multiple objects). It specifically probes multi-view consistency (detecting the Janus problem) and the ability of 2D diffusion guidance to translate into coherent 3D structures. Use when the user wants to benchmark on T$^3$ Bench, or asks about evaluating this task. Reports Multi-view Quality (ImageReward), Alignment (GPT-4).
Evaluates a unified text-to-text transformer's ability to generalize across diverse NLP tasks including language understanding, summarization, question answering, and machine translation. The protocol tests the efficacy of pre-training objectives, data scaling, and consistent text-to-text fine-tuning pipelines. Use when the user wants to benchmark on GLUE, SuperGLUE, CNN/Daily Mail, SQuAD, WMT, or asks about evaluating this task. Reports GLUE average score.
Evaluates NLP models' ability to identify and mask personally identifiable information (direct and quasi-identifiers) in legal texts while preserving non-sensitive content. It probes both privacy protection (full coverage of masking spans) and information utility (minimizing unnecessary masking of non-identifying entities). Use when the user wants to benchmark on TAB corpus, or asks about evaluating this task. Reports ER_{di}.
Evaluates the predictive performance of tabular machine learning models across 51 real-world datasets under standardized, reproducible protocols. It probes how hyperparameter tuning, nested cross-validation, and post-hoc ensembling affect peak performance and efficiency trade-offs. Use when the user wants to benchmark on TabArena, or asks about evaluating this task. Reports predictive performance.
Tests a model's ability to verify the truthfulness of a factual statement given a table. It probes structured reasoning capabilities by requiring the model to cross-reference table contents with a claim and output a binary label. Use when the user wants to benchmark on TabFact, or asks about evaluating this task. Reports binary classification accuracy.
Evaluates the robustness of tabular machine learning models when the set of available features dynamically changes in open environments. It measures performance degradation across classification and regression tasks under varying degrees of feature removal (20% to 100%). Use when the user wants to benchmark on TabFSBench datasets, or asks about evaluating this task. Reports performance gap (Δ).
Evaluates the ability of retrieval-augmented large language models to perform in-context learning on tabular data for classification and regression tasks. It probes how well non-parametric retrieval of support instances scales with dataset size and compares against numeric-based and classic tabular baselines. Use when the user wants to benchmark on Held-out Tabular Benchmark, or asks about evaluating this task. Reports AUROC, NMAE.
Evaluates a model's ability to answer questions about tabular data using various reasoning strategies. It probes factual retrieval, numerical reasoning, and complex multi-step table understanding across different difficulty levels. Use when the user wants to benchmark on Penguins in a Table, TableBench, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates graph-based machine learning models for sequence labeling (BIESO) and table row detection on handwritten historical register books. The protocol tests the models' ability to segment table rows and label cell boundaries using pre-extracted textline and column features rather than raw images. Use when the user wants to benchmark on Dataset1, Dataset2, or asks about evaluating this task. Reports F1 score.
Evaluates table structure recognition (TSR) and cell detection capabilities by comparing two tokenization schemes (OTSL vs. HTML) on transformer-based image-to-sequence models. It measures how well the model predicts table layouts and cell bounding boxes across diverse document types. Use when the user wants to benchmark on PubTabNet, FinTabNet, PubTables-1M, or asks about evaluating this task. Reports Tree Edit Distance score (TEDs).
Probes multimodal large language models' ability to perform spatially grounded reasoning over complex hierarchical tables. It specifically evaluates performance degradation across three cognitive levels (Perception, Reasoning, Analysis) and measures how explicit spatial anchoring mitigates perceptual overload and spatial attention failure. Use when the user wants to benchmark on TableVision, or asks about evaluating this task. Reports exact-match Accuracy (%).
Evaluates deep learning models on table structure recognition (TSR) and table content recognition (TCR) by predicting LaTeX token sequences from tabular images. It probes the model's ability to accurately reconstruct table layouts and textual content under varying aspect ratios and sequence lengths. Use when the user wants to benchmark on TabLeX, or asks about evaluating this task. Reports EMA.
This protocol evaluates the faithfulness of feature attributions for LLM-based tabular classifiers. It measures how well an attribution method ranks features by sequentially masking them in importance order and tracking the resulting drop in the model's predicted class probability. Use when the user wants to benchmark on Adult Income, Heart Disease, or asks about evaluating this task. Reports faithfulness.
Evaluates the performance of traditional machine learning, deep learning, attention-based, and contrastive learning methods on tabular classification tasks. It probes how data characteristics (dimensionality, difficulty) influence the optimal learning strategy and compares different masking/filling strategies used in contrastive learning. Use when the user wants to benchmark on OpenML Tabular Benchmark, or asks about evaluating this task. Reports F1 score.
Evaluates the predictive performance and stability of 32 deep learning and tree-based tabular models across a large collection of diverse tabular datasets. It probes how well different architectures handle classification and regression tasks, and how dataset characteristics influence method rankings. Use when the user wants to benchmark on LAMDA-TALENT Benchmark, or asks about evaluating this task. Reports average_rank.
Evaluates automated tabular data cleaning pipelines by measuring downstream classification accuracy and calibration when processed by a Tabular Foundation Model (TabPFN v2). It probes whether cleaning strategies can effectively align dirty data distributions with the model's learned prior to improve predictive performance. Use when the user wants to benchmark on OpenML CC18 Benchmark Suite (D1–D10), or asks about evaluating this task. Reports accuracy.
This evaluation probes the robustness and relative performance of tabular machine learning models when subjected to expert-level, dataset-specific preprocessing pipelines rather than standardized baselines. It specifically measures how feature engineering, hyperparameter optimization, and test-time adaptation shift model rankings and close performance gaps across real-world competition datasets. Use when the user wants to benchmark on Kaggle competition datasets (MBGM, BPCCM, HQC, SCTP, PSSDP...
Evaluates the effectiveness of deep image embedding clustering methods compared to traditional clustering algorithms on heterogeneous tabular datasets. It probes whether architectures designed for spatial image data can effectively learn representations for low-dimensional, non-spatial tabular data. Use when the user wants to benchmark on malware, mice, vehicle, olive, dermatology, breast cancer, Ecoli, or asks about evaluating this task. Reports clustering accuracy.
Evaluates feature selection methods by measuring downstream neural network performance on tabular datasets containing controlled extraneous features. It probes whether selected features improve or maintain predictive accuracy for classification and reduce error for regression tasks. Use when the user wants to benchmark on ALOI (AL), California Housing (CA), Covertype (CO), Eye Movements (EY), Gesture (GE), Helena (HE), Higgs 98k (HI), House 16K (HO), Jannis (JA), Otto Group Product Classifica...