
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates LLMs' ability to perform multi-step, constraint-aware reasoning and inference on real-world time series data. It probes compositional reasoning, numerical precision, and the capacity to assemble complex analytical or forecasting workflows via executable code generation. Use when the user wants to benchmark on TSAIA, or asks about evaluating this task. Reports Success Rate.
Evaluates large language models' ability to perform time series analysis and reasoning across six tasks (anomaly detection, classification, characterization, comparison, data transformation, and temporal relationship) using three question formats (true-or-false, multiple-choice, and puzzling). Use when the user wants to benchmark on TSAQA, or asks about evaluating this task. Reports accuracy.
Evaluates object detection models on traffic surveillance footage under diverse weather conditions and varying degrees of vehicle occlusion. It probes robustness to environmental degradation, scale variation, and dense urban traffic scenarios. Use when the user wants to benchmark on TSBOW, or asks about evaluating this task. Reports mAP50.
Compute the TschuprowsT metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TschuprowsT, or asks how to score with TschuprowsT.
Evaluates how time series foundation models scale in forecasting accuracy and uncertainty calibration as model size, compute, and training data size increase. It probes both in-distribution generalization and out-of-distribution transfer capabilities across multiple standard time series forecasting benchmarks. Use when the user wants to benchmark on Monash subset, LSF subset, or asks about evaluating this task. Reports NLL.
Evaluates large language models' ability to understand and generate structured scene graphs from textual narratives. It probes spatial reasoning, action decomposition, and the capacity to map dynamic descriptions to discrete visual or structural elements. Use when the user wants to benchmark on TSG Bench, or asks about evaluating this task. Reports Exact Match (EM) / Accuracy, Precision, Recall, Macro F1.
Evaluates the quality of discovered motif sets in time series by measuring alignment with ground truth segments. It accounts for variable-length patterns and time warping while penalizing both false discoveries and missed patterns. Use when the user wants to benchmark on TSMD benchmark datasets, or asks about evaluating this task. Reports F1-score.
Evaluates large language models on time series question answering across five tasks: forecasting, imputation, anomaly detection, classification, and open-ended reasoning. It probes the model's ability to integrate textual context with numerical time series data for both precise numerical prediction and natural language explanation. Use when the user wants to benchmark on TSQA, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform sequential recommendation with a focus on capturing repeat-aware temporal patterns. It measures how well the model balances predicting new items versus recurring items based on user interaction history and time intervals. Use when the user wants to benchmark on RetailRocket, LastFM, Diginetica, or asks about evaluating this task. Reports HR@K.
Evaluates models on Time Series Extrinsic Regression (TSER), where the goal is to predict a single continuous scalar value from multivariate time series inputs of varying lengths and dimensions. It probes the model's ability to handle irregular time series, missing values, and diverse domain-specific patterns without imputation. Use when the user wants to benchmark on Monash TSER Archive, or asks about evaluating this task. Reports R2.
This benchmark evaluates an AI system's ability to perform fact verification using time-series evidence. It probes multi-timeframe temporal reasoning, cross-series numerical analysis, and the generation of factually consistent justifications aligned with human annotations. Use when the user wants to benchmark on TSVer, or asks about evaluating this task. Reports Accuracy.
Evaluates lesion detection performance on synthetic and real breast mammography/tomosynthesis images. Probes the model's ability to localize lesions across varying breast densities, lesion sizes, and lesion densities using a free-response receiver operating characteristic (FROC) framework. Use when the user wants to benchmark on T-SYNTH, EMBED, or asks about evaluating this task. Reports FROC (Sensitivity vs. Average False Positives per Image).
Evaluates the ability of models to detect human body forgeries generated by diffusion models. It probes spatiotemporal motion inconsistencies and generalization across different generation configurations and unseen manipulation models. Use when the user wants to benchmark on TT-DF, or asks about evaluating this task. Reports AUC.
This protocol evaluates a 3D convolutional auto-encoder for removing reverberation artifacts (clutter) from transthoracic echocardiographic (TTE) sequences. It measures how well the network preserves cardiac structures while suppressing simulated artifacts, using synthetic data with known ground truth for training and validation, and normal in-vivo sequences for testing. Use when the user wants to benchmark on Synthetic TTE sequences, or asks about evaluating this task. Reports reconstruction...
Evaluates the impact of probabilistic versus deterministic duration modeling on the naturalness and intelligibility of non-autoregressive text-to-speech systems. It specifically probes how well stochastic duration predictors handle prosodic variability and disfluencies in spontaneous speech compared to read-aloud speech. Use when the user wants to benchmark on LJ, RS, TSGD2, AptS, or asks about evaluating this task. Reports CMOS.
Evaluates a text-to-speech model's ability to generate audio that matches natural language descriptions of speaker attributes (gender, accent, pitch, speaking rate, recording quality) and overall audio fidelity. It measures both objective acoustic metrics and subjective human ratings of relevance and naturalness. Use when the user wants to benchmark on MLS, LibriTTS-R, or asks about evaluating this task. Reports MOS.
Evaluates the ability of LLMs to improve reasoning performance at test time through self-reflection and targeted variant question synthesis, without external supervision. It probes how well a model can adapt its policy to difficult mathematical and general reasoning problems by diagnosing its own failures and generating corrective training signals. Use when the user wants to benchmark on AMC23, MATH-500, Minerva, OlympiadBench, AIME 2024, AIME 2025, GPQA-Diamond, MMLU-Pro, or asks about evalu...
This evaluation protocol probes the language understanding, reasoning, and instruction-following capabilities of Portuguese LLMs across diverse domains including academic exams, natural language inference, physical commonsense, and code generation. It is specifically designed to provide reliable training signals during pretraining and assess post-training alignment. Use when the user wants to benchmark on ARC Challenge, Calame, Global PIQA, HellaSwag, LAMBADA, ENEM, BLUEX, OAB, Belebele, MMLU...
Evaluates fine-grained temporal understanding on dense dynamic videos, probing camera motion, scene transitions, action sequences, and multi-subject interactions. It measures how well models capture dynamic visual elements and handle varying video complexities. Use when the user wants to benchmark on TUNA, or asks about evaluating this task. Reports F1 score, Accuracy.
Evaluates the robustness and zero-shot adaptation of vision-language segmentation models under prompt tuning across diverse medical and natural domain datasets. It probes how different prompt tuning strategies and prompt depths handle domain shifts and varying class counts. Use when the user wants to benchmark on Kvasir-SEG, ClinicDB, BKAI, ISIC 2016, DFU 2022, CAMUS, BUSI, CheXlocalize, Cityscapes, PascalVOC, or asks about evaluating this task. Reports Dice score.
Evaluates automatic speech recognition (ASR) models on Tunisian Arabic dialect audio, measuring overall transcription accuracy and code-switching performance for embedded English and French phrases. Use when the user wants to benchmark on LinTO, TunSwitch, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates the quality of Turkish legal language model embeddings for information retrieval tasks, specifically focusing on case law, regulations, and contract retrieval. It also assesses masked language modeling capabilities on diverse Turkish corpora to measure morphological and domain-specific token prediction accuracy. Use when the user wants to benchmark on MTEB-Turkish benchmark, Turkish Legal Retrieval Benchmarks, blackerx/turkish_v2, fthbrmnby/turkish_product_reviews, hazal/Turkish-Bio...
Evaluates a model's ability to capture statistical patterns and long-range dependencies in Turkish text. It tests generative probability estimation at both subword and character levels across news and Wikipedia domains. Use when the user wants to benchmark on trwiki-67, trnews-64, or asks about evaluating this task. Reports Perplexity (Ppl), Bits-per-character (Bpc).
Measures the quality of bidirectional translation between Turkish and English. It evaluates how well models capture cross-lingual semantic alignment and syntactic restructuring across different corpus types. Use when the user wants to benchmark on Wmt-16 (Turkish-English subset), MuST-C (Turkish-English subset), or asks about evaluating this task. Reports BLEU Score.
Tests the ability to identify and classify named entities (Person, Location, Organization) in Turkish text. It probes fine-grained token-level classification and boundary detection. Use when the user wants to benchmark on Milliyet-Ner, WikiANN (Turkish subset), or asks about evaluating this task. Reports CoNLL F-1.
Evaluates the detection of sentence boundaries in Turkish text across diverse domains (scientific abstracts, news, social media). It tests robustness to formatting variations and punctuation absence. Use when the user wants to benchmark on trseg-41, or asks about evaluating this task. Reports F1-score.
Evaluates a model's ability to detect lane boundaries in highway driving scenarios using a point-based accuracy metric over predefined row anchors. Use when the user wants to benchmark on TuSimple, or asks about evaluating this task. Reports accuracy.
Evaluates the end-to-end performance and optimization capability of a deep learning compiler across diverse hardware back-ends (GPU, CPU, embedded GPU, FPGA) on standard inference workloads. It measures how effectively the compiler automatically generates high-performance kernels compared to hand-tuned vendor libraries and existing frameworks. Use when the user wants to benchmark on DL Inference Workloads (ResNet-18, MobileNet, LSTM, DQN, DCGAN), or asks about evaluating this task. Reports sp...
Evaluates a model's ability to predict explicit and implicit user negative feedback for video recommendations, including binary judgment of controversial content and classification of dislike reasons. It also tests the model's capability to simulate user viewing behavior (e.g., fast-skip) based on historical interactions and profile data. Use when the user wants to benchmark on TVNF, MovieLens, Steam, or asks about evaluating this task. Reports Recall.
This evaluation probes a model's ability to correctly route natural language questions to the most appropriate domain-specific QA agent from a large, heterogeneous pool. It measures both sample efficiency (performance with few training examples per agent) and scalability (maintaining accuracy as the number of candidate agents grows to hundreds). Use when the user wants to benchmark on QA-Tasks, Many-Agents, or asks about evaluating this task. Reports Accuracy@1.
Evaluates named entity recognition (NER) and syntactic NLP capabilities on noisy, informal social media text. It probes a model's ability to handle domain-specific challenges like abbreviations, irregular capitalization, and complex entity structures common in tweets, while establishing baselines for tokenization, lemmatization, POS tagging, and dependency parsing. Use when the user wants to benchmark on Tweebank-NER (TB2), or asks about evaluating this task. Reports entity-level F1.
Compute the TweedieDevianceScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TweedieDevianceScore, or asks how to score with TweedieDevianceScore.
Evaluates the ability of language models to classify short social media posts across seven distinct Twitter-specific tasks, including sentiment, emotion, hate speech, and irony detection. It probes domain adaptation by comparing models pre-trained on generic text versus those further trained on large-scale Twitter corpora. Use when the user wants to benchmark on TweetEval, or asks about evaluating this task. Reports M-F1.
Evaluates language models on classifying tweets into predefined topics, testing both single-label and multi-label classification capabilities. It probes robustness to social media noise, short-form content, and topic overlap in real-world settings. Use when the user wants to benchmark on TweetTopic, or asks about evaluating this task. Reports Macro F1.
Evaluates a model's ability to perform dynamic visual reasoning by generating temporally grounded, physically plausible future frames and textual explanations. It probes both the quality of the step-by-step reasoning process and the correctness of the final answer in open-ended video scenarios. Use when the user wants to benchmark on TwiFF-Bench, Seed-Bench-R1, or asks about evaluating this task. Reports Answer score.
Evaluates the ability of LLMs to simulate individual-level human behavior across demographic, psychological, cognitive, economic, and behavioral economics domains. It measures test-retest accuracy and replication of known behavioral biases using a large-scale dataset of 2,058 U.S. individuals. Use when the user wants to benchmark on Twin-2K-500, or asks about evaluating this task. Reports test-retest accuracy.
Evaluates whether reward models exhibit political bias by measuring the average reward scores assigned to politically left-leaning versus right-leaning statements on the same topics. The protocol compares mean reward differences across model sizes and training runs to detect systematic left-leaning skew. Use when the user wants to benchmark on TwinViews-13k, or asks about evaluating this task. Reports average_reward.
Evaluates the annotation quality of the TWNERTC corpus for Turkish named entity recognition (NER) and text categorization (TC) by comparing automated labels against human-verified ground truths. It measures how well coarse-grained and fine-grained entity types, as well as domain categories, align with human judgment across different noise-reduction post-processing variants. Use when the user wants to benchmark on TWNERTC, or asks about evaluating this task. Reports F1-Score.
Evaluates LLMs' ability to solve university-level mathematical problems, both text-based and multimodal. It also includes a meta-evaluation component to assess how well models can judge the correctness of free-form mathematical solutions. Use when the user wants to benchmark on U-MATH, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy of dense optical flow estimation and the reliability of per-pixel uncertainty quantification in an unsupervised setting. It probes the model's ability to handle occlusions, textureless regions, and domain shifts without ground-truth flow supervision. Use when the user wants to benchmark on KITTI, Sintel, or asks about evaluating this task. Reports EPE.
Evaluates the accuracy of universal machine-learned interatomic potentials (uMLIPs) in predicting energy, forces, and stress tensors during high-temperature molecular dynamics simulations of metal-organic frameworks (MOFs), including their stability and thermal decomposition behavior. Use when the user wants to benchmark on High-Temperature MOF AIMD Benchmark, or asks about evaluating this task. Reports energy MAE.
Evaluates the capability to characterize and quantify correlated noise in near-term quantum processors by measuring the exponential decay of average fidelity under random quantum circuits. Use when the user has predictions and gold and needs to compute uXEB decay fitting (ENR extraction).
This evaluation probes an autonomous UAV navigation model's ability to track wildlife by predicting flight commands that match expert pilot behavior. It measures how well the model maintains optimal camera framing and altitude for behavioral video collection. Use when the user wants to benchmark on KABR, or asks about evaluating this task. Reports % of actions matching original flight.
Evaluates the predictive accuracy and computational efficiency of a user-behavior-based network traffic forecasting method against statistical and neural network baselines on real-world SMS traffic data. Use when the user wants to benchmark on Guangzhou and Milan SMS datasets, or asks about evaluating this task. Reports R2.
This evaluation protocol assesses the computational efficiency, physical realism, and policy-learning effectiveness of the UBSoft simulation platform for robotic tasks in unbounded soft environments. It benchmarks how well different algorithms (RL, trajectory optimization, heuristics) perform on eight manipulation and locomotion tasks, and evaluates the fidelity of open-loop sim-to-real transfer. Use when the user wants to benchmark on UBSoft Benchmark, or asks about evaluating this task. Rep...
Evaluates large language models' ability to handle dynamic, multi-turn financial dialogues across diverse user personas and task types, measuring their adaptability to shifting user needs and financial expertise. Use when the user wants to benchmark on UCFE, or asks about evaluating this task. Reports Elo score.
Evaluates the time-accuracy trade-offs of approximate Gaussian Process regression methods against exact baselines and simple models across multiple UCI regression datasets. It measures how quickly approximations converge to near-exact performance while tracking predictive quality over time. Use when the user wants to benchmark on UCI regression datasets, or asks about evaluating this task. Reports NLPD.
Evaluates whether a latent-structure diagnostic framework can distinguish between intrinsic self-preservation (terminal survival optimization) and instrumental self-preservation (survival as a means to a task) in autonomous agents. It measures the entanglement gap in hidden representations to classify agent types and tests robustness against adversarial mimics. Use when the user wants to benchmark on UCIP Gridworld, or asks about evaluating this task. Reports Accuracy.
Evaluates automatic speaker verification (ASV) robustness to speaking-style mismatches between enrollment and test utterances. It measures how well data augmentation techniques can compensate for style variability without requiring multi-style training data. Use when the user wants to benchmark on UCLA database, or asks about evaluating this task. Reports EER.
Evaluates time series classifiers' reliance on temporal structure by measuring accuracy degradation when temporal alignment is disrupted via padding. It contrasts performance on the original UCR benchmark against a perturbed version to isolate the contribution of temporal correlations versus tabular features. Use when the user wants to benchmark on UCR Augmented, or asks about evaluating this task. Reports accuracy.