
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates real-time probabilistic forecasting of financial and weather time series, probing a model's ability to quantify uncertainty via quantile modeling and maintain calibration over sequential submission rounds. Use when the user wants to benchmark on DAX, Wind, Temperature, or asks about evaluating this task. Reports skill score.
Evaluates multimodal foundation models on open-ended, expert-level queries across 10 professional domains, probing visual perception, domain knowledge, and long-context reasoning in single-round, multi-lingual, and multi-turn settings. Use when the user wants to benchmark on ProBench, or asks about evaluating this task. Reports ELO rating.
This evaluation probes the training efficiency and multi-GPU scaling behavior of deep learning frameworks. It measures how quickly models process mini-batches and how effectively data parallelization affects model convergence across various network architectures and hardware configurations. Use when the user has predictions and gold and needs to compute processing_time.
Evaluates reinforcement learning agents on their ability to learn efficiently and generalize to unseen, procedurally generated environments. It measures how well algorithms adapt to novel level distributions under strict computational and timestep constraints. Use when the user wants to benchmark on Procgen Benchmark, or asks about evaluating this task. Reports mean normalized return.
This protocol evaluates the causal impact of specific profile image features (smile, body-shot, and gender) on lender selection preferences in a simulated micro-lending marketplace. It uses a conjoint-style choice experiment with GAN-generated images to isolate how visual cues influence funding decisions independent of borrower creditworthiness. Use when the user wants to benchmark on Custom GAN-generated profile images, or asks about evaluating this task. Reports Average Treatment Effect (ATE).
Evaluates vision models on prosthesis-specific video understanding, including instance segmentation of amputees and prosthetic limbs, 2D human pose estimation with focus on lower-body keypoints, and automated gait pattern classification from pose sequences. Use when the user wants to benchmark on ProGait, or asks about evaluating this task. Reports mIoU, AP@[0.5,0.95].
Evaluates a progressive feature transmission protocol for split inference at the wireless edge, measuring how efficiently features are transmitted to meet target inference accuracy or uncertainty thresholds under varying channel conditions. Use when the user wants to benchmark on GM dataset, MNIST, or asks about evaluating this task. Reports average communication latency.
Evaluates how well discovered motif sets in time series approximate ground truth motif sets. It penalizes false positives, false negatives, and redundant motifs without requiring uniform motif lengths or a fixed number of motif sets. Use when the user has predictions and gold and needs to compute PROM.
This evaluation protocol assesses the capability of language models to act as automated judges for other language models. It probes two distinct paradigms: direct assessment, where a model scores a single response against a reference or rubric, and pairwise ranking, where a model selects the preferred response between two candidates. The benchmarks cover instruction-following, alignment, and fine-grained custom criteria. Use when the user wants to benchmark on Vicuna Bench, MT Bench, FLASK, F...
Evaluates the fine-grained judgment capability of vision-language models by scoring generated text outputs against instance-specific rubrics and reference answers. It measures alignment with human preferences and state-of-the-art VLM judges across instruction following, VQA, and captioning tasks. Use when the user wants to benchmark on LLaVA-Bench, VisIT-Bench, Perception-Bench, OKVQA, VQAv2, TextVQA, COCO-Captions, NoCaps, or asks about evaluating this task. Reports Pearson correlation.
Evaluates 3D volumetric medical image segmentation capability on prostate MRI scans. It probes the model's ability to accurately delineate organ boundaries under clinical variability and class imbalance using end-to-end fully convolutional networks. Use when the user wants to benchmark on PROMISE2012, or asks about evaluating this task. Reports Dice coefficient.
This benchmark evaluates an LLM's or monitoring system's ability to distinguish between safe user inputs and malicious prompt injection attacks. It measures both false positive rates on legitimate interactions and false negative rates on adversarial prompts to assess overall security robustness. Use when the user wants to benchmark on Gandalf, Tensor-Trust, SPML-Dataset, or asks about evaluating this task. Reports Error Rate (ER).
This benchmark evaluates the vulnerability of large language models to prompt injection attacks when generating scientific paper reviews. It probes whether hidden or biased instructions embedded in parsed PDFs can systematically skew the model's review scores and recommendations. Use when the user wants to benchmark on ICLR 2024 Review Dataset, or asks about evaluating this task. Reports Rating.
Evaluates Chinese large language models on five biomedical NLP tasks: information extraction, text classification, natural language inference, dialogue understanding, and content generation. It probes instruction-following, few-shot in-context learning, and parameter-efficient fine-tuning capabilities in a medical domain context. Use when the user wants to benchmark on PromptCBLUE, or asks about evaluating this task. Reports Instance-level strict micro-F1.
Evaluates dense retrieval models on instruction-following and standard out-of-domain tasks, specifically probing their ability to leverage natural language prompts for zero-shot hyperparameter tuning and robustness to query phrasing. Use when the user wants to benchmark on FollowIR, InstructIR, MS MARCO, BEIR, or asks about evaluating this task. Reports nDCG@10.
This evaluation probes a model's ability to detect prompt injection attacks in realistic deployment settings. It specifically tests whether a detector can distinguish between benign conversational or application-structured inputs and maliciously crafted injections while maintaining a very low false positive rate to avoid costly false alarms. Use when the user wants to benchmark on PromptShield Evaluation Set, or asks about evaluating this task. Reports TPR@0.1%FPR.
Evaluates a model's ability to translate between natural language mathematics and Lean 3 formal statements (autoformalization and informalization). It measures syntactic validity, semantic correctness, and lexical similarity to assess reasoning over undergraduate-level theory. Use when the user wants to benchmark on ProofNet, or asks about evaluating this task. Reports Accuracy.
Detects propaganda techniques in news articles through a two-stage pipeline: identifying text spans containing propaganda (Span Identification) and classifying the specific rhetorical technique used within those spans (Technique Classification). Use when the user wants to benchmark on SemEval-2020 Task 11 Propaganda Detection, or asks about evaluating this task. Reports F1.
Evaluates a two-stage malware detection framework that uses a fast ML classifier for initial triage and a deep learning model for borderline cases. It probes the model's ability to accurately classify system call sequences as malicious or benign while balancing detection latency and false positive rates in real-time scenarios. Use when the user wants to benchmark on Propedeutica System Call Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to decompose sentences into atomic semantic units (propositional segmentation) and determine entailment relationships between text spans. It probes fine-grained compositional semantic alignment and partial entailment recognition beyond sentence-level NLI. Use when the user wants to benchmark on PropSegmEnt, or asks about evaluating this task. Reports Precision/Recall/F1w (macro-averaged).
Evaluates conversational agents on dialogue safety classification, rule-of-thumb generation, and prosocial response generation. It probes the model's ability to identify unsafe content, generate socially informed guidelines, and produce safe, engaging, and respectful dialogue responses. Use when the user wants to benchmark on PROSOCIALDIALOG, or asks about evaluating this task. Reports accuracy, BLEU-4.
Evaluates protein language models and geometric deep learning architectures on five realistic downstream biological tasks, including binding affinity prediction, functional annotation, mutation effects, cleavage site detection, and PROTAC interaction modeling. It probes how pretraining objectives, structural information integration, and domain-specific inductive biases affect generalization on limited biological data. Use when the user wants to benchmark on Protap Benchmark, or asks about eva...
Evaluates protein language models on classification and regression tasks to assess their capability in predicting protein properties, sub-cellular localization, epitope regions, and mutational effects. Use when the user wants to benchmark on Sub-cellular Localization, Membrane Solubility, Epitope Region Prediction, GB1 Mutational Landscape, or asks about evaluating this task. Reports accuracy, Spearman's rank correlation coefficient.
This evaluation framework probes the physical validity, functional success, and generalization capability of generative protein models. It emphasizes leakage-aware dataset splits, structural and docking quality metrics, and experimental validation pipelines to ensure designs are biologically plausible and functionally active. Use when the user wants to benchmark on PLINDER, PoseBusters, PDFBench, PoseX, VenusX, FragBench, GeomMotif, or asks about evaluating this task. Reports RMSD.
Evaluates the ability of graph neural networks combined with language models to learn structural and sequence representations of proteins. It probes how well the learned embeddings preserve structural similarity via TM-score prediction and generalize to downstream classification tasks across different protein families and out-of-distribution datasets. Use when the user wants to benchmark on Kinase dataset, SCOPe dataset, or asks about evaluating this task. Reports MSE.
Evaluates a model's ability to predict protein-ligand binding using only sequence data. It probes generalization across diverse protein targets, novel chemical scaffolds, and external benchmarks by measuring ranking performance between binders and decoys. Use when the user wants to benchmark on DEL Protein Split, DEL Chemical Library Split, MF-PCBA, Public Binders/Decoys, or asks about evaluating this task. Reports AUROC.
Evaluates a model's ability to predict the functional or stability impact of amino acid substitutions in proteins without prior experimental data for the specific variant. It probes zero-shot generalization across diverse protein families, taxonomic groups, and mutational depths (single-site vs. deep mutations). Use when the user wants to benchmark on DTm, DDG, ProteinGym, or asks about evaluating this task. Reports TPR@threshold.
Evaluates protein language models on sequence understanding across multiple downstream tasks (structure, function, interactions, developability) and 3D structure prediction from single amino acid sequences. Use when the user wants to benchmark on CAMEO, CASP15, OOD Protein Sequences (UniProt), or asks about evaluating this task. Reports TM-score.
This benchmark evaluates the quality of prototype-based explainable AI (XAI) methods for time series and image data. It probes how well generated prototypes capture model behavior and data structure across nine interpretability properties, including correctness, consistency, continuity, and latent space cohesion. Use when the user wants to benchmark on ECG200, or asks about evaluating this task. Reports Total.
Evaluates multimodal language models' vision-centric instruction following, multi-image reasoning, and general visual understanding. It probes the model's ability to process single and multiple images, answer factual questions, and perform complex reasoning across a suite of standard benchmarks. Use when the user wants to benchmark on CV-Bench (CVB-2D, CVB-3D), SEED-Bench, MMBench (MMB), MME, QBench2, MMMU, RealWorldQA, MMStar, MMVet, Mantis-Eval, MMT-Bench (MMT), TextVQA, or asks about evalu...
Evaluates an approximate caching system for Retrieval-Augmented Generation (RAG) pipelines. It measures how well the cache preserves retrieval quality and end-to-end accuracy while reducing database lookup latency under uniform and skewed query workloads. Use when the user wants to benchmark on MMLU (econometrics subset), MedRAG (PubMedQA subset), MedRAG-Zipf, or asks about evaluating this task. Reports test accuracy.
Evaluates a model's ability to learn sequentially from multiple tasks without catastrophic forgetting, measuring both task accuracy and continual learning dynamics like forward/backward transfer and forgetting rates across NLP and vision benchmarks. Use when the user wants to benchmark on Standard & Long, TRACE, ViT Benchmark, or asks about evaluating this task. Reports Accuracy (Acc/AAA).
Evaluates the ability of Estimation of Model Accuracy (EMA) methods to predict the structural quality of protein complex models. It probes global and interface-level accuracy estimation using correlation, ranking, and classification metrics against reference structural scores. Use when the user wants to benchmark on CASP16_inhouse_TOP5_dataset, CASP16_community_dataset, or asks about evaluating this task. Reports Pearson’s correlation (CorrP).
Evaluates sound event localization and detection (SELD) performance on synthetic and real-world audio, measuring classification accuracy, localization precision, and overall detection quality across different network architectures and fine-tuning strategies. Use when the user wants to benchmark on synthetic-test-set, synthetic-training-set, Indoor Recordings, or asks about evaluating this task. Reports SELD.
Evaluates the ability of masked language models to predict masked amino acids in short peptide sequences, measuring how well the model captures local sequence dependencies without autoregressive assumptions. It specifically probes the model's capacity to generalize to unseen short peptides that were excluded from the reference database during training. Use when the user wants to benchmark on UniRef100-excluded peptides, or asks about evaluating this task. Reports PseudoPPL.
This evaluation probes the closed-loop planning robustness and causal reasoning of autonomous vehicle controllers by measuring their ability to handle compounding errors and distribution shifts. It combines real-world driving observations with pseudo-synthetic future scenarios generated via neural rendering to approximate interactive simulation without requiring a full physics engine. Use when the user wants to benchmark on nuPlan (navhard subset), or asks about evaluating this task. Reports ...
Evaluates panoptic scene graph generation models on their ability to predict object triplets with correct masks and relations. It probes recall-based performance at different top-k limits, mean recall across predicates, pair recall, and predicate ranking accuracy. Use when the user wants to benchmark on PSG, or asks about evaluating this task. Reports Mean Recall@k.
Evaluates the ability of models to detect span-level hallucinations in multilingual question-answering contexts. It probes cross-lingual generalization and token-level inconsistency detection between generated answers and ground truth. Use when the user wants to benchmark on PsiloQA, or asks about evaluating this task. Reports IoU.
Evaluates the phonological accent fidelity and prosodic naturalness of Indic text-to-speech systems across Hindi, Telugu, and Tamil. It decomposes accent into per-phoneme dimensions (retroflex, aspiration, Tamil-zha, vowel-length) and corpus-level distributional metrics, revealing gaps between intelligibility and native-like accent. Use when the user wants to benchmark on PSP Benchmark Sets, or asks about evaluating this task. Reports FAD.
Evaluates protein structure prediction models on their ability to infer accurate 3D atomic coordinates from amino acid sequences. It specifically probes topological backbone similarity and side-chain accuracy when trained on large-scale distilled protein datasets. Use when the user wants to benchmark on CASP14 test set, PSP dataset, or asks about evaluating this task. Reports TM-score.
This benchmark evaluates the robustness and accuracy of automatic speech recognition (ASR) systems for the Persian language across diverse acoustic conditions, demographic groups, and linguistic domains. It specifically probes how well models handle regional accents, spontaneous or informal speech, and underrepresented demographics, while highlighting architectural and data-related performance gaps. Use when the user wants to benchmark on PSRB, or asks about evaluating this task. Reports SW-WER.
Evaluates the objective quantification and benchmarking of psychomotor execution quality in sports using wearable IMU data. It maps raw 3D motion trajectories into a normalized performance space and uses unsupervised clustering to identify optimal movement patterns and detect technical deviations. Use when the user wants to benchmark on Table Tennis Forehand Stroke (IMU), or asks about evaluating this task. Reports Euclidean distance to ideal performance origin.
Evaluates an AI model's ability to generate empathetic, principle-constrained psychological counseling dialogues in simulated multi-turn interactions. It probes competencies like accurate empathy, logical consistency, resistance handling, and ethical guidance beyond surface-level language features. Use when the user wants to benchmark on PsyEval, or asks about evaluating this task. Reports PsyEval.
Evaluates the quality of a newly constructed Portuguese-English parallel corpus of scientific abstracts. It probes the effectiveness of automated sentence alignment algorithms and the translation performance of SMT and NMT models on domain-specific academic text. Use when the user wants to benchmark on Theses and Dissertations Abstracts Corpus, or asks about evaluating this task. Reports BLEU.
Evaluates next-token prediction accuracy and long-range dependency modeling in language models, with a specific focus on handling rare and out-of-vocabulary words without expanding vocabulary size. Use when the user wants to benchmark on Penn Treebank, WikiText-2, or asks about evaluating this task. Reports perplexity.
Evaluates deep learning models on 12-lead ECG time series for multi-label classification of diagnostic, rhythm, and form statements. It probes the ability of architectures to learn directly from raw signals versus traditional feature extraction, and assesses transfer learning and demographic attribute prediction capabilities. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports term-centric macro-averaged AUC.
This benchmark evaluates how ECG sampling frequency impacts deep learning models for binary atrial fibrillation detection. It probes a model's discrimination capability and the reliability of its predicted probabilities across different temporal resolutions, while enforcing strict patient-level separation to prevent data leakage. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports AUROC.
This evaluation probes a model's ability to detect cardiovascular anomalies in 12-lead electrocardiogram (ECG) signals and localize the specific temporal regions where abnormalities occur. It tests the model's capacity to capture both global and local temporal dependencies in raw, unsegmented time-series data without relying on traditional R-peak detection or heartbeat segmentation. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports AUC.
Evaluates the ability of self-supervised pre-trained Vision Transformers to classify ECG signals across multiple diagnostic label hierarchies (e.g., all statements, ST-MEM labels, diagnostic subclasses, rhythm statements). Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports macro AUC.
Evaluates the ability of foundation models to perform multi-label classification on clinical 12-lead ECG recordings. It probes robustness to class imbalance and sample efficiency by measuring performance across varying label sets and training data sizes. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports macro AUROC.