
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates large language models' ability to extract structured information from multi-modal sources (text, images, audio) into valid JSON formats, isolating schema compliance from value accuracy. Use when the user wants to benchmark on Multi-Source Structured Output Benchmark, or asks about evaluating this task. Reports correct_value_extraction.
Evaluates how different prompting strategies (baseline, zero-shot, chain-of-thought, and automated optimizers) affect the accuracy, ranking stability, and variance of language model performance across multiple knowledge and reasoning benchmarks. Use when the user wants to benchmark on MMLU-Pro, GSM8K, MedCalc-Bench, GPQA, HeadQA, MedBullets, Medec, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to predict architectural elements (walls, doors, windows) and room layouts within indoor 3D scenes. It tests the model's capacity for structured scene understanding and spatial reasoning by comparing predicted layouts against ground-truth annotations. Use when the user wants to benchmark on Structured3D, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to localize a specific object or event in a video based on a natural language query. It measures both temporal localization (identifying the correct start and end timestamps) and spatial localization (predicting accurate bounding box trajectories across the video frames). Use when the user wants to benchmark on VidSTG, HCSTVG-v1&v2, or asks about evaluating this task. Reports m_vIoU.
Evaluates unsupervised style transfer models on their ability to transform text between two styles (e.g., Shakespeare vs. modern English, formal vs. informal) while preserving semantic meaning and maintaining linguistic quality. Use when the user wants to benchmark on Shakespeare author imitation dataset (Xu et al., 2012), Formality transfer dataset (Rao and Tetrault, 2018), or asks about evaluating this task. Reports J(A,S,F).
Evaluates a model's ability to perform sequential product recommendation in an e-commerce setting by predicting the next item in a user session. It specifically probes how well the model leverages historical interaction sequences, visual style embeddings, and shopping cart data to rank relevant products. Use when the user wants to benchmark on Style4Rec E-commerce Dataset, or asks about evaluating this task. Reports HR@5.
Evaluates speech language models' ability to control four paralinguistic dimensions (emotion, speed, volume, pitch) across multi-turn dialogues. It probes whether models can follow gradational style-intensity instructions while preserving fixed semantic content and maintaining coherent intensity trajectories across turns. Use when the user wants to benchmark on StyleBench, or asks about evaluating this task. Reports dimension-specific metrics.
Evaluates a visual-to-audio dubbing model's ability to generate emotionally consistent, speaker-identical speech that aligns temporally with video lip movements. It probes multi-scale style learning across unseen speakers, reference audio variations, and phoneme-level lip-sync accuracy. Use when the user wants to benchmark on V2C-Animation, GRID, or asks about evaluating this task. Reports WER.
Evaluates zero-shot segmentation accuracy on cellular and subcellular microscopy imagery, and assesses the downstream reliability of extracted morphological features for drug hit validation in high-content screening assays. Use when the user wants to benchmark on Cell segmentation datasets, Hit validation datasets, or asks about evaluating this task. Reports Dice Score (DSC), Z'-factor.
Evaluates the quality of distributed subgraph representations learned by subgraph2vec on graph classification and clustering tasks. It probes the model's ability to capture structural and semantic similarities in chemical/biological graphs and Android application control-flow graphs. Use when the user wants to benchmark on MUTAG, PTC, PROTEINS, NCI1, NCI109, CLONE260, TRAIN10K, TEST10K, or asks about evaluating this task. Reports Accuracy.
Evaluates the precision of statistical estimators (Empirical Bayes, Synthetic Regression, Direct Training) for estimating model performance on data subgroups with limited observations. It probes how well these methods reduce mean squared error and produce reliable confidence intervals when benchmarking LLMs, vision models, and tabular classifiers on niche tasks. Use when the user wants to benchmark on LLM MC QA tasks, Computer Vision tasks (LAION CLIP benchmark), COCO Captions, Tabular Fairne...
This benchmark evaluates text anonymization methods by measuring both span-level masking accuracy and subject-level privacy leakage. It probes whether anonymized text successfully prevents adversarial LLMs from inferring personal identifiable information (PII) and sensitive attributes, while maintaining text utility. Use when the user wants to benchmark on PANORAMA, TAB, or asks about evaluating this task. Reports CPR.
Evaluates models' ability to classify six subjective linguistic features (Assertive, Cautious, Optimistic, Specific, Clear, Relevant) in financial earnings call question-and-answer transcripts. It probes how well models capture nuanced, tone-based, and domain-specific communication cues beyond factual content. Use when the user wants to benchmark on SubjECTive-QA, or asks about evaluating this task. Reports weighted F1 score.
Evaluates the faithfulness of image attribution methods by measuring how prediction confidence changes as important regions are removed or added. It also probes the ability to identify specific image regions that cause model misclassifications. Use when the user wants to benchmark on Celeb-A, VGG-Face2, CUB-200-2011, or asks about evaluating this task. Reports Deletion AUC.
Evaluates the ability of machine learning and dynamical models to forecast subseasonal temperature and precipitation over the contiguous United States at 3–4 and 5–6 week lead times. It benchmarks predictive accuracy against operational baselines and tests robustness to high noise and spatial heterogeneity in climate data. Use when the user wants to benchmark on SubseasonalClimateUSA, or asks about evaluating this task. Reports % IMPROVEMENT OVER MEAN DEB. CFSV2 RMSE.
Evaluates the efficiency and compactness of subword tokenizers on Hindi, English, and code-mixed text by measuring how many tokens are generated per word and how frequently words are split into multiple tokens. Use when the user has predictions and gold and needs to compute Subword Fertility (SF).
Evaluates the robustness and generalization of 3D policy learning models for robotic manipulation across varying environmental conditions, temporal horizons, and real-world interference. It probes spatial understanding, fine-grained pose control, and resilience to domain randomization and lighting changes. Use when the user wants to benchmark on RoboTwin 2.0, ManiSkill2, Real-World Manipulation, or asks about evaluating this task. Reports Success Rate (%).
Evaluates the sum spectral efficiency of a hybrid centralized-distributed precoding scheme in fronthaul-constrained cell-free massive MIMO networks, comparing it against fully centralized and fully distributed baselines under varying fronthaul capacities and antenna configurations. Use when the user has predictions and gold and needs to compute sum spectral efficiency (sum SE).
This evaluation probes a model's ability to compress long documents into concise summaries while preserving topical coverage, cross-sentence coherence, and factual consistency under strict token budgets. It measures how well extractive or generative methods balance semantic relevance with structural discourse cues across diverse domains. Use when the user wants to benchmark on CNN/DailyMail, GovReport, arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-2.
Evaluates the factual consistency of human reference summaries across popular abstractive summarization benchmark datasets. It probes whether widely used datasets contain systematic factual errors or low-abstraction artifacts that compromise their validity as training and evaluation standards. Use when the user wants to benchmark on CNN/DM, XSUM, XL-Sum (English), or asks about evaluating this task. Reports Factuality Score.
Evaluates the correlation between automatic summarization metrics and LLM-as-a-Judge models against human judgments across five quality criteria (coherence, consistency, fluency, relevance, 5W1H) in Spanish and Basque. Use when the user wants to benchmark on BASSE, or asks about evaluating this task. Reports Spearman's $ ho$.
Compute the SumMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SumMetric, or asks how to score with SumMetric.
This benchmark evaluates how well automatic summarization metrics align with human judgments across multiple quality dimensions. It probes whether standard n-gram, embedding-based, and reference-less metrics reliably predict human-perceived coherence, consistency, fluency, and relevance of generated summaries. Use when the user wants to benchmark on SummEval, or asks about evaluating this task. Reports Kendall’s tau.
Evaluates the long-term driving consistency, trajectory quality, and lane-change efficiency of autonomous driving agents in varying traffic densities on simulated highways. Use when the user wants to benchmark on SUMO Highway Scenarios, or asks about evaluating this task. Reports average achieved speed.
Compute sunhill/cider via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sunhill/cider.
Compute sunhill/clip_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sunhill/clip_score.
Compute sunhill/spice via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sunhill/spice.
Evaluates a decentralized multi-agent reinforcement learning algorithm that selectively shares high-temporal-difference-error experiences. It probes cooperative and competitive multi-agent coordination, credit assignment, and communication efficiency in anonymous environments with separate per-agent reward signals. Use when the user wants to benchmark on PettingZoo (Pursuit, Battle, Adversarial-Pursuit), or asks about evaluating this task. Reports total mean episode reward.
Evaluates instruction-following and cross-task generalization capabilities of language models on a massive, diverse benchmark of 1,616 NLP tasks spanning 76 task types and 55 languages. It measures how well models trained on a mix of tasks perform on unseen tasks when given natural language instructions. Use when the user wants to benchmark on Super-NaturalInstructions, or asks about evaluating this task. Reports human evaluation metric.
Evaluates a model's ability to perform spatial super-resolution on global weather forecast data, specifically upscaling temperature and cloud coverage maps from 1° to 0.5° resolution. It measures pixel-wise reconstruction accuracy against high-resolution ground truth. Use when the user wants to benchmark on GraphCast-ERA5 Paired Dataset, or asks about evaluating this task. Reports MSE.
Evaluates pre-trained speech models on downstream spoken language understanding tasks. It probes the model's ability to classify spoken intents, fill semantic slots in transcriptions, and detect specific keywords in audio. Use when the user wants to benchmark on SUPERB, or asks about evaluating this task. Reports test accuracy.
Evaluates speech models on three downstream tasks: Intent Classification, Keyword Spotting, and Automatic Speech Recognition. It specifically probes noise robustness by comparing performance on clean speech versus speech corrupted with out-of-distribution CHiMe3 background noise. Use when the user wants to benchmark on SUPERB, or asks about evaluating this task. Reports Accuracy (Acc%).
Evaluates the effectiveness of a proactive validation system for cloud AI infrastructure in detecting hardware defects, selecting optimal benchmark subsets, and balancing validation cost against system reliability. It probes the ability of automated criteria and selection algorithms to distinguish healthy nodes from degraded ones while minimizing downtime and maximizing GPU utilization. Use when the user wants to benchmark on Cluster Benchmark Dataset, or asks about evaluating this task. Repo...
Evaluates deep chemical reasoning capabilities of LLMs using expert-curated, entity-masked multiple-choice problems. It probes both final-answer accuracy and the fidelity of the reasoning process against expert-annotated solution paths, while also assessing the impact of multimodal inputs on complex chemical problem-solving. Use when the user wants to benchmark on SUPERChem-A11, SUPERChem-release, SUPERChem-holdout, SUPERChem-100, Multimodal-Essential Subset, or asks about evaluating this tas...
This benchmark evaluates vision-language models' ability to function as intelligent agents for AI smart glasses in real-world egocentric scenarios. It probes capabilities in object detection, multi-hop reasoning, retrieval-augmented generation, and accurate answer formulation based on visual context and external knowledge. Use when the user wants to benchmark on SuperGlasses, or asks about evaluating this task. Reports accuracy.
Evaluates general-purpose language understanding across eight diverse tasks including coreference resolution, question answering, and natural language inference. It probes a model's ability to handle complex reasoning, sample-efficient learning, and transfer learning beyond standard GLUE capabilities. Use when the user wants to benchmark on SuperGLUE, or asks about evaluating this task. Reports SuperGLUE score.
Evaluates a language model's ability to follow diverse natural language instructions across various NLP tasks in a zero-shot setting. It measures how well the model generalizes to unseen tasks without in-context examples. Use when the user wants to benchmark on SUPER-NATURALINSTRUCTIONS, or asks about evaluating this task. Reports ROUGE-L.
Evaluates the ability of a predictor model to estimate the performance of instruction-following language models on unseen tasks, using only the task instruction as input. It probes the fundamental challenge of third-party model transparency and controllability at the task level. Use when the user wants to benchmark on SuperNI, or asks about evaluating this task. Reports RMSE.
Evaluates the ability of a gradient-free, locally-updated spiking neural network to classify images using population-level spike agreement metrics. It probes whether replacing backpropagation with supervised Spike Agreement-Dependent Plasticity (SADP) and Cohen’s κ can achieve competitive vision and biomedical classification performance while maintaining biological plausibility and hardware compatibility. Use when the user wants to benchmark on MNIST, Fashion-MNIST, CIFAR-10, LC25000, Brain M...
Evaluates models' ability to generate calibrated probabilistic forecasts of supply chain disruptions from raw news text. It probes temporal generalization, uncertainty quantification, and the prioritization of high-risk signals for decision-making. Use when the user wants to benchmark on Supply Chain Disruption Forecasting Dataset, or asks about evaluating this task. Reports Brier score.
This benchmark evaluates fine-grained spatial understanding and reasoning capabilities of vision-language models in real-world driving scenarios. It probes six distinct spatial dimensions: orientation (Yaw), pixel-level localization, depth estimation, pairwise distance, lateral ordering, and front-back relations. Use when the user wants to benchmark on SURDS, or asks about evaluating this task. Reports Score.
Evaluates a NeRF model's ability to reconstruct reflective scenes with high visual fidelity and accurate geometry. It specifically probes the model's capacity to separate Lambertian and specular appearance components while enforcing surface regularisation to resolve shape-radiance ambiguity. Use when the user wants to benchmark on Shiny Objects, Shiny Real, Koala, or asks about evaluating this task. Reports PSNR.
Evaluates multi-modal large language models on hierarchical spatiotemporal reasoning in surgical videos across five clinical dimensions: causal action ordering, cue-action alignment, affordance mapping, micro-transition localization, and anomaly onset tracking. The benchmark tests a model's ability to progressively narrow its focus from global video comprehension to fine-grained frame-level localization while maintaining logical consistency through a chain-of-thought protocol. Use when the us...
Evaluates a model's ability to predict the next item in a user's interaction sequence by leveraging long-term historical behavior and filtering out noise. It probes the model's capacity to handle varying sequence lengths and efficiently model dynamic user preferences over time. Use when the user wants to benchmark on Taobao, Kuaishou, or asks about evaluating this task. Reports GAUC.
Evaluates the correctness and efficiency of gradient-based versus boolean logic-based methods for identifying feature-parameter interactions and transferring trained weights when new features are added to a reinforcement learning model. It measures how well each mapping technique preserves model performance and computational speed during architectural surgery. Use when the user has predictions and gold and needs to compute Interactions Found.
Evaluates zero-shot and linear-probe transferability of surgical vision-language models across laparoscopic and robotic video modalities. Probes hierarchical procedural understanding (phase, step, action) and object-centric recognition (tools, instrument-verb-target triplets) under varying temporal contexts and data regimes. Use when the user wants to benchmark on Cholec80, AutoLaparo, GraSP, SARRARP50, CholecT50, or asks about evaluating this task. Reports video-wise F1-score.
Evaluates zero-shot surgical video generation models using a four-tiered Surgical Plausibility Pyramid. It probes the model's ability to maintain visual realism while correctly simulating domain-specific surgical causality, instrument handling, tissue feedback, and clinical intent over time. Use when the user wants to benchmark on SurgVeo benchmark, or asks about evaluating this task. Reports Visual Perceptual Plausibility.
This benchmark evaluates language-guided spatial understanding and reasoning in complex 3D scenes. It probes a model's ability to reason about relative positions, narrative/parametric perspectives, and absolute distances without relying on semantic shortcuts or explicit object names in the prompts. Use when the user wants to benchmark on SURPRISE3D, or asks about evaluating this task. Reports Accuracy (A25/A50).
Probes the ability of surrogate-assisted genetic algorithms to optimize continuous 2D functions under tight evaluation budgets, simulating interactive recommendation scenarios where user feedback is sparse and dynamic. Use when the user wants to benchmark on Bohachevsky, Ackley, and Schwefel benchmark functions, or asks about evaluating this task. Reports Best Fitness.
Evaluates a model's ability to detect and temporally localize anomalous events in long, untrimmed surveillance videos using only video-level labels. It probes the model's robustness to high intra-class variation, ambiguous normal-anomalous boundaries, and varying lighting/occlusion conditions. Use when the user wants to benchmark on Surveillance Anomaly Dataset, or asks about evaluating this task. Reports AUC.