
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the ability of conformal prediction frameworks to produce statistically valid prediction sets with instance-level uncertainty quantification for encoder-only transformers, measuring both classification accuracy and calibration efficiency across standard NLP benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, or asks about evaluating this task. Reports Test Accuracy.
This benchmark evaluates a model's ability to accurately answer conference submission checklist questions based on manuscript content. It specifically probes long-form document understanding, retrieval-augmented generation (RAG) effectiveness, and the model's capacity to reflect on ethical considerations, reproducibility, and societal impacts. Use when the user wants to benchmark on ConfReady Evaluation Set, or asks about evaluating this task. Reports Accuracy.
Compute the ConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ConfusionMatrix, or asks how to score with ConfusionMatrix.
Evaluates a model's ability to perform Named Entity Recognition (NER) across multiple languages, specifically testing its robustness to out-of-domain text, orthographic variations, and cross-lingual transfer when trained on noisy Wikipedia-derived data. Use when the user wants to benchmark on CoNLL 2002/2003 NER, or asks about evaluating this task. Reports Exact F1.
Evaluates a model's ability to jointly detect mentions and resolve coreferential chains in English text without relying on external syntactic parsers or hand-crafted features. It measures how well the model groups word spans into entity clusters based on contextual and structural cues. Use when the user wants to benchmark on CoNLL-2012 (English), or asks about evaluating this task. Reports F1.
Evaluates a model's ability to locate and connect dots in sequential order across various visual patterns. It probes precise spatial reasoning and the capacity to generate non-destructive SVG overlays that explain the reasoning process. Use when the user wants to benchmark on Connect-the-Dots, or asks about evaluating this task. Reports Accuracy.
Evaluates a multi-metric layer pruning method (Consensus) on image classification models, measuring trade-offs between computational efficiency (FLOPs reduction) and predictive performance (accuracy drop), while also assessing robustness against adversarial and out-of-distribution attacks. Use when the user wants to benchmark on CIFAR-10, ImageNet, CIFAR-10.2, CIFAR-C, ImageNet-C, or asks about evaluating this task. Reports Δ Acc. (difference in accuracy).
Evaluates the reproducibility and deterministic behavior of generative AI models (diffusion and LLMs) by measuring the likelihood of identical outputs given identical prompts and seeds. It probes the stability of model inference and training under decentralized, heterogeneous hardware conditions. Use when the user has predictions and gold and needs to compute consensus.
Compute the consensus_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute consensus_score, or asks how to score with consensus_score.
Evaluates LLM generalization and functional consistency by measuring how well models preserve core functionality after iterative, reversible transformations. It probes cumulative error and path-specific divergence across multi-step transformation sequences without relying on static benchmarks. Use when the user wants to benchmark on ConsistencyChecker (Dynamic), or asks about evaluating this task. Reports forest-level consistency score (C3(F)).
This evaluation probes how different instructional guidelines (constitutions) shape AI-generated medical dialogues across specific socio-communicative dimensions like empathy, information gathering, and decision-making. It measures human preference for dialogue quality under varying constitutional constraints. Use when the user wants to benchmark on Custom AI-generated medical dialogues, or asks about evaluating this task. Reports Bradley-Terry preference rate.
Evaluates the adversarial robustness of tabular deep learning models under realistic, domain-aware constraints. It measures how easily an attacker can flip model predictions while respecting feature mutability, boundaries, types, and relational constraints across ten progressively restricted threat models. Use when the user wants to benchmark on phishing, credit scoring, botnet detection, or asks about evaluating this task. Reports adversarial_label_flip.
Evaluates a graph convolutional neural network's ability to predict optimal next-node preferences for constrained shortest path problems with mandatory waypoints, to accelerate constraint programming solvers. Use when the user wants to benchmark on Maneuver benchmark, Exploration benchmark, or asks about evaluating this task. Reports Number of instances resolved with proof of optimality.
Evaluates the trustworthiness and accuracy of LLM-generated structured outputs (JSON) against a ground truth or expected schema. It probes the model's ability to detect per-field and per-document errors in data extraction tasks without requiring labeled data. Use when the user wants to benchmark on Four real-world datasets (unspecified in excerpt), or asks about evaluating this task. Reports rating.
Evaluates vision-language models on construction site safety inspection tasks, including image captioning, safety rule violation detection, reasoning, and visual grounding of specific objects. Use when the user wants to benchmark on ConstructionSite 10k, or asks about evaluating this task. Reports IoU.
Evaluates the success rate of imitation learning policies in contact-rich manipulation tasks requiring precise force control, slip detection, and in-hand pose estimation. It compares vision-only baselines against visuo-tactile policies with and without temporal-aware contrastive pretraining. Use when the user wants to benchmark on Contact-Rich Manipulation Tasks, or asks about evaluating this task. Reports success rate.
Measures the extent to which multimodal evaluation benchmarks are contaminated by pre-training data, assessing both visual similarity and textual inference leakage to quantify data contamination risks. Use when the user has predictions and gold and needs to compute image-only contamination rate.
Evaluates how language models merge conflicting generated and retrieved contexts in open-domain QA. It probes whether models exhibit a systematic bias toward generated contexts over retrieved ones when only one context contains the correct answer. Use when the user wants to benchmark on NQ-CC, TQA-CC, or asks about evaluating this task. Reports DiffGR.
Evaluates a model's ability to integrate essential natural language context with numerical time series data to produce accurate forecasts. It probes multimodal reasoning, constraint satisfaction, and the capacity to leverage textual information for improving time series prediction. Use when the user wants to benchmark on CiK, or asks about evaluating this task. Reports RCRPS.
This benchmark evaluates session-based recommendation models by predicting the immediate next item in a user's interaction sequence. It probes the model's ability to capture sequential dependencies and adapt to evolving user preferences and new items in both static and continuously updating environments. Use when the user wants to benchmark on MOOC, News, and RecSys Challenge Datasets, or asks about evaluating this task. Reports HR@k.
Evaluates how well ASR models and Large Audio Language Models (LALMs) leverage contextual world knowledge and linguistic reasoning to transcribe speech containing named entities. It tests performance across ten domains under three context conditions: no context, coarse-grained domain labels, and fine-grained technical terms. Use when the user wants to benchmark on ContextASR-Bench, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates conversational context recall and utilization in voice interaction models, specifically measuring how well they remember and respond to past user and system utterances in multi-turn dialogues. It also probes the robustness of retrieval-augmented generation (RAG) when applied to speech-based models. Use when the user wants to benchmark on ContextDialog, or asks about evaluating this task. Reports GPT Score.
Evaluates whether integrating multimodal contextual metadata into pre-trained time series forecasting models improves prediction accuracy. It probes the model's ability to align external covariates with historical time series data to enhance forecast precision across multiple domains and horizons. Use when the user wants to benchmark on Synthetic ARMA(2,2), PEMS-SF, ETT (ETTm2), ECL, Beijing AQ, Store Sales, Monash (Bitcoin), Bitcoin + News, or asks about evaluating this task. Reports MSE.
Evaluates zero-shot text-to-video retrieval for contextual advertising. It probes a model's ability to rank relevant video content based on natural language queries using multimodal signals (vision, audio, captions, metadata). Use when the user wants to benchmark on Val-1, Val-2, or asks about evaluating this task. Reports P@K.
Evaluates speech-to-text systems on their ability to correctly recognize domain-specific custom vocabulary (e.g., company names, products) in real-world earnings call audio. It probes how well models leverage provided keyword contexts (local vs. global/noisy) to improve keyword recognition without introducing transcription artifacts. Use when the user wants to benchmark on Contextual Earnings-22, or asks about evaluating this task. Reports keyword F-score.
Evaluates information retrieval systems by measuring system-level performance metrics (dead links, response time, redundancy) and user-perceived relevance across different query topics and rank positions. Use when the user wants to benchmark on Custom IR Evaluation Corpus, or asks about evaluating this task. Reports Relevance Judgments.
Probes a multimodal large language model's ability to infer and localize objects within human-AI interaction contexts (e.g., cloze tests, captioning, QA) using open-vocabulary object names, rather than fixed class sets. Use when the user wants to benchmark on CODE, or asks about evaluating this task. Reports Acc@1.
Evaluates a model's ability to detect sarcasm in contextual settings across Reddit comments, tweets, and multi-turn dialogues. It probes the capacity to capture sentiment incongruity and contextual cues rather than relying on surface-level lexical features. Use when the user wants to benchmark on SARC 2.0, Twitter, Sarcasm Corpus V2 Dialogues, or asks about evaluating this task. Reports F1-Score.
Evaluates the ability of large multimodal models to sequentially learn new instruction-following tasks without catastrophically forgetting previously acquired capabilities. It measures both retained performance on old tasks and the degree of forgetting across sequential training stages. Use when the user wants to benchmark on Flickr30k, TextCaps, VQA v2, OCR-VQA, GQA, VizWiz, TextVQA, or asks about evaluating this task. Reports Average performance ($A_t$).
Evaluates a model's ability to learn sequentially across multiple tasks without catastrophic forgetting. It measures how well the model retains accuracy on previously learned tasks while adapting to new ones. Use when the user wants to benchmark on Split MNIST, Permuted MNIST, Split CIFAR-10/100, or asks about evaluating this task. Reports average classification accuracy.
Evaluates continual learning techniques for malware classification under domain, class, and task incremental settings. It measures how well models adapt to evolving malware distributions without catastrophic forgetting. The protocol compares complex CL methods against simple baselines like joint replay. Use when the user wants to benchmark on Drebin, EMBER, or asks about evaluating this task. Reports Mean accuracy.
Evaluates a model's ability to retain knowledge from previously learned tasks while continuously training on new ones, and measures how past knowledge facilitates learning new tasks and improves performance on old ones. Use when the user has predictions and gold and needs to compute Average Performance (AP).
This benchmark evaluates a model's ability to sequentially learn a mix of visual understanding and generation tasks without catastrophically forgetting previously acquired knowledge. It specifically probes intra-modal retention (maintaining performance on earlier tasks) and inter-modal stability (preventing updates for one modality from degrading the other). Use when the user wants to benchmark on ScienceQA, TextVQA, GQA, VizWiz, ImageNet, CustomConcept101, or asks about evaluating this task....
Evaluates a model's ability to perform continual learning in Named Entity Recognition (CL-NER) by incrementally learning new entity types while mitigating catastrophic forgetting of previously learned types. It specifically probes how well the model handles the 'Other-class' (miscellaneous/old entities) during incremental training and maintains performance across sequential learning steps. Use when the user wants to benchmark on OntoNotes5, i2b2, CoNLL2003, or asks about evaluating this task....
Evaluates a model's ability to maintain safety alignment and task performance during sequential continual fine-tuning across multiple domains. It probes whether gradient-based sample selection prevents safety degradation (elastic reversion) and catastrophic forgetting while preserving general capabilities. Use when the user wants to benchmark on AdvBench, HarmBench, TruthfulQA, ARC-C, BoolQ, HellaSwag, Winogrande, GSM8K, MedMCQA, Squad_v2, or asks about evaluating this task. Reports ASR.
Compute the ContinuousRankedProbabilityScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ContinuousRankedProbabilityScore, or asks how to score with ContinuousRankedProbabilityScore.
Evaluates generalizable Neural Radiance Field (NeRF) methods for novel view synthesis, specifically probing their ability to generalize from synthetic training data to real-world indoor and outdoor scenes. It measures rendering quality and geometric consistency across different domain gaps. Use when the user wants to benchmark on 3D-FRONT, ScanNet, DTU, LLFF, Google Scanned Object, or asks about evaluating this task. Reports PSNR.
Evaluates sequential recommendation models on Amazon review datasets, measuring ranking quality for next-item prediction. It specifically probes performance on long-tail items, sequence sparsity, and robustness to noisy inputs. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Amazon Tools, Amazon Office, or asks about evaluating this task. Reports Recall@20.
Evaluates controllable image generation based on visual conditions (segmentation masks, edges, depth maps) by measuring how closely the generated image's extracted conditions match the input conditions. It tests spatial and structural controllability. Use when the user wants to benchmark on ControlNet++ dataset, or asks about evaluating this task. Reports mIoU (Seg. Mask).
This evaluation probes a model's ability to rewrite conversational search queries to maximize retrieval effectiveness. It measures how well reformulated questions help both sparse and dense retrievers locate relevant passages across different dialogue contexts, including initial turns and topic shifts. Use when the user wants to benchmark on QReCC, TopiOCQA, or asks about evaluating this task. Reports MRR.
Evaluates open-domain chatbot capabilities in persona-driven multi-turn conversations, measuring response quality via automatic metrics and human judgments of engagement and persona consistency. Use when the user wants to benchmark on PERSONA-CHAT, or asks about evaluating this task. Reports Engagingness.
Evaluates conversational dense retrieval models on their ability to rank relevant documents given multi-turn conversational queries. It probes context capture, few-shot learning effectiveness, and robustness to noisy conversation history compared to query rewriting baselines. Use when the user wants to benchmark on TREC CAsT, OR-QuAC, or asks about evaluating this task. Reports NDCG@3, MRR@5.
Evaluates knowledge graph link prediction by measuring a model's ability to infer missing entities (head or tail) from subject-relation triples. It probes the expressiveness of 2D convolutional embeddings and tests robustness against test-set leakage via inverse relations. Use when the user wants to benchmark on WN18, FB15k, YAGO3-10, Countries, FB15k-237, WN18RR, or asks about evaluating this task. Reports MRR.
Probes a model's ability to reconstruct reply relationships in multi-party, entangled text conversations. It requires identifying which message responds to which, handling simultaneous conversations, and distinguishing reply edges from system or directed messages. Use when the user wants to benchmark on IRC Conversation Disentanglement Corpus, or asks about evaluating this task. Reports F1 score for reply-edge prediction.
Assesses the linguistic quality and conversational realism of LLM-synthesized Knowledge Graph QA turns. Probes fluency, factual relevance, diversity, and grammatical correctness across varied interaction styles and noise augmentations. Use when the user wants to benchmark on ConvKGYarn, or asks about evaluating this task. Reports Fluency, Relevance, Diversity, Grammar & Agreement.
Evaluates a convolutional neural network's ability to detect seismic events versus background noise and classify their geographic origin using raw waveform data. It probes the model's generalization to unseen temporal periods and non-repeating seismic events. Use when the user wants to benchmark on Oklahoma Seismic Dataset (OGS), or asks about evaluating this task. Reports detection accuracy.
Evaluates conversational memory capabilities across six dimensions: recalling user facts, tracking assistant statements, abstaining when information is missing, inferring preferences, handling changing facts, and making implicit connections. It specifically tests the ability to synthesize evidence distributed across multiple conversation turns. Use when the user wants to benchmark on ConvoMem, or asks about evaluating this task. Reports accuracy.
Evaluates conversational query reformulation (CQR) by measuring how effectively a model rewrites multi-turn queries into standalone search queries that retrieve relevant passages. It probes the model's ability to optimize rewrites using only retrieval signals, without human annotations or LLM distillation. Use when the user wants to benchmark on TopiOCQA, QReCC, or asks about evaluating this task. Reports MRR@3.
Evaluates the ability of protein language models to adapt to evolving biological databases through continual pretraining. It probes how well models maintain performance on high-quality sequence validation, predict mutation fitness effects, and generalize across diverse protein understanding tasks over time. Use when the user wants to benchmark on UniProt Validation Set, ProteinGym, PEER, DGEB, or asks about evaluating this task. Reports Spearman correlation.
Evaluates the ability of Multimodal Large Language Models to generate factually grounded captions and answers while suppressing object-level hallucinations. It probes visual grounding, reasoning consistency, and alignment with human or GPT-4 preferences across multiple reasoning and perception benchmarks. Use when the user wants to benchmark on CHAIR, POPE, MMBench, MME, or asks about evaluating this task. Reports POPE F1 Score.