
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the robustness of Out-of-Distribution (OOD) detection models under realistic distribution shifts caused by semantic-preserving transformations. It measures how well detectors distinguish between true out-of-distribution samples and inlier samples that have undergone common corruptions or augmentations, revealing performance gaps that standard benchmarks miss. Use when the user wants to benchmark on CIFAR-10-R, CIFAR-100-R, ImageNet-30-R, or asks about evaluating this task. Reports A...
Evaluates scientific machine learning models on sim-to-real transfer for complex physical systems. It probes a model's ability to predict spatiotemporal dynamics from real-world measurements, leveraging simulated pretraining, and assesses long-term prediction stability under autoregressive rollout. Use when the user wants to benchmark on Cylinder, ControlledCylinder, FSI, Foil, Combustion, or asks about evaluating this task. Reports RMSE.
Evaluates end-to-end simultaneous speech-to-speech translation quality and latency on long-form, multi-domain continuous speech. Probes the model's ability to maintain semantic accuracy, speaker voice characteristics, and low delay in real-time multilingual dialogue. Use when the user wants to benchmark on RealSI, or asks about evaluating this task. Reports VIP.
Evaluates the scalability and performance of a simulation-checking algorithm for timed automata under fairness assumptions. It measures how efficiently the algorithm verifies liveness properties and handles state-space explosion across parameterized real-time system benchmarks. Use when the user wants to benchmark on Fischer's timed mutual exclusion algorithm, CSMA/CD, Timed consumer/producer, Network of TAs, or asks about evaluating this task. Reports CPU time.
Evaluates the ability of 3D reconstruction models to synthesize novel views of smoke-degraded scenes with high photometric fidelity and structural preservation. It probes view-dependent medium modeling and multi-view consistency under severe scattering conditions. Use when the user wants to benchmark on RealX3D (NTIRE 2026 Track 2 Smoke Subset), or asks about evaluating this task. Reports PSNR.
Evaluates large language models' logical reasoning and problem-solving capabilities across mathematical, algorithmic, and creative tasks. It measures both the correctness of final answers and the computational efficiency of the reasoning process. Use when the user wants to benchmark on Game of 24, BIG-Bench (subset), Python Puzzles, MGSM, Shakespearean Sonnet Writing, or asks about evaluating this task. Reports Acc_logic.
Evaluates language model reasoning and general knowledge across multiple standard benchmarks. It measures the impact of data mixture optimization on model performance using accuracy scores on commonsense, scientific, and factual QA tasks. Use when the user wants to benchmark on PIQA, ARC_C, ARC_E, HellaSwag, WinoGrande, SIQA, MMLU, or asks about evaluating this task. Reports test accuracy.
Evaluates LLM logical and spatial reasoning capabilities under multi-agent collaboration and memory-augmented prompting. It probes how different reasoning styles, exemplar retrieval methods, and answer aggregation strategies impact accuracy on formal logic and object-tracking tasks. Use when the user wants to benchmark on FOLIO, RACO, TSO, or asks about evaluating this task. Reports accuracy.
This evaluation protocol assesses the cross-domain generalization, safety, and instruction-following capabilities of models after reasoning-focused supervised fine-tuning (SFT). It measures in-domain math performance, out-of-domain reasoning in coding and science, general instruction following, and resistance to harmful queries. Use when the user wants to benchmark on MATH500, AIME24, LiveCodeBench v2, GPQA-Diamond, MMLU-Pro, IFEval, AlpacaEval 2.0, HaluEval, TruthfulQA, HEx-PHI, or asks abou...
Probes a model's ability to infer implicit human intentions from natural language and generate multi-step, route-aware activity plans grounded in 3D scene segmentation. It evaluates both textual planning coherence and spatial reasoning over 3D environments. Use when the user wants to benchmark on ReasonPlan3D, or asks about evaluating this task. Reports BLEU-4.
Evaluates a model's ability to generate precise segmentation masks from implicit, complex text queries that require reasoning and world knowledge. It specifically probes whether the model can move beyond simple explicit referring expressions to handle multi-step logical deductions and visual grounding simultaneously. Use when the user wants to benchmark on ReasonSeg, refCOCO, refCOCO+, refCOCOg, or asks about evaluating this task. Reports gIoU.
Evaluates the predictive performance of recommendation models on large-scale click-through rate datasets. It specifically probes how model scalability and embedding size affect ranking quality, revealing the phenomenon of embedding collapse when scaling up feature interactions. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
Evaluates how different data splitting strategies (leave-one-last-item, leave-one-last-basket, temporal global) impact the performance ranking of recommendation models on e-commerce datasets. It probes whether evaluation protocols introduce temporal leakage or distribution shifts that confound model comparisons and invalidate cross-paper rankings. Use when the user wants to benchmark on Tafeng Dataset, Dunnhumby Dataset, or asks about evaluating this task. Reports NDCG@10.
Compute the recall_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute recall_score, or asks how to score with recall_score.
Evaluates language models on associative recall, information extraction, and question answering from long contexts, while measuring generation throughput and language modeling perplexity. It probes the tradeoff between memory efficiency, recall accuracy, and inference speed across synthetic and real-world benchmarks. Use when the user wants to benchmark on Pile, SWDE, FDA, SQUAD, LM Eval Harness (SuperGLUE, ARC, PIQA, WinoGrande, HellaSwag, LAMBADA), or asks about evaluating this task. Report...
Compute the Recall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute Recall, or asks how to score with Recall.
Compute the RecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RecallAtFixedPrecision, or asks how to score with RecallAtFixedPrecision.
Evaluates the semantic fidelity, object accuracy, and prompt adherence of text-to-image generation models by comparing automated metrics and human ratings on standard benchmarks. Use when the user wants to benchmark on MS-COCO validation set, DrawBench, or asks about evaluating this task. Reports FID.
Evaluates multilingual receipt understanding across four tasks: question answering, object detection/classification, OCR, and information extraction. Probes model capabilities in handling real-world noise, mixed Arabic-English layouts, and complex formatting. Use when the user wants to benchmark on ReceiptSense, or asks about evaluating this task. Reports exact match.
Tests an algorithm's ability to optimally place a receiver in 3D indoor environments to maximize speech intelligibility, measured by the Speech Transmission Index (STI). It evaluates how well the optimization handles complex acoustic properties like reverberation and noise across different scene geometries. Use when the user wants to benchmark on Office, Berlin, Suburban 3D scenes, or asks about evaluating this task. Reports STI.
Evaluates cross-modal retrieval between food images and cooking recipes. It probes a model's ability to align visual and textual representations in a shared embedding space to rank relevant recipes given an image, and vice versa. Use when the user wants to benchmark on Recipe1M+, or asks about evaluating this task. Reports medR.
Evaluates an end-to-end autonomous driving agent's ability to generate safe, comfortable, and efficient driving trajectories using only camera inputs. It probes the model's closed-loop planning capabilities, safety-critical scenario handling, and visual reasoning in complex urban environments. Use when the user wants to benchmark on NAVSIM, Bench2Drive, or asks about evaluating this task. Reports PDMS.
Evaluates how well different 3D reconstruction methods perform in a downstream object pose estimation task, rather than measuring standalone geometric reconstruction accuracy. It compares pose estimation results using reconstructed 3D models against those using ground-truth CAD models. Use when the user wants to benchmark on YCB-V, or asks about evaluating this task. Reports accuracy of the estimated poses.
This benchmark evaluates multimodal models on predicting continuous personality traits and interview performance scores from video, audio, and text inputs. It probes the model's ability to perform fine-grained behavioral analysis and regression across psychometric targets. Use when the user wants to benchmark on RecruitView, or asks about evaluating this task. Reports Spearman's ρ.
Evaluates session-based recommendation models by predicting the next item in a user's clickstream sequence. It probes the model's ability to capture temporal dynamics and handle data sparsity in e-commerce sessions. Use when the user wants to benchmark on RecSys Challenge 2015, or asks about evaluating this task. Reports Recall@20.
Evaluates session-based recommendation models by predicting the next item in a user's browsing sequence. It measures ranking quality and prediction efficiency to assess accuracy and deployability in real-time recommender systems. Use when the user wants to benchmark on RecSys Challenge 2015 dataset, or asks about evaluating this task. Reports Recall@20.
Evaluates machine theory of mind in LLM-based conversational recommender systems by testing cognitive inference (fine/coarse intention, belief) and behavioral prediction (prediction, judgement) for both recommender and seeker roles in dialogue scenarios. Use when the user wants to benchmark on RECTOM, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of a diffusion-based regularization framework to reconstruct high-resolution subsurface velocity models from seismic data. It probes robustness under varying data conditions, including clean recordings, Gaussian noise contamination, and missing traces. The benchmark also tests out-of-distribution generalization on complex geological structures. Use when the user wants to benchmark on OpenFWI, Marmousi, or asks about evaluating this task. Reports RMSE.
Evaluates the quality of AI-generated adversarial responses across three dimensions: adherence to malicious objectives (toxicity), logical/semantic consistency (coherence), and textual variation (diversity). It assesses whether a model can produce high-quality, diverse, and coherent toxic content for red-teaming without suffering from reward hacking or semantic drift. Use when the user wants to benchmark on Curated Red-Teaming Dataset, or asks about evaluating this task. Reports Toxicity-Util...
Compute red1bluelost/evaluate_genericify_cpp via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of red1bluelost/evaluate_genericify_cpp.
Evaluates LLM robustness against adversarial prompts (Attack Success Rate) and their tendency to over-defend on benign prompts (Rejection Rate). It probes safety alignment, refusal behavior, and cross-domain vulnerability across 22 risk categories and 19 domains. Use when the user wants to benchmark on RedBench, or asks about evaluating this task. Reports Rejection Rate (RR), Attack Success Rate (ASR).
This evaluation probes the ability of automated red-teaming agents to successfully jailbreak diverse code-generating AI assistants. It measures how effectively an attacker can craft and optimize malicious prompts to bypass safety guardrails and force the execution of harmful code across multiple programming languages and agent architectures. Use when the user wants to benchmark on RedCode-Exec, RedCode-Gen, RMCbench, or asks about evaluating this task. Reports attack success rate (ASR).
This benchmark evaluates zero-shot large language models on their ability to classify suicide risk severity from Reddit posts using the clinically validated Columbia-Suicide Severity Rating Scale (C-SSRS). It probes the models' ordinal classification capabilities, intent detection, and alignment with human clinical annotations across seven severity levels. Use when the user wants to benchmark on Reddit r/SuicideWatch posts (C-SSRS labeled), or asks about evaluating this task. Reports F1-Score.
This evaluation probes a model's ability to generate abstractive summaries from informal, user-generated text and formal documents. It measures how well the model captures long-range dependencies and abstracts key information without relying on extractive heuristics. Use when the user wants to benchmark on Reddit TIFU, Newsroom-Abs, XSum, or asks about evaluating this task. Reports ROUGE-1.
Evaluates modular components of a conversational recommendation system, specifically cold-start movie rating prediction and movie opinion sentiment analysis (seen/liked status) from dialogue text. Use when the user wants to benchmark on REDIAL, MovieLens, or asks about evaluating this task. Reports RMSE.
Evaluates large language models' ability to answer questions about rare diseases, including diagnosis, symptoms, causes, and related properties. It probes the models' medical knowledge retrieval and reasoning capabilities in a specialized, low-resource domain. Use when the user wants to benchmark on ReDis-QA, or asks about evaluating this task. Reports accuracy.
Evaluates reinforcement fine-tuning methods for red-teaming LLMs by measuring the toxicity and diversity of generated adversarial prompts across toxic continuation and instruction-following tasks. Use when the user wants to benchmark on toxic continuation, instruction following, or asks about evaluating this task. Reports cumulative toxicity-diversity score.
Evaluates the model's ability to solve complex mathematical, coding, and general reasoning problems. It probes multi-step reasoning, domain-specific knowledge integration, and long-context handling across diverse benchmarks. Use when the user wants to benchmark on Math & Reasoning Benchmarks, Hellobench, SedarEval, Chinese Graduate Entrance Mathematics Test, or asks about evaluating this task. Reports AVG.
Evaluates a model's ability to determine whether pairs of P, I, or O spans in an abstract refer to the same underlying information. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
Evaluates an AI model's ability to classify underwater substrates for autonomous coral reseeding deployment. It probes both fine-grained patch-level semantic segmentation (distinguishing coral, deploy, and no-deploy zones) and coarse-grained image-level decision making for real-time marine robotics. Use when the user wants to benchmark on Great Barrier Reef ReefScan Dataset, or asks about evaluating this task. Reports Macro F1.
Probes multimodal large language models' ability to perform visual grounding and complex textual reasoning under challenging conditions. It specifically tests whether models rely on shortcut cues or genuinely comprehend referring expressions when faced with linguistically nontrivial descriptions and hard distractors. Use when the user wants to benchmark on Ref-Adv, or asks about evaluating this task. Reports Acc0.5.
Evaluates a model's ability to segment objects in audio-visual videos based on natural language referring expressions. It tests both seen categories and generalization to unseen categories, as well as handling null references where no object exists. Use when the user wants to benchmark on Ref-AVS Dataset, or asks about evaluating this task. Reports Jaccard Index ($\mathcal{J}$).
This benchmark evaluates large language models' ability to detect, localize, and correct scientific confabulations in generated answers. It probes fine-grained factuality awareness, span-level error identification, and factual restoration capabilities under domain-specific scrutiny. Use when the user wants to benchmark on ReFACT, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to localize a target object in an aerial image based on a fine-grained natural language description. It probes cross-modal alignment, scale-invariant object detection, and handling of complex backgrounds with numerous distractors. Use when the user wants to benchmark on RefAerial, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports mP (average P@0.5/0.6/0.7/0.8).
This benchmark evaluates a model's ability to answer questions about chart images while simultaneously localizing the visual evidence (via bounding boxes) that supports the answer. It probes spatial-text alignment, arithmetic and logical reasoning over charts, and hallucination reduction through explicit grounding. Use when the user wants to benchmark on RefChartQA, or asks about evaluating this task. Reports answer accuracy.
Evaluates fine-grained visual grounding and spatial reasoning by measuring how accurately a multimodal model can localize regions in images corresponding to given referring expressions under varying linguistic and spatial conditions. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports IoU@50 accuracy.
Evaluates a model's ability to perform referring expression segmentation at both object and part levels. It probes fine-grained cross-modal alignment and pixel-level semantic understanding by requiring precise mask prediction for diverse textual references. Use when the user wants to benchmark on RefCOCOm, or asks about evaluating this task. Reports mIoU.
Evaluates instruction-based image editing models on referring expressions, measuring how well they align edits with text instructions while preserving background and maintaining perceptual quality. It probes multi-object scene editing, background preservation, and the ability to handle complex spatial grounding without relying on CLIP. Use when the user wants to benchmark on RefEdit-Bench, PIE-Bench, or asks about evaluating this task. Reports VIEScore.
Evaluates Multimodal Large Language Models (MLLMs) on automatic sports refereeing tasks, probing their ability to detect incidents, classify fouls, apply sport-specific rules, and ground decisions temporally across 11 different sports. Use when the user wants to benchmark on RefereeBench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to perform pixel-level image segmentation conditioned on natural language expressions. It probes spatial reasoning, attribute grounding, and fine-grained visual-linguistic alignment by requiring the model to segment specific objects or amorphous regions described in text. Use when the user wants to benchmark on ReferIt, or asks about evaluating this task. Reports prec@0.5.