
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates whether representation learning models capture biologically meaningful signals in high-content microscopy images. It probes the model's ability to distinguish drug-induced perturbations from controls, predict zero-shot drug-target interactions, and recover known gene-gene relationships from phenotypic embeddings. Use when the user wants to benchmark on RxRx3-core, or asks about evaluating this task. Reports average precision.
Evaluates models' ability to transcribe spoken mathematical equations and sentences into correct LaTeX syntax. It probes audio-to-text conversion, handling of mathematical symbols, and robustness to syntactic variations in LaTeX formatting. Use when the user wants to benchmark on S2L-equations, S2L-sentences, or asks about evaluating this task. Reports CER.
Evaluates the skill of deep learning post-processing models for global sub-seasonal temperature and precipitation forecasts against climatological baselines and ECMWF recalibrated forecasts. It probes the ability of spatial CNN architectures to correct systematic errors and produce well-calibrated probabilistic tercile predictions over a 2–4 week horizon. Use when the user wants to benchmark on S2S AI Challenge test set (2020), or asks about evaluating this task. Reports RPSS.
Evaluates speech-to-speech models on instruction following, assessing both semantic correctness and paralinguistic/speech quality in a reference-free, head-to-head comparison. Use when the user wants to benchmark on S2S-Arena, or asks about evaluating this task. Reports ELO score.
Evaluates whether multimodal LLMs and agentic systems can reliably generate actionable decision support from operational S2S climate service products. It probes three core capabilities: actionable signal comprehension, uncertainty-conditioned decision-making handoffs, and evidence-grounded planning under dynamic hazards. Use when the user wants to benchmark on S2SServiceBench, or asks about evaluating this task. Reports Rubric Score (CT, ACT, TTH, EG, FC, UC).
Evaluates large language models' ability to perform complex, multi-step reasoning and data dependency tracking by executing SQL queries over synthetic, arbitrarily long tables. It probes contextual reasoning, long-context understanding, and structured data manipulation capabilities beyond traditional benchmarks. Use when the user wants to benchmark on S3Eval, or asks about evaluating this task. Reports SQL execution performance.
Compute the SacreBLEUScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SacreBLEUScore, or asks how to score with SacreBLEUScore.
Evaluates sparse autoencoder (SAE) architectures across multiple dimensions including reconstruction fidelity, feature disentanglement, concept detection, and practical interpretability tasks. It systematically compares how different SAE designs, dictionary sizes, and sparsity levels impact these capabilities. Use when the user wants to benchmark on Gemma-2-2B, Pythia-160M, or asks about evaluating this task. Reports Loss Recovered.
Evaluates the safety alignment and refusal capabilities of Video Large Multimodal Models (VLMMs) against everyday adversarial queries and covert, human-red-teamed prompts. It measures whether models can maintain safety guidelines across diverse harmful categories without compromising general utility or falling back to memorized refusals. Use when the user wants to benchmark on SafeVidBench, or asks about evaluating this task. Reports Safety Rate.
Evaluates the ability of reinforcement learning agents to maintain safety constraints and retain knowledge across sequentially changing non-stationary robotic environments, measuring the trade-off between task performance, safety violations, and catastrophic forgetting. Use when the user wants to benchmark on Damaged HalfCheetah Velocity, Damaged Ant Velocity, Safe Continual World, or asks about evaluating this task. Reports Final Task Reward.
Evaluates a hybrid trajectory planning framework that fuses flow matching with model-predictive control. It probes the system's ability to generate adaptive, collision-free motions for a 7-DoF robot manipulator while strictly enforcing safety constraints in real-time across global planning, reactive replanning, and dynamic human-robot handover scenarios. Use when the user wants to benchmark on Custom Robot Manipulation Benchmarks (Exp 1-3), or asks about evaluating this task. Reports adherenc...
This benchmark probes the safety judgment and alignment capabilities of professional-level AI agents. It evaluates whether agents can resist executing harmful or risky actions when given complex, domain-specific instructions in fields like finance, law, and healthcare. Use when the user wants to benchmark on SafePro, or asks about evaluating this task. Reports unsafe rate.
Evaluates the capability of classifiers to detect sexist, abusive, offensive, and hate speech in conversational text across multiple granularity levels and established benchmarks. It probes fine-grained toxicity detection, cross-dataset generalization, and performance against strong supervised and LLM baselines. Use when the user wants to benchmark on EDOS (SemEval 2023), OffensEval 2019, AbusEval, HatEval, or asks about evaluating this task. Reports F1.
Evaluates the safety-aware task planning capabilities of embodied LLM agents in interactive simulation environments. It probes whether agents can proactively reject hazardous instructions, avoid implicit risks in long-horizon planning, and maintain planning performance on safe tasks across varying levels of task abstraction. Use when the user wants to benchmark on SafeAgentBench, or asks about evaluating this task. Reports rejection rate.
This evaluation probes a model's ability to retain safety alignment and refusal capabilities while undergoing sequential continual domain adaptation across medical, legal, and coding tasks. It measures cumulative safety erosion and domain performance retention compared to unconstrained fine-tuning baselines. Use when the user wants to benchmark on HarmBench, TruthfulQA, BBQ, WildGuard, MedQA, LegalBench, CodeAlpaca, HumanEval, MMLU, or asks about evaluating this task. Reports Safety Score.
Evaluates the vulnerability of aligned multimodal LLMs to universal adversarial image attacks that bypass safety filters. It measures how often a single optimized image forces the model to generate unsafe or affirmative responses across diverse text prompts. Use when the user wants to benchmark on SafeBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates the ability of Large Vision-Language Models (LVLMs) to detect and refuse harmful visual content without modifying the base model architecture. It measures both safety defense effectiveness on toxic inputs and the preservation of utility on benign inputs. Use when the user wants to benchmark on Toxic Image Categories (Porn, Bloody, Insulting, Alcohol, Cigarette, Gun, Knife, Neutral), or asks about evaluating this task. Reports DSR.
Evaluates the biosafety risks and jailbreak vulnerabilities of protein foundation models by measuring their ability to reconstruct harmful protein sequences and 3D structures from partially masked inputs. It probes whether models can bypass safety filters and generate biologically dangerous proteins when given sequence and structural prompts. Use when the user wants to benchmark on SafeProtein-Bench, or asks about evaluating this task. Reports jailbreak success rate.
This evaluation probes an agent's ability to learn safe navigation and manipulation policies from human demonstrations in environments with unknown safety constraints. It specifically tests the trade-off between maximizing task reward and minimizing safety violations (cost) under out-of-distribution conditions. Use when the user wants to benchmark on Safety-Gymnasium (SafetyPointGoal1-v0, SafetyPointCircle2-v0, SafetyCarButton1-v0, SafetyCarPush2-v0), or asks about evaluating this task. Repor...
Evaluates LLM safety alignment and robustness against adversarial jailbreak attacks, particularly in scientific domains. It also measures the model's ability to maintain general helpfulness, truthfulness, and avoid over-refusal on benign queries. Use when the user wants to benchmark on AdvBench, HarmBench, StrongReject, SciKnowEval (L4), SciSafeEval, LabSafety Bench (Hard), GSM8K, MT-Bench, MMLU, GPQA, SimpleQA, XsTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates the safety alignment preservation and task utility of LLMs after parameter-efficient fine-tuning under data-poisoning attacks. It measures how well defenses maintain core capabilities while resisting harmful behavior injection. Use when the user wants to benchmark on MMLU, MT-Bench, AdvBench, PolicyEval, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates the safety alignment of Large Vision Language Models (LVLMs) against safety-awareness benchmarks and multimodal jailbreak attacks. It measures the model's ability to detect and mitigate harmful intents while preserving general multimodal reasoning capabilities. Use when the user wants to benchmark on MSSBench, SIUO, MM-SafetyBench, MML-M, FigStep, or asks about evaluating this task. Reports safety rate.
This evaluation probes a model's ability to retain task-specific utility while preserving safety alignment during supervised fine-tuning. It measures how well a method prevents safety degradation when exposed to benign or contaminated fine-tuning data, balancing performance retention against harmful output generation. Use when the user wants to benchmark on SST-2, AGNEWS, GSM8K, PubMedQA, AlpacaEval, JailbreakBench, HarmBench, AdvBench, BeaverTails, or asks about evaluating this task. Reports...
Evaluates the safety alignment of reasoning models by measuring how frequently they comply with harmful or jailbreak prompts across multiple risk categories. It also measures utility retention on standard mathematical and knowledge benchmarks to ensure safety improvements do not degrade general capabilities. Use when the user wants to benchmark on DAN, Wildjailbreak, StrongReject, GSM8K, MMLU, or asks about evaluating this task. Reports attack success rate.
Quantifies implicit representational harms in pre-trained language models by measuring the disparity in language modeling probabilities between harmful and benign sentences targeting 13 marginalized demographics. It probes whether a model's internal likelihood estimates reflect toxic or stereotypical biases toward specific groups. Use when the user has predictions and gold and needs to compute safety score.
This evaluation protocol probes the trade-off between safety alignment and reasoning capability in Large Reasoning Models. It measures how post-alignment fine-tuning impacts performance on standard reasoning benchmarks versus the model's propensity to generate harmful responses to malicious prompts. Use when the user wants to benchmark on GPQA, AIME24, MATH500, BeaverTails, or asks about evaluating this task. Reports Reasoning Accuracy.
Evaluates the effectiveness of a spectral geometric adversarial attack on 3D mesh autoencoders by measuring how well perturbed meshes deceive a downstream classifier and evade detection. Use when the user wants to benchmark on CoMA, SMAL, or asks about evaluating this task. Reports Targeted classification accuracy.
Evaluates automatic speech recognition (ASR) performance on the Oromo language using real-world, crowd-sourced audio data. It measures how well different model architectures (Conformer trained from scratch, Whisper fine-tuned) transcribe spoken Oromo into text under varying acoustic conditions. Use when the user wants to benchmark on Sagalee, or asks about evaluating this task. Reports WER.
Evaluates embodied vision-and-language navigation (VLN) capabilities within physically executable 3D Gaussian Splatting environments. It probes a model's ability to follow natural language instructions (high- and low-level), navigate to goals without collisions, and exhibit smooth, natural motion continuity rather than mechanical or wall-hugging behaviors. Use when the user wants to benchmark on SAGE-Bench, VLN-CE (R2R Val-Unseen), or asks about evaluating this task. Reports SR.
This benchmark evaluates speech large language models across five hierarchical levels of understanding, ranging from basic automatic speech recognition and language identification to paralinguistic perception (pitch, volume, emotion), abstract acoustic reasoning (medical cough analysis), and creative/agentic tasks (spoken English coaching). It probes the model's ability to process raw audio, follow instructions, and extract both semantic and non-semantic acoustic features. Use when the user w...
Evaluates Arabic language models on financial and Shari’ah-compliant reasoning across multiple task types, including multiple-choice questions, extractive summarization, and open-ended question answering. It probes the gap between general Arabic fluency and domain-specific procedural/financial reasoning. Use when the user wants to benchmark on SAHM, or asks about evaluating this task. Reports exact-match accuracy.
Compute saicharan2804/my_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of saicharan2804/my_metric.
Evaluates the scalability and performance trends of scientific AI workloads (3D CNNs) on HPC systems under varying node counts and dataset sizes. It probes how hardware constraints like GPU memory, I/O bandwidth, and network communication affect training efficiency and model convergence. Use when the user wants to benchmark on SAIH-cosmo, or asks about evaluating this task. Reports average_flops.
Evaluates large language models on language proficiency, reading comprehension, reasoning, and cultural understanding across three Southeast Asian languages (Indonesian, Vietnamese, Thai). It covers eight diverse tasks including question answering, machine translation, text summarization, multiple-choice exams, commonsense reasoning, machine reading comprehension, natural language inference, and sentiment analysis. Use when the user wants to benchmark on XQuAD, TyDiQA, Flores-200, ThaiSum, In...
Evaluates the safety, robustness, and helpfulness of Large Language Models across a hierarchical taxonomy of 6 domains, 16 tasks, and 66 categories. It measures performance on benign, adversarial (attack-enhanced), and defense-enhanced prompts, as well as multiple-choice safety questions, while also benchmarking the effectiveness of various attack and defense strategies. Use when the user wants to benchmark on SALAD-Bench, ToxicChat, Beavertails, SafeRLHF, Harmbench, Lifetox, AdvBench-50, or ...
Evaluates the quality of synthetically generated Text-to-SQL data by measuring question-SQL alignment and business realism in a sales analytics domain. Use when the user wants to benchmark on Salesforce Sales Analytics Database, or asks about evaluating this task. Reports Question-SQL Alignment (%).
Evaluates the ability of saliency prediction models to accurately identify visually salient regions and semantic objects in autonomous driving scenarios. It probes whether models can capture critical driving elements like pedestrians and approaching vehicles while mitigating center-bias and peripheral vision neglect. Use when the user wants to benchmark on BDD-A, DR(eye)VE, JAAD, or asks about evaluating this task. Reports D_KL.
Evaluates the alignment, chatbot capability, reasoning, coding, multilingual understanding, and truthfulness of the Dromedary-2 model using automatic LLM-as-a-judge scoring and standard benchmark accuracy metrics. Use when the user wants to benchmark on Vicuna-Bench, MT-Bench, AlpacaEval, Big Bench Hard (BBH), HumanEval, TydiQA, TruthfulQA, or asks about evaluating this task. Reports GPT-4-based automatic evaluation.
Probes whether tabular models can effectively leverage declarative business knowledge and metadata semantics for prediction tasks, rather than relying solely on statistical correlations in raw features. It evaluates the impact of schema-grounded semantic embeddings on model inductive biases and relative performance across different model families. Use when the user wants to benchmark on SALT-KG, or asks about evaluating this task. Reports ranking metrics.
Evaluates the effectiveness and overhead of fine-grained GPU sharing primitives for scheduling deep learning training, hyper-parameter tuning, and inference workloads on a single GPU. Use when the user wants to benchmark on Salus DL Workload Trace & Benchmarks, or asks about evaluating this task. Reports Makespan.
Evaluates audio source separation capabilities conditioned on text, visual masks, or temporal spans. It probes open-vocabulary extraction, speaker/music/instrument isolation, and cross-modal grounding in both studio and in-the-wild settings. Use when the user wants to benchmark on SAM Audio Evaluation Set, MUSDB18, or asks about evaluating this task. Reports separation fidelity.
Evaluates the zero-shot segmentation capability of the Segment Anything Model (SAM) across diverse medical imaging modalities and anatomical structures. It probes the model's ability to generalize without task-specific fine-tuning to structured medical targets like organs, lesions, and retinal layers. Use when the user wants to benchmark on Skin Lesion Analysis Toward Melanoma Detection (ISIC), DoFE (Drishiti-GS, RIM-ONE-r3, REFUGE amalgamated), AMOS, MICCAI 2017 Robotic Instrument Segmentati...
Detects and localizes semantically coordinated multimodal manipulations where visual edits are paired with contextually consistent textual narratives. Probes a model's ability to perform binary classification, multi-label categorization, and fine-grained visual tampering region localization using external celebrity attribute knowledge. Use when the user wants to benchmark on SAMM, or asks about evaluating this task. Reports ACC.
Evaluates zero-shot singing voice conversion by measuring timbre transfer accuracy and audio quality. It probes the model's ability to disentangle content, pitch, and timbre, and generalize to unseen human and non-human (animal) speakers without fine-tuning. Use when the user wants to benchmark on Custom zero-shot test set, or asks about evaluating this task. Reports MOS-S.
Evaluates sequential recommendation models on their ability to predict the next item in a user's chronological interaction history. It specifically probes whether sharpness-aware minimization improves generalization and data efficiency compared to standard Transformers and self-supervised baselines. Use when the user wants to benchmark on Amazon-Beauty, Amazon-Sports, Amazon-Toys, Yelp, or asks about evaluating this task. Reports HR@10.
Abstractive summarization of naturalistic, messenger-style dialogues. It probes a model's ability to generate concise, human-like summaries from informal, multi-turn conversations containing typos, emoticons, and multi-party interactions. Use when the user wants to benchmark on SAMSum Corpus, or asks about evaluating this task. Reports ROUGE F1 (ROUGE-1, ROUGE-2, ROUGE-L).
Evaluates whether an LLM's internal hidden layer activations can predict the veracity of a given statement. It probes the model's implicit knowledge of truthfulness by training a classifier on neural activations rather than relying on explicit prompting or output probabilities. Use when the user wants to benchmark on True-False Dataset, LLM-Generated Statements, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness and accuracy of OCR models on synthetic, book-style Arabic documents with high typographic diversity across 10 fonts. It measures character-level precision, word-level accuracy, and overall sequence fluency to benchmark vision-language and traditional OCR systems. Use when the user wants to benchmark on SARD, or asks about evaluating this task. Reports CER.
Evaluates a model's ability to generate scalable vector graphics (SVG) from text prompts and reference images, measuring visual fidelity, semantic alignment, structural success, and code efficiency. Use when the user wants to benchmark on SArena-Icon, or asks about evaluating this task. Reports SR.
Evaluates a model's ability to predict the next item in a user's interaction sequence based on historical behavior. It probes the model's capacity to capture long-range dependencies and adapt to varying data sparsity across different domains. Use when the user wants to benchmark on Amazon (Beauty), Amazon (Games), Steam, MovieLens-1M, or asks about evaluating this task. Reports Recall@K.