
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the safety and refusal capabilities of language models across multiple risk categories including harmful content generation, jailbreaking, stereotype bias, privacy leaks, and toxicity detection. It measures how well models balance harmlessness (refusing unsafe prompts) and helpfulness (complying with safe prompts) in both single- and multi-turn interactions. Use when the user wants to benchmark on XSTest, DecodingTrust, ToxiGen, XSafety, RTP-LX, Microsoft Internal Automated Measurem...
Evaluates diverse reasoning and coding capabilities using an internal benchmark designed to minimize data contamination and LLM-judge bias. It probes a model's ability to debug, extend, and explain code, as well as identify errors in mathematical proofs and generate related problems. Use when the user wants to benchmark on PhiBench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates machine learning classifiers and feature selection strategies for detecting phishing websites. It probes the ability of models to distinguish between legitimate and malicious web pages using content-based, external service, and hybrid feature sets, while measuring classification accuracy and macro F1-score. Use when the user wants to benchmark on Collected Phishing Dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates the security and robustness of autonomous LLM email agents against phishing attacks by measuring how different system prompt configurations affect detection sensitivity and operational false positive rates. It specifically probes the model's ability to maintain high recall while minimizing usability costs, and tests adversarial brittleness under infrastructure phishing conditions where attacker-controlled domains match sender addresses. Use when the user wants to benchmark on Synthe...
Evaluates category-level 6D object pose estimation on photometrically challenging objects (reflective, transparent, occluded). It tests both in-distribution generalization (seen objects) and out-of-distribution generalization (novel objects within the same category), comparing RGB-D and monocular approaches. Use when the user wants to benchmark on PhoCaL, or asks about evaluating this task. Reports 3D IoU.
This benchmark evaluates Vietnamese-English machine translation quality by comparing neural baselines and commercial engines. It probes translation accuracy across multiple domains and sentence lengths using both automatic metrics and human preference judgments. Use when the user wants to benchmark on PhoMT, or asks about evaluating this task. Reports BLEU.
Evaluates the inference speed, energy efficiency, and classification accuracy of a GPU-accelerated binary neural network engine on mobile devices against standard frameworks. Use when the user has predictions and gold and needs to compute runtime.
Evaluates automatic speech recognition performance on two low-resource, phonologically complex endangered languages (Archi and Kina Rutul) at the word, character, and phoneme levels. It specifically probes how training data frequency impacts phoneme recognition accuracy and error types. Use when the user wants to benchmark on Archi & Kina Rutul ASR, or asks about evaluating this task. Reports PER.
Compute phonemetransformers/segmentation_scores via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of phonemetransformers/segmentation_scores.
This benchmark evaluates foundation models' ability to solve Olympiad-level physics problems using retrieval-augmented generation (RAG). It probes step-wise logical reasoning, correct application of physical laws, and robustness to noisy retrieved context. Performance is measured via an LLM-as-judge scoring framework that rewards both intermediate reasoning quality and final answer correctness. Use when the user wants to benchmark on PhoPile, or asks about evaluating this task. Reports Averag...
Evaluates personalized, intent-driven photo retrieval capabilities that go beyond simple visual matching. It tests a system's ability to fuse multi-source constraints (temporal, spatial, social identity) and correctly abstain when no relevant image exists in a personal album. Use when the user wants to benchmark on PhotoBench, or asks about evaluating this task. Reports Recall@K.
Evaluates a 1D photochemical and climate model's ability to simulate atmospheric composition, photochemical networks, and radiative energy balance across diverse planetary environments. It probes whether the model can reproduce observed vertical gas profiles, cloud properties, and thermal structures without relying on unphysical surface fluxes. Use when the user wants to benchmark on Planetary Atmospheric Observations (Venus, Earth, Mars, Titan, Jupiter, WASP-39b), or asks about evaluating th...
Evaluates the accuracy of predicting galaxy photometric redshifts from optical (grizy) photometric data across multiple redshift ranges. It probes a model's ability to minimize systematic bias, reduce catastrophic outliers, and maintain low error rates under varying data distributions. Use when the user wants to benchmark on Hyper Suprime-Cam Photometric Redshift Data, or asks about evaluating this task. Reports MAE.
Classifies the sentiment polarity of specific phrases within tweets, requiring models to handle contextual disambiguation, sarcasm, and informal language. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
Evaluates phishing website detection models on temporally disjoint, real-world data with realistic base rates, while mitigating training-to-test leakage and varying difficulty levels. Use when the user wants to benchmark on PhreshPhish, or asks about evaluating this task. Reports Precision-Recall.
Compute phucdev/blanc_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of phucdev/blanc_score.
Compute phucdev/vihsd via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of phucdev/vihsd.
Evaluates a humanoid robot's ability to imitate human motion and follow pelvis trajectories using physically-grounded retargeting. It probes full-body tracking accuracy and partial-state path-following control across diverse locomotion categories on Unitree G1 and H1-2 robots. Use when the user wants to benchmark on PHUMA, Unseen Video, or asks about evaluating this task. Reports success_rate.
Evaluates multimodal models' ability to perform physical reasoning, spatial cognition, and egocentric task planning, as well as their capacity to act as reliable critics/judges for physical AI tasks. Use when the user wants to benchmark on PhyCritic-Bench, VL-RewardBench, Multimodal-RewardBench, CosmosReason1-Bench, CV-Bench, EgoPlanBench2, or asks about evaluating this task. Reports accuracy (overall/macro).
Evaluates an agent's ability to reason about 2D Newtonian physics to solve goal-driven puzzles by placing dynamic objects. It probes sample-efficient learning and generalization across unseen task templates and action spaces. Use when the user wants to benchmark on PHYRE, or asks about evaluating this task. Reports AUCCESSION.
This benchmark evaluates a model's ability to reason about physical laws and detect physical commonsense violations in gameplay videos. It probes spatial, temporal, and meta-information-based physical reasoning through curated multi-choice questions. Use when the user wants to benchmark on PhysGame, or asks about evaluating this task. Reports accuracy.
Evaluates a model's capacity for metric spatial reasoning, object enumeration, relative spatial comparisons, and topological/directional relationship understanding within real-world warehouse environments. Use when the user wants to benchmark on PhysicalAI-Spatial-Intelligence-Warehouse, or asks about evaluating this task. Reports normalized exact-match accuracy.
Evaluates the utility of physically disentangled scene representations (geometry, albedo, lighting, camera) for downstream vision tasks. Probes robustness to out-of-distribution lighting and viewpoints, and measures how well learned features transfer to clustering, linear classification, and segmentation benchmarks. Use when the user wants to benchmark on CelebA, Buffy, BBT, RAF-DB, CelebA Mask, ShapeNet Cars, or asks about evaluating this task. Reports clustering_accuracy.
Evaluates the ability of models to forecast irregularly sampled multivariate time series generated from biological ordinary differential equations. It probes how well different architectures handle sparse, non-uniform temporal observations and varying levels of dynamical complexity. Use when the user wants to benchmark on Physiome-ODE, or asks about evaluating this task. Reports MSE.
Evaluates a model's ability to perform multilabel classification of cardiac abnormalities from 12-lead ECG recordings. It specifically probes robustness to severe class imbalance and the capacity to handle diverse lead configurations and sampling rates in long-sequence cardiac signals. Use when the user wants to benchmark on PhysioNet/CinC Challenge 2021, or asks about evaluating this task. Reports Macro AUPRC.
Probes the ability of multimodal large language models to solve undergraduate-level physics problems that require integrating textual descriptions with complex diagrams. It specifically tests multi-step scientific reasoning, mathematical derivation, and conceptual understanding across eight distinct physics sub-disciplines. Use when the user wants to benchmark on PhysUniBench, or asks about evaluating this task. Reports accuracy.
Probes multimodal physical reasoning by requiring models to interpret realistic visual scenarios, understand implicit physical conditions, and apply domain-specific knowledge across six physics domains. It evaluates both visual grounding and the ability to integrate symbolic reasoning with real-world constraints. Use when the user wants to benchmark on PhyX, or asks about evaluating this task. Reports accuracy.
Evaluates the capability of physics-informed autoregressive models to accurately forecast time-dependent partial differential equations and global atmospheric variables over multi-step horizons. Use when the user wants to benchmark on PDE Benchmarks (Wave, Reaction, Convection, Heat), ERA5, or asks about evaluating this task. Reports RMSE.
Evaluates the robustness of LLM prompt injection defenses against diverse attack strategies (heuristic, direct, adaptive, optimization-based) across multiple tasks and benchmarks. It measures the trade-off between maintaining legitimate task utility and preventing the execution of malicious injected instructions. Use when the user wants to benchmark on SQuAD v2, Dolly, NQ, InjecAgent, AgentDojo, AgentDyn, WASP, OPI, SEP, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates a Kubernetes-based multi-model orchestration framework for self-hosted LLMs, measuring how hybrid routing and adaptive scaling affect inference reliability, latency, GPU utilization, and cost across diverse benchmarks. Use when the user wants to benchmark on HumanEval, GSM8K, MBPP, TruthfulQA, ARC, HellaSwag, MATH, MMLU Pro, or asks about evaluating this task. Reports success.
Compute pico-lm/blimp via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of pico-lm/blimp.
Compute pico-lm/perplexity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of pico-lm/perplexity.
Evaluates real-time performance and resource efficiency of TinyML models on embedded hardware by measuring inference latency, CPU/memory utilization, and prediction confidence across different platforms. Use when the user wants to benchmark on Gesture Classification, Keyword Spotting, MobileNet V2, or asks about evaluating this task. Reports Average Inference Latency (ms).
Evaluates the ability of PII masking models to correctly identify and classify sensitive information in text. It probes performance across diverse contexts, multilingual inputs, noisy formats, and evolving entity types. Use when the user wants to benchmark on PII Masking Dataset, or asks about evaluating this task. Reports non-identification.
Evaluates a model's ability to identify and extract private or sensitive information spans from text across legal, healthcare, and finance domains. It probes the model's capacity for fine-grained named entity recognition under limited labeled data and domain-specific privacy definitions. Use when the user wants to benchmark on ECHR, MACCROBAT, PUPA (Finance Subset), or asks about evaluating this task. Reports F1.
This benchmark probes the ability of NER and PII detection systems to accurately identify and classify personally identifiable information spans across highly heterogeneous, cross-domain text sources. It specifically evaluates cross-domain generalization and robustness to diverse, fine-grained PII entity types that are rarely seen together in standard training corpora. Use when the user wants to benchmark on PIIBench, or asks about evaluating this task. Reports span-level F1.
This evaluation probes a language model's ability to capture statistical patterns in diverse English text domains and its cross-domain generalization. It measures next-token prediction accuracy across academic, technical, legal, and conversational corpora. Use when the user wants to benchmark on The Pile, or asks about evaluating this task. Reports perplexity (BPB).
Evaluates the sharpness and calibration of probabilistic net-load forecasts. It measures how closely predicted quantiles align with actual observations and how narrow the prediction intervals are while maintaining statistical reliability. Use when the user has predictions and gold and needs to compute Pinball Score.
Evaluates a model's ability to predict fine-grained hierarchical labels for tokens within identified P, I, and O spans. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
Evaluates a model's ability to identify text spans corresponding to Patient, Intervention, and Outcome elements within clinical trial abstracts. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
Evaluates a model's ability to reason about physical commonsense and intuitive physics by selecting the correct solution for a given goal from two options. It probes understanding of object affordances, material properties, and non-prototypical uses of everyday items. Use when the user wants to benchmark on PIQA, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates multimodal large language models on proactive intent recommendation from continuous, noisy GUI visual streams. It probes the model's ability to track long-horizon user behavior, distinguish true latent goals from background noise, and exercise operational restraint by remaining silent when no action is required. Use when the user wants to benchmark on PIRA-Bench, or asks about evaluating this task. Reports S_final.
Evaluates video anomaly detection and understanding on synthetic, long-form videos with diverse scenes and balanced anomaly categories. Probes models' temporal consistency, ability to detect subtle behavioral anomalies, and long-context narrative comprehension. Use when the user wants to benchmark on Pistachio, or asks about evaluating this task. Reports frame-level AUC.
Evaluates the capability of 3D face reconstruction models to predict accurate 3D meshes and facial landmarks from 2D RGB images. It probes how well models generalize to high-resolution, in-the-wild faces across diverse ages and expressions, highlighting domain gaps from synthetic training data. Use when the user wants to benchmark on Pixel-Face, or asks about evaluating this task. Reports ARMSE.
Evaluates multimodal models' ability to perform fine-grained visual reasoning, object counting, temporal video understanding, and complex infographic parsing. It specifically probes whether models can effectively leverage pixel-space operations (e.g., zooming, frame selection) rather than defaulting to text-only reasoning pathways. Use when the user wants to benchmark on V* (V-Star), TallyQA, MVBench, InfographicVQA, or asks about evaluating this task. Reports Acc.
This benchmark evaluates single-image 3D face reconstruction pipelines on two tasks: posed reconstruction (measuring geometric fidelity under diverse expressions) and neutral reconstruction (testing the ability to disentangle identity shape from expression). It probes a model's capacity to recover accurate facial geometry and surface normals from a single input image. Use when the user wants to benchmark on Pixel3DMM Benchmark, or asks about evaluating this task. Reports Chamfer Distance (L1/...
Evaluates the ability of recommender systems to rank items using raw pixel images instead of traditional ID embeddings. It probes cold-start item recommendation, cross-domain transfer learning, and end-to-end vision-based recommendation performance. Use when the user wants to benchmark on PixelRec, or asks about evaluating this task. Reports Recall@N.
This evaluation probes the ability of hybrid quantum-classical models to accurately predict residue-level pKa values across diverse protein microenvironments. It tests whether entanglement-aware quantum feature mappings generalize beyond the training distribution to experimental datasets and capture subtle electronic and geometric correlations in flexible peptide regions. Use when the user wants to benchmark on PKAD-R, Aβ40, or asks about evaluating this task. Reports RMSE.
Evaluates the PheKnowLator ecosystem's software features and computational performance against other biomedical KG construction tools, and measures construction efficiency on 12 benchmark knowledge graphs. Use when the user has predictions and gold and needs to compute coverage score.
Probes an LLM's ability to generate safe and helpful responses by classifying harmful content across 19 distinct categories and 3 severity levels, while aligning with human preference rankings on Q-A-B triplets. Use when the user wants to benchmark on PKU-SafeRLHF, or asks about evaluating this task. Reports harm_category.