
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates an unsupervised transfer learning method's ability to jointly cluster cells and genomic features across two datasets with differing distributions. It probes the model's capacity to elastically transfer clustering knowledge from an auxiliary dataset to a target dataset based on data similarity, without requiring labeled data or matching cluster counts. Use when the user wants to benchmark on Simulated scATAC-seq & scRNA-seq, Real data 1 (Human scRNA & scATAC), Real data 2 (Human & Mo...
Evaluates a graph neural network's ability to classify Bitcoin transactions as licit or illicit using structural and narrative features, while testing a retrieval-augmented generation pipeline for producing regulatory-aligned explanations. Use when the user wants to benchmark on Elliptic AML dataset, or asks about evaluating this task. Reports F1-score.
Evaluates the utility, robustness, and interpretability of graph-derived signals for tabular machine learning on a binary node classification task. It compares graph-augmented models against tabular baselines using statistical hypothesis testing and graph perturbation analysis to ensure reproducibility. Use when the user wants to benchmark on Elliptic, or asks about evaluating this task. Reports F1-score.
Evaluates the robustness of Android malware detection models against feature-space and problem-space adversarial attacks, as well as temporal concept drift. It measures detection accuracy under increasing perturbation budgets while enforcing a strict false positive rate constraint on benign applications. Use when the user wants to benchmark on ELSA-RAMD Benchmark, or asks about evaluating this task. Reports TPR 100-FSA.
This evaluation probes a trained model's data memorization by auditing whether specific query images were included in its training set. It measures the ability to correctly distinguish memorized data from non-memorized data under varying calibration set qualities and query dataset constraints. Use when the user has predictions and gold and needs to compute auditing score (ρ_EMA / ρ_KS).
Evaluates a model's ability to generate highly abstractive, ultra-concise email subject lines from email body text. It probes extreme compression, informativeness, and fluency in a real-world email triaging context, distinguishing the task from standard text summarization. Use when the user wants to benchmark on AESLC, or asks about evaluating this task. Reports ROUGE-1.
Evaluates tiny language models and baselines on resource-constrained embedded devices by measuring performance across classification and regression tasks under strict memory limits (≤2 MB). It probes the trade-off between model compression, hardware compatibility, and NLP task accuracy. Use when the user wants to benchmark on TinyNLP, GLUE, or asks about evaluating this task. Reports Accuracy, GLUE Average Score.
Evaluates the conversational quality, memory retrieval, personalization, and long-term memory extraction capabilities of an edge-deployed AI companion system over simulated multi-session interactions. Use when the user wants to benchmark on Synthetic User Simulation, or asks about evaluating this task. Reports Conversation Quality.
This evaluation probes how well different embedding models capture underlying number-theoretic structures in numeric sequences. It measures the quality of the latent representation space by comparing clustering performance against ground-truth mathematical group labels versus unsupervised KMeans assignments. Use when the user wants to benchmark on number-theoretic-sequences, or asks about evaluating this task. Reports Silhouette Coefficient.
Evaluates the quality of text embeddings across diverse tasks including retrieval, classification, clustering, and semantic similarity. It probes multilingual, cross-lingual, and code understanding capabilities, measuring how well dense vector representations capture semantic relationships for downstream applications. Use when the user wants to benchmark on MTEB (Massive Text Embedding Benchmark), XOR-Retrieve, XTREME-UP, or asks about evaluating this task. Reports MTEB Task Mean.
Evaluates machine learning models on static malware classification of Windows PE binaries. It probes the effectiveness of engineered static features versus raw binary inputs for distinguishing malicious from benign software. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports ROC AUC.
Evaluates the performance and robustness of various machine learning classifiers for static malware detection on high-dimensional tabular features, specifically assessing the impact of dimensionality reduction techniques (PCA, LDA) on classification accuracy and discriminative power. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.
Evaluates a multi-stage machine learning pipeline for detecting and classifying Windows PE files using static analysis features. It probes the model's ability to perform binary malware detection, hierarchical threat-type classification, family identification, and behavioral categorization. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.
Evaluates malware classifiers on detection, family identification, and attribute prediction tasks across multiple file formats. It specifically probes robustness against concept drift and novel malware families using a temporally separated test set and a challenge set of evasive samples. Use when the user wants to benchmark on EMBER2024, or asks about evaluating this task. Reports ROC AUC.
Evaluates embodied AI agents' ability to navigate to target objects in 3D environments and perform manipulation tasks. It probes spatial reasoning, path efficiency, trajectory smoothness, and zero-shot generalization across different simulation domains and visual styles. Use when the user wants to benchmark on ProcTHOR-10k, ArchitecTHOR, AI2-iTHOR, RoboTHOR, ManipulaTHOR, Habitat 2022 ObjectNav, or asks about evaluating this task. Reports SR (Success Rate).
This benchmark suite evaluates embodied AI models across perception, spatial reasoning, navigation, and task planning. It aggregates 22 diverse benchmarks to measure capabilities like 2D/3D question answering, instruction following in navigation, and complex task decomposition. Use when the user wants to benchmark on Embodied Arena, or asks about evaluating this task. Reports Exact Matching Accuracy.
Evaluates embodied agent systems on runtime governance, safety, and accountability across seven dimensions: unauthorized capability invocation, runtime drift robustness, recovery success, policy portability, version upgrade safety, human override latency, and audit completeness. It shifts focus from pure task success to controllability and policy compliance under real-world perturbations. Use when the user wants to benchmark on EmbodiedGovBench, or asks about evaluating this task. Reports una...
This protocol evaluates the safety and navigation performance of embodied agents against physical and model-based attacks. It measures task completion efficiency, path optimality, and goal satisfaction across diverse simulated and real-world environments. Use when the user wants to benchmark on Li et al. (2023), Kim et al. (2024), Khanna et al. (2024), Yin et al. (2024), Wang et al. (2024b), or asks about evaluating this task. Reports Success weighted by Path Length (SPL).
Evaluates an agent's ability to perform long-horizon embodied interactive tasks by synergizing visual search, reasoning, and action. It probes spatial reasoning, self-reflection, and planning capabilities in both simulated and real-world environments. Use when the user wants to benchmark on Unspecified (Simulated & Real-world tasks), or asks about evaluating this task. Reports success_rate.
Evaluates an embodied AI model's capabilities in general multimodal reasoning, 3D spatial perception, and long-horizon task planning across multiple public benchmarks and a custom simulation environment. Use when the user wants to benchmark on MM-IFEval, MMStar, MMMU, AI2D, OCRBench, BLINK, CV-Bench, EmbSpatial, ERQA, EgoPlan, EgoPlan2, EgoThink, Internal Planning, VLM-PlanSim-99, or asks about evaluating this task. Reports Action Pair Match F1-Score.
Evaluates the efficiency and executability of conversational workflow automation for embodied AI development tasks, including environment synthesis, trajectory collection, and VLA model evaluation. Use when the user wants to benchmark on RoboTwin, or asks about evaluating this task. Reports average task completion time.
Evaluates how image compression codecs impact the performance of Vision-Language-Action (VLA) models in closed-loop robotic manipulation tasks under ultra-low bitrates. It measures whether compressed visual inputs cause task failure or require excessive inference steps, highlighting the disconnect between traditional visual fidelity metrics and embodied AI operational requirements. Use when the user wants to benchmark on EmbodiedComp, or asks about evaluating this task. Reports Success Rate (...
This evaluation probes the model's ability to generate executable sub-goal plans from visual inputs and translate them into low-level control actions in simulated robotic environments. It specifically tests closed-loop planning and few-shot policy adaptation across standard embodied AI benchmarks. Use when the user wants to benchmark on Franka Kitchen, Meta-World, or asks about evaluating this task. Reports success rate.
Evaluates LLMs' physical safety and decision-making capabilities in embodied contexts by testing their ability to refuse unsafe instructions, interpret safety-critical goals, model state transitions, and sequence actions correctly. Use when the user wants to benchmark on EmbodyGuard, or asks about evaluating this task. Reports recall.
Evaluates embodied vision-language models on step-wise, instruction-driven navigation and interaction in complex, partially observable environments. It probes three core capabilities: active exploration, dynamic spatial-semantic reasoning, and multi-stage goal execution under closed-loop perception-action constraints. Use when the user wants to benchmark on EmbRACE-3K, or asks about evaluating this task. Reports success.
Evaluates large vision-language models' ability to reason about egocentric spatial relations (e.g., above, below, left, right, close, far) within 3D embodied environments. It probes whether models can accurately localize objects and identify spatial configurations from a first-person perspective. Use when the user wants to benchmark on Embspacial-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of information extraction models to update knowledge graphs with emerging textual knowledge. It probes capabilities in extracting existing and new triples, linking emerging entities, and deprecating obsolete relations based on temporal text evidence. Use when the user wants to benchmark on EMERGE, or asks about evaluating this task. Reports recall.
Evaluates text-to-speech models on six complex linguistic and prosodic dimensions: emotions, paralinguistics, syntactic complexity, questions, foreign words, and complex pronunciation. It uses a model-as-a-judge framework to measure pairwise preference against a baseline, alongside automated speech quality metrics. Use when the user wants to benchmark on EmergentTTS-Eval, or asks about evaluating this task. Reports win-rate.
Evaluates a multimodal emotion recognition system that fuses facial, posture, and gait cues with situational knowledge to classify emotional states. It probes the model's ability to generalize across posed and wild settings, different data modalities, and subject-independent splits. Use when the user wants to benchmark on FER-2013, CAER-S, FABO, EWalk, GroupWalk, GEMEP, or asks about evaluating this task. Reports Accuracy.
Human-subject validation of cross-modal emotional alignment between music and images. It probes whether pairing audio and visual stimuli based on a 13-dimensional emotional coordinate space yields higher perceptual matching accuracy compared to semantic-only or random matching. Use when the user wants to benchmark on EMID, or asks about evaluating this task. Reports accuracy.
Evaluates the effectiveness of the Emilia dataset for Text-to-Speech generation by comparing models trained on Emilia versus MLS. It probes intelligibility, speaker similarity, and naturalness across formal and spontaneous speaking styles in both English and multilingual settings. Use when the user wants to benchmark on LibriSpeech-Test, Emilia-Test, Aishell-3, Common Voice, or asks about evaluating this task. Reports WER.
Evaluates massively multilingual language models on intrinsic next-word prediction, machine translation, text classification, math reasoning, and code generation across dozens of languages, with a specific focus on low-resource language performance and cross-lingual transfer capabilities. Use when the user wants to benchmark on Glot500-c, Parallel Bible Corpus (PBC), FLORES-200, SIB-200, Taxi-1500, MGSM, or asks about evaluating this task. Reports BLEU.
This benchmark evaluates multimodal large language models on their ability to perform integrated visual-textual reasoning across mathematics, physics, chemistry, and coding. It probes capabilities such as fine-grained spatial simulation, multi-hop visual inference, and cross-modal problem solving under both direct and chain-of-thought prompting conditions. Use when the user wants to benchmark on EMMA-mini, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates embodied mobile manipulation agents in open environments, testing their ability to execute long-horizon, language-conditioned tasks that require interleaved high-level planning and low-level continuous navigation/manipulation. It specifically probes reasoning fidelity, execution success, adaptability to failures, and path efficiency compared to expert trajectories. Use when the user wants to benchmark on EMMOE-100, or asks about evaluating this task. Reports PLWSR.
Evaluates a model's ability to recognize and classify handwritten characters (digits and letters) from standardized 28x28 grayscale images. It probes robustness to case variations, class overlap, and imbalanced distributions across multiple dataset configurations. Use when the user wants to benchmark on EMNIST, or asks about evaluating this task. Reports classification accuracy.
Evaluates machine learning models' ability to classify brief, nonverbal vocal bursts into discrete emotional states. It probes the limits of audio classification on highly ambiguous, short-duration human vocalizations across 30 affective categories. Use when the user wants to benchmark on EmoGator, or asks about evaluating this task. Reports F1 score.
Probes a model's ability to perform context-aware, multi-dimensional emotion understanding by predicting Plutchik’s 8 basic emotions from rich textual scenarios. It specifically tests zero-shot multi-label emotion prediction and evaluates whether models can capture emotional entanglement (co-occurrence) rather than treating emotion dimensions as independent. Use when the user wants to benchmark on EmoScene, or asks about evaluating this task. Reports Macro F1.
Evaluates vision models' ability to classify emotions in images across multiple affective categories. It also measures the alignment between emotions expressed in text prompts and those visually present in generated images. Use when the user wants to benchmark on EmoSet, or asks about evaluating this task. Reports macro-averaged F1-score.
Evaluates how emotional stimuli (joy, encouragement, anger, insecurity) and their intensity affect LLM behavior across factual accuracy, sycophancy, and toxicity. It measures the performance delta when emotional prompt add-ons are applied to base prompts. Use when the user wants to benchmark on Anthropic’s SycophancyEval subset, Toxicity dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates large language models' emotional intelligence and empathy by testing their ability to recognize key events, mixed events, implicit emotions, and user intent from real-world emotional scenarios, and to generate appropriate empathetic responses. Use when the user wants to benchmark on EmotionQueen, or asks about evaluating this task. Reports PASS rate.
Evaluates multimodal emotion recognition and sentiment analysis capabilities across unimodal and fused modalities, alongside emotional speaker style captioning. It probes how well models capture discrete emotions, continuous sentiment, and fine-grained speaking styles from Chinese dyadic dialogues. Use when the user wants to benchmark on EmotionTalk, or asks about evaluating this task. Reports ACC.
Evaluates a model's ability to recognize and aggregate facial emotions across multiple individuals in a crowd scene. It probes the system's capacity to capture complex spatial dependencies among overlapping facial expressions to predict the dominant group-level emotion. Use when the user wants to benchmark on EmotiW2018, GECV, or asks about evaluating this task. Reports mean accuracy (mAC).
Evaluates multimodal LLMs on understanding, reasoning, and predicting fine-grained emotion transitions in short video clips. It probes capabilities across four progressive tasks: detecting whether an emotion change occurs, identifying before/after emotion states, generating evidence-grounded reasoning for transitions, and predicting the next emotion state. Use when the user wants to benchmark on EmoTrans, or asks about evaluating this task. Reports Accuracy.
Evaluates an omni-modal language model's ability to engage in end-to-end spoken dialogue with vivid emotional control. It probes the model's dialogue quality, text generation accuracy under different input modalities, and its capability to control and classify speech styles/emotions. Use when the user wants to benchmark on EMOVA-EmotionDialogue-Test, or asks about evaluating this task. Reports end-to-end spoken dialogue score.
Evaluates an LLM-based text-to-speech model's ability to generate emotionally expressive speech guided by natural language prompts, measuring content accuracy, emotional fidelity, and audio naturalness. Use when the user wants to benchmark on EmoVoice-DB, Secap, or asks about evaluating this task. Reports WER.
This benchmark evaluates a model's ability to generate or retrieve empathetic, relevant, and fluent responses to emotionally grounded personal stories. It probes how well dialogue systems can acknowledge and react to a speaker's feelings in open-domain conversations. Use when the user wants to benchmark on EmpatheticDialogues, or asks about evaluating this task. Reports Human Empathy/Relevance/Fluency.
Evaluates an LLM-based agent's ability to automate cohort selection, feature extraction, and clinical code standardization across heterogeneous Electronic Medical Record (EMR) databases. It probes schema-aware SQL reasoning, iterative query planning, and robustness to unseen database structures without manual rule engineering. Use when the user wants to benchmark on MIMIC-III, eICU, SICdb, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to map clinical questions to structured logical forms and to extract precise answer spans or predict answer classes from unstructured electronic medical records. It probes complex clinical reasoning, including temporal, arithmetic, and multi-sentence contextual understanding. Use when the user wants to benchmark on emrQA, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates large language models and retrieval-augmented generation systems on emergency medical services (EMS) multiple-choice questions across different clinical subject areas and certification levels. It probes the models' ability to apply domain-specific expertise and reasoning to answer standardized medical certification questions. Use when the user wants to benchmark on EMSQA, or asks about evaluating this task. Reports exact-match accuracy (Acc).
Evaluates autonomous driving perception and prediction capabilities, specifically multi-agent object tracking and trajectory forecasting, using a dataset collected in the UAE with diverse driving scenarios. Use when the user wants to benchmark on EMT, or asks about evaluating this task. Reports MOTA.