
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates LLM agents' ability to solve mathematical reasoning and code generation tasks by leveraging a shared, contrastively distilled memory system. It probes cross-agent knowledge transfer, reasoning invariance extraction, and task-aware memory retrieval efficiency. Use when the user wants to benchmark on MATH500, GSM8K, MBPP, HumanEval, or asks about evaluating this task. Reports Accuracy (%).
Evaluates multimodal vision-language models on understanding memes across multiple languages and semantic categories. It probes cross-modal reasoning, cross-lingual transfer, and the ability to generalize across diverse tasks like harm detection, misinformation, and humor/sarcasm. Use when the user wants to benchmark on MemeLens Unified Benchmark, or asks about evaluating this task. Reports Macro-F1.
Probes a model's ability to generate actionable, natural-language feedback to improve the memorability of a photograph at capture time. It evaluates both the effectiveness of the feedback in increasing image memorability and the linguistic coherence of the suggestions. Use when the user wants to benchmark on MemBench, or asks about evaluating this task. Reports IR.
Compute the MemorizationInformedFrechetInceptionDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MemorizationInformedFrechetInceptionDistance, or asks how to score with MemorizationInformedFrechetInceptionDistance.
Evaluates four core memory competencies in LLM agents: accurate retrieval, test-time learning, long-range understanding, and selective forgetting. It transforms long-context datasets into session-based multi-turn interactions to simulate real-world memory accumulation and retrieval. Use when the user wants to benchmark on MemoryAgentBench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates how well LLM-based systems retain and utilize both declarative and procedural memory across diverse domains and task formats. It specifically probes continual learning capabilities by measuring performance improvements when systems process explicit and implicit user feedback over multiple interaction sessions. Use when the user wants to benchmark on MemoryBench (Domain & Task Format Partitions), or asks about evaluating this task. Reports LLM-as-Judge score.
Evaluates long-horizon robotic manipulation capabilities of vision-language-action models under non-Markovian dynamics. It probes the model's ability to maintain and retrieve perceptual and semantic memory over extended task horizons using only third-person visual observations and language instructions. Use when the user wants to benchmark on SimplerEnv-Bridge, SimplerEnv-Fractal, LIBERO, Real-world Manipulation, or asks about evaluating this task. Reports success rate.
Evaluates models on classifying social media memes for sentiment, emotion intensity (humour, sarcasm, offensiveness), and motivation. It probes the ability of text-only and multi-modal architectures to perform both binary and ordinal/multi-class classification across multiple related subtasks. Use when the user wants to benchmark on Memotion 2.0, or asks about evaluating this task. Reports weighted F1.
Evaluates multimodal understanding of internet memes by classifying five emotion categories (humor, sarcasm, offense, motivation) and overall sentiment from combined image and text inputs. Use when the user wants to benchmark on Memotion Analysis Dataset, or asks about evaluating this task. Reports F1 score.
Evaluates AudioLLMs on multi-turn, persona-conditioned spoken dialogue generation. It probes the model's ability to maintain speaker consistency, track conversation context, and generate contextually appropriate text responses to audio inputs in Arabic (MSA) and English. Use when the user wants to benchmark on MENA SpeechBank, or asks about evaluating this task. Reports Average Rubric Score (ARS).
Evaluates the reliability of machine translation evaluation metrics across reference-based, quality estimation, and LLM-as-a-judge paradigms when applied to non-literal content such as internet slang, idioms, and literary expressions. Use when the user wants to benchmark on MENT, or asks about evaluating this task. Reports Composite Meta Score.
Evaluates the safety, clinical adherence, and crisis response quality of LLM-powered mental health chatbots against expert-defined guidelines. It probes the model's ability to provide evidence-based advice, identify health risks, maintain consistent crisis intervention, provide appropriate resources, and empower users. Use when the user wants to benchmark on Institute for Future Health Mental Health Query Set, or asks about evaluating this task. Reports TotalScore.
Evaluates large language models' ability to perform psychiatric diagnostic decision-making using DSM-5 criteria. It probes their capacity to handle information incompleteness, perform differential diagnosis among overlapping disorders, and calibrate diagnostic commitment under varying prompt constraints. Use when the user wants to benchmark on MentalBench, or asks about evaluating this task. Reports accuracy (exact match).
Evaluates LLMs' capacity to generate empathetic, relevant, and contextually appropriate responses in mental health counseling scenarios. It probes the model's ability to demonstrate active listening, emotional validation, safety awareness, and ethical boundary adherence through a multi-dimensional rubric. Use when the user wants to benchmark on MentalChat16K, or asks about evaluating this task. Reports MentalHealth Counseling Metrics (7 dimensions).
Evaluates a model's ability to detect all types of entity mentions (named and general) in both clean written text and noisy spoken/transcribed speech. It specifically probes handling of ambiguous, nested, and context-dependent terms across different data modalities. Use when the user wants to benchmark on Wikipedia, Transcribed Speech, ASR Output, or asks about evaluating this task. Reports F1-measure.
Evaluates large language models on Russian-language instruction following across 21 tasks spanning 11 skill domains, including problem-solving, exam-based questions, and ethical diagnostics. It probes zero-shot and few-shot capabilities under strict black-box conditions to measure alignment with human performance and prevent data leakage. Use when the user wants to benchmark on MERA, or asks about evaluating this task. Reports accuracy.
Evaluates document understanding models on token classification tasks using structured school transcripts. It probes the model's ability to accurately label tokens based on layout and text features across English and Spanish documents with varying templates and layouts. Use when the user wants to benchmark on MERIT, or asks about evaluating this task. Reports Token Classification.
This evaluation probes the effectiveness of automated, quality-filtered 3D vehicle datasets for text-to-3D generative modeling. It measures how fine-tuning a base model on curated meshes improves multi-view consistency and perceptual alignment compared to caption- or aesthetic-score-based filtering. Use when the user wants to benchmark on MeshFleet, CarCaption3K, CarCaption800, or asks about evaluating this task. Reports CLIP-S.
Determines the overall sentiment polarity of an entire tweet, addressing class imbalance, slang, and informal social media text. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
Evaluates information retrieval models on a large-scale, dialectally diverse Spanish dataset. It probes the ability of lexical and dense retrieval models to rank relevant Wikipedia documents for real-world Spanish search queries without fine-tuning. Use when the user wants to benchmark on MessIRve, or asks about evaluating this task. Reports nDCG@10.
Evaluates few-shot image classification performance across 40 diverse datasets spanning 10 domains. It probes a model's ability to adapt to new classes with limited labeled examples (1, 5, 10, 20-shot) in within-domain settings. Use when the user wants to benchmark on Meta-Album, or asks about evaluating this task. Reports average accuracy.
Evaluates the efficiency and representational fidelity of LLM inference benchmarking methodologies by quantifying how well a reduced set of experimental parameters can accurately predict system performance compared to exhaustive testing. It measures the trade-off between computational cost and the accuracy of projected latency and throughput metrics. Use when the user has predictions and gold and needs to compute efficiency_metric.
Evaluates few-shot meta-learners and transfer learning baselines on their ability to generalize across heterogeneous vision tasks including classification, semantic segmentation, keypoint localization, and regression. It specifically probes cross-task knowledge transfer, in-distribution versus out-of-distribution robustness, and the comparative effectiveness of single-task versus multi-task meta-training protocols. Use when the user wants to benchmark on Meta Omnium, or asks about evaluating ...
Evaluates the downstream performance of language models pre-trained on data selected by various quality-based methods compared to random sampling. It probes how different data curation strategies impact general knowledge, commonsense reasoning, and reading comprehension capabilities. Use when the user wants to benchmark on ARC-Challenge, ARC-Easy, SciQ, HellaSwag, SIQA, WinoGrande, RACE, OpenbookQA, or asks about evaluating this task. Reports average accuracy.
Evaluates small language models' capability to adapt to and execute diverse tool-use tasks, including REST API calls, complex SQL queries, web navigation, and command-line interactions. It measures both functional correctness and efficiency under few-shot prompting and lightweight adaptation. Use when the user wants to benchmark on Gorilla APIBench (BFCL V4), Spider 2.0 (Enterprise Subset), WebArena, InterCode (Bash & CTF), or asks about evaluating this task. Reports Execution Success Rate (SR).
Assesses multi-task robotic manipulation capabilities across varying difficulty levels (Easy, Medium, Hard, Very Hard) in simulation to evaluate robustness and generalization. Use when the user wants to benchmark on Meta-World, or asks about evaluating this task. Reports success rate (%).
Evaluates multilingual models' ability to detect metaphorical expressions at the token level and interpret them within a Natural Language Inference (NLI) framework across English and Spanish. It probes cross-lingual transfer, domain generalization, and the impact of metaphorical content on model reasoning. Use when the user wants to benchmark on Meta4XNLI, or asks about evaluating this task. Reports accuracy.
Evaluates few-shot image classification models on real-world domains using a standardized N-way K-shot framework. It probes the ability of models to quickly adapt to new classes with limited labeled examples and generalize across diverse image domains. Use when the user wants to benchmark on MetaDL meta-datasets (1-5), or asks about evaluating this task. Reports average rank.
This benchmark evaluates a robotic framework's ability to fold various garments according to language instructions. It probes the model's spatial-temporal trajectory generation, action prediction accuracy, and generalization across different garment categories and unseen language prompts. Use when the user wants to benchmark on MetaFold dataset, CLOTH3D, or asks about evaluating this task. Reports Success Rate.
Evaluates the few-shot adaptation capability of language models on a diverse set of NLP tasks. It measures how well a model generalizes to unseen tasks after being fine-tuned on automatically extracted few-shot examples from web tables. Use when the user wants to benchmark on Min et al. (2021) Tasks, CROSSFIT, UNIFIEDQA, or asks about evaluating this task. Reports mean Dev Tasks score.
Evaluates reinforcement learning agents on their ability to learn and generalize across multiple robotic manipulation tasks. It tests multi-task RL by measuring performance on a shared set of training tasks, and meta-RL by measuring rapid adaptation to completely unseen tasks from the same distribution. Use when the user wants to benchmark on Meta-World, or asks about evaluating this task. Reports average expected return.
Evaluates the capability of hyperspectral image processing models to detect and segment methane plumes on resource-constrained satellite hardware. It probes the trade-off between detection accuracy (precision, recall, F1) and computational efficiency (runtime) across different spectral enhancement filters and lightweight neural networks. Use when the user wants to benchmark on STARCOP, or asks about evaluating this task. Reports F1.
Evaluates a neural weather model's ability to generate high-resolution, fully dense forecasts of precipitation and surface variables over CONUS from sparse observational data, extending lead times up to 24 hours. Use when the user wants to benchmark on MRMS & OMO Weather Network, or asks about evaluating this task. Reports CRPS.
Evaluates named entity recognition capabilities on biomedical and social media text. It probes the model's ability to identify and classify domain-specific entities (e.g., diseases, drugs, vaccines) in both informal tweets and formal scientific abstracts under fully-supervised and few-shot learning conditions. Use when the user wants to benchmark on METS-CoV, BioRED, or asks about evaluating this task. Reports Micro F1.
Evaluates a model's ability to perform binary classification for identifying conclusion sentences in Hebrew audit reports, and to rank sentence pairs by semantic similarity for hierarchical conclusion allocation. Use when the user wants to benchmark on MevakerConcSen, PS (Parallel Sentences), or asks about evaluating this task. Reports F1, Kendall Rank Correlation (KRC).
This benchmark probes few-shot multimodal word learning under referential uncertainty by testing cross-situational reasoning, semantic bootstrapping, and pragmatic inference. It evaluates how well vision-language and language models generalize from limited examples to name attributes, objects, relations, numbers, and pragmatic concepts compared to human baselines. Use when the user wants to benchmark on MEWL, or asks about evaluating this task. Reports accuracy.
Evaluates a training-free, dynamic multi-expert aggregation framework for multimodal reasoning. It tests the system's ability to select specialized pre-trained experts and synthesize their outputs across video, audio, 3D, and medical domains without fine-tuning. Use when the user wants to benchmark on Video-MMMU, MMAU, SQA3D, M3D, or asks about evaluating this task. Reports accuracy.
Evaluates noise-robust expressive speech-to-speech translation (S2ST) across English-Spanish and Spanish-English directions. It measures how well systems preserve naturalness and expressive style when translating speech under clean and artificially noisy conditions. Use when the user wants to benchmark on mExpresso / mDRAL, or asks about evaluating this task. Reports Naturalness MOS.
Evaluates multilingual and monolingual bi-encoder models on a FAQ retrieval task across 21 languages. It probes cross-lingual knowledge transfer, semantic robustness to lexical changes, and the impact of training data distribution on retrieval performance. Use when the user wants to benchmark on MFAQ, or asks about evaluating this task. Reports MRR.
Evaluates clinical LLMs for demographic bias and fairness across race and gender groups under varying levels of clinical context. It measures how model predictions for ED triage and opioid prescription shift when demographic cues are introduced or clinical information is reduced. Use when the user wants to benchmark on ED-Triage Pool, Opioid Analgesic Recommendation Pool, or asks about evaluating this task. Reports Fairness-Accuracy Balance (FAB) score.
Evaluates large language models on multilingual financial misinformation detection across nine tasks in English, Chinese, Greek, and Bengali. It probes the model's ability to identify false or misleading financial claims, handle numerical sensitivity and reversed causality, and generalize across diverse linguistic settings. Use when the user wants to benchmark on MFMDBench, or asks about evaluating this task. Reports Macro-F1.
Compute mfumanelli/geometric_mean via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mfumanelli/geometric_mean.
This evaluation protocol assesses the effectiveness of a genre-audience reformulation data augmentation strategy on large language model pretraining. It measures downstream task performance across a suite of reasoning, knowledge, and language understanding benchmarks to determine if augmented data improves generalization and mitigates data repetition degradation. Use when the user wants to benchmark on ARC, HellaSwag, Winogrande, MMLU, GSM8K, CSQA, OpenBookQA, PIQA, TriviaQA, or asks about ev...
Evaluates automatic speech recognition and Arabic dialect identification in uncontrolled, real-world settings with high dialectal and genre diversity. It probes robustness to orthographic variability and low-resource conditions by measuring transcription accuracy against multiple human references. Use when the user wants to benchmark on MGB-3, or asks about evaluating this task. Reports MR-WER.
Compute mgfrantz/roc_auc_macro via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mgfrantz/roc_auc_macro.
Evaluates the ability of a multiscale Grassmann manifold framework to cluster single-cell RNA-seq data compared to standard dimensionality reduction and clustering baselines. It probes how well non-Euclidean subspace representations preserve cellular structure and handle varying noise levels across different dataset scales. Use when the user wants to benchmark on GSE75748time, GSE94820, GSE67835, GSE75748cell, GSE109979, GSE84133human1, GSE84133human2, GSE84133human4, GSE57249, or asks about ...
Evaluates the cycle-accurate simulation fidelity and performance of a multi-GPU simulator against real hardware. It probes the simulator's ability to model microarchitectural components (ALU, L1/L2 caches, DRAM) and cross-GPU memory access patterns under unified memory systems. Use when the user wants to benchmark on MGMark, or asks about evaluating this task. Reports execution time.
Evaluates the ability of classifiers to detect and categorize stereotypes across multiple intersecting dimensions (race, gender, profession, religion) in text. It also probes cross-dataset generalization and quantifies generative bias deviation in LLMs. Use when the user wants to benchmark on MGS Dataset (MGSD), StereoSet, CrowsPairs, or asks about evaluating this task. Reports Macro F1 Score.
Evaluates the robustness of machine-generated text detectors against adversarial perturbations such as editing, paraphrasing, prompting, and co-generation. It measures how well detectors maintain binary classification performance when texts are intentionally modified to evade detection. Use when the user wants to benchmark on News-style MGT dataset, or asks about evaluating this task. Reports TPR@FPR.
Evaluates the reasoning, tool-use, and memory-augmented planning capabilities of agents on complex multi-hop QA and visual question answering tasks. It probes how well models leverage episodic memory, test-time learning, and reflection to improve answer accuracy over iterative search trajectories. Use when the user wants to benchmark on FVQA-test, InfoSeek, MMSearch, SimpleVQA, LiveVQA, In-house 1, In-house 2, HotpotQA, 2WikiMultiHopQA, SimpleQA, GAIA (text-only subset), or asks about evaluat...