
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates long-term conversational memory, temporal reasoning, and factual consistency across multi-session dialogues and noisy search-augmented contexts. Probes an agent's ability to retrieve, resolve temporal conflicts, and answer complex queries over extended interaction histories. Use when the user wants to benchmark on LOCOMO, LongMemEval, SealQA-Hard, or asks about evaluating this task. Reports LOCOMO Overall Accuracy.
Evaluates the ability of LLMs to use external tools by testing three progressively complex capabilities: direct API slot-filling, API retrieval from a catalog, and multi-step planning combined with retrieval and calling. It measures how well models understand instructions, locate relevant functions, and generate correctly formatted calls with valid parameters. Use when the user wants to benchmark on API-Bank, or asks about evaluating this task. Reports API call correctness.
Evaluates the retrieval accuracy and ranking quality of query-based API recommendation systems for Java APIs at both class and method levels. It also measures how query reformulation techniques impact recommendation performance. Use when the user wants to benchmark on APIBench-Q, or asks about evaluating this task. Reports Success Rate@k.
Evaluates the ability of molecular graph machine learning models and fingerprint-based methods to predict binary pesticide toxicity to honey bees. It specifically probes domain generalization by testing performance on structurally novel compounds and temporally separated data rather than random splits. Use when the user wants to benchmark on ApisTox, or asks about evaluating this task. Reports MCC.
Evaluates the quality and transferability of semi-supervised multi-view graph embeddings for Android applications across classification, clustering, and link prediction tasks. Probes whether multi-view and semi-supervised learning improve embedding accuracy and scalability compared to unimodal baselines. Use when the user wants to benchmark on Batch malware detection, Online malware detection, Malware familial clustering, Clone detection, App recommendation, or asks about evaluating this task...
This benchmark probes vision-language models' ability to infer non-observable, culturally grounded structured metadata (culture, period, origin, creator) from images of heritage artifacts. It evaluates whether models can go beyond visual perception to perform semantic alignment with museum annotations, revealing culture-dependent reasoning capabilities and potential biases. Use when the user wants to benchmark on Appear2Meaning, or asks about evaluating this task. Reports exact match accuracy.
Evaluates zero-shot generalization of action recognition models to appearance-free videos generated by warping noise or random dots with optical flow, testing reliance on motion cues over static shape and texture. Use when the user wants to benchmark on UCF5, AFD5, AFF5, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to generate correct Python code from natural language problem descriptions. It measures functional correctness by executing generated programs against a large bank of automated test cases, rather than relying on text-similarity metrics like BLEU. Use when the user wants to benchmark on APPS, or asks about evaluating this task. Reports strict_accuracy.
This benchmark evaluates automatic speech recognition (ASR) systems on their ability to transcribe long-form, spontaneous call-center dialogues across 14 English accents. It specifically probes robustness to non-standard accents, conversational speech patterns, and sensitivity to audio segmentation strategies. Use when the user wants to benchmark on AppTek Call-Center Dialogues, or asks about evaluating this task. Reports WER.
Evaluates the ability of code language models to automatically generate correct patches for real-world Java bugs. It probes whether pre-trained or fine-tuned models can produce syntactically valid and semantically correct code that passes developer-written test suites and survives manual verification. Use when the user wants to benchmark on Defects4J v1.2, Defects4J v2.0, QuixBugs, HumanEval-Java, or asks about evaluating this task. Reports plausible_patch.
This protocol evaluates an LLM's ability to predict a paper's future scientific impact based on its text and peer reviews, and its ability to iteratively revise the manuscript to maximize that predicted impact. It probes the model's capacity for rubric discovery, agentic text editing, and alignment with human expert preferences. Use when the user wants to benchmark on ICLR & NeurIPS Peer Review Dataset, or asks about evaluating this task. Reports MAE, Improvement Score ($\Delta S$).
This evaluation probes the operational reliability and safety alignment of large language models under repeated inference. It specifically measures how stochastic decoding and sampling depth expose intermittent safety failures, refusal inconsistencies, and guardrail instability that single-generation benchmarks typically mask. Use when the user wants to benchmark on APST Safety Prompt Set (AIR-BENCH Equivalent), or asks about evaluating this task. Reports empirical failure probability.
Evaluates machine learning models' ability to predict Air Quality Index (AQI) across Indian cities using historical pollutant concentrations and location metadata. It probes the models' capacity to capture temporal patterns and spatial variations in air pollution, particularly around agricultural burning events. Use when the user wants to benchmark on Indian Air Quality Monitoring Dataset (22 stations), or asks about evaluating this task. Reports R².
This benchmark evaluates audio question answering models on their ability to correctly answer standard multiple-choice questions and, crucially, to detect and reject unanswerable cases. It specifically probes three failure modes: missing correct options, categorical mismatches between questions and answers, and questions irrelevant to the audio input. Use when the user wants to benchmark on AQUA-Bench, or asks about evaluating this task. Reports conditional accuracy (CA).
Evaluates deep learning models' ability to classify marine species from underwater images under challenging environmental conditions like turbidity, low illumination, and occlusion. It probes robustness to visual distortions, class imbalance, and fine-grained feature discrimination in complex aquatic scenes. Use when the user wants to benchmark on AQUA20, or asks about evaluating this task. Reports Accuracy.
Evaluates a 2B-parameter vision-language model's visual understanding, knowledge reasoning, and text reading capabilities across a comprehensive suite of standard multimodal benchmarks. It measures how well the model handles general VQA, mathematical reasoning, and document comprehension tasks. Use when the user wants to benchmark on MMBench, MMStar, MMMU, MathVista, HallusionBench, AI2D, OCRBench, MMVet, or asks about evaluating this task. Reports accuracy.
Evaluates large language models on appellate review tasks for criminal judgments, specifically detecting, classifying, and correcting legal errors in finalized court decisions. It probes fine-grained legal reasoning, diagnostic accuracy, and the ability to generate legally valid corrections. Use when the user wants to benchmark on AR-Bench, or asks about evaluating this task. Reports Accuracy (Acc), Macro F1 (MaF1).
Evaluates a model's ability to track multiple visual objects in video sequences using auditory referring expressions instead of text. It probes cross-modal alignment, robustness to varying weather and video quality conditions, and handling of complex, unconstrained traffic dynamics. Use when the user wants to benchmark on Echo-KITTI, Echo-KITTI+, Echo-BDD, or asks about evaluating this task. Reports HOTA.
Evaluates large language models' ability to perform commonsense reasoning within specific Arab cultural contexts. It probes region-specific grounding by testing models across 13 Arab countries and 12 cultural domains, measuring how well they understand local norms, habits, and social scenarios in both Arabic and English. Use when the user wants to benchmark on ArabCulture, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of Automatic Speech Recognition (ASR) models to accurately transcribe spoken Arabic from real-world telephonic calls. It probes robustness to dialectal diversity, variable audio quality, and background noise typical of call-domain environments. Use when the user wants to benchmark on Arabic Call Domain Benchmark, or asks about evaluating this task. Reports WER.
This benchmark evaluates a model's ability to identify check-worthy factual claims within Arabic social media posts. It probes the system's capacity to filter out non-factual or irrelevant content and prioritize claims that require verification based on public interest and potential impact. Use when the user wants to benchmark on Arabic Check-Worthiness Dataset, or asks about evaluating this task. Reports P@30.
This benchmark evaluates a model's ability to classify the veracity of Arabic social media claims as true or false. It probes factual consistency and reasoning against reliable sources in a binary classification setting. Use when the user wants to benchmark on Arabic Claim Verification Dataset, or asks about evaluating this task. Reports Macro-F1.
This benchmark tests a system's ability to retrieve relevant evidence snippets from a large pool of web pages for a given Arabic claim. It evaluates ranking performance in a retrieval setting where only a small fraction of snippets actually contain verifying evidence. Use when the user wants to benchmark on Arabic Evidence Retrieval Dataset, or asks about evaluating this task. Reports P@10.
Evaluates the semantic textual similarity (STS) capability of Arabic text embedding models, specifically testing how Matryoshka Representation Learning and hybrid loss training preserve semantic alignment across different embedding dimensions. Use when the user wants to benchmark on MTEB Arabic STS (STS17, STS22, STS22-v2), or asks about evaluating this task. Reports STS correlation (Pearson/Spearman, scaled 0-100).
ArabicMMLU probes massive multitask language understanding in Modern Standard Arabic across 40 educational subjects spanning STEM, social sciences, humanities, Arabic language, and culturally specific domains. It evaluates models on cross-lingual transfer, cultural localization, and robustness to linguistic phenomena like negation across primary, middle, high school, and university levels. The benchmark specifically tests how well models handle Arabic-specific knowledge and exam-style multipl...
Evaluates Arabic-centric and cross-lingual text embedding models across multiple linguistic, cultural, and domain-specific capabilities. It probes how well models capture dialectal variations, regional cultural knowledge, and specialized domain terminology in Arabic. Use when the user wants to benchmark on ArabicMTEB, or asks about evaluating this task. Reports Avg..
Evaluates LLMs' capabilities in understanding and generating dialectal Arabic (Levantine, Egyptian, Gulf) and assessing cultural awareness. It probes dialect identification, text generation, cognitive reasoning, and machine translation across dialects diverging from Modern Standard Arabic. Use when the user wants to benchmark on AraDiCE, or asks about evaluating this task. Reports F1 score.
This benchmark evaluates Arabic language models on healthcare-related question answering, specifically probing their ability to classify mental health conditions and generate culturally appropriate medical advice. It tests both discriminative capabilities (multi-label classification and multiple-choice selection) and generative capabilities (open-ended response generation) in clinical and mental health contexts. Use when the user wants to benchmark on AraHealthQA, or asks about evaluating thi...
Evaluates the quality of English-to-Arabic translations for text-to-SQL tasks and measures the execution accuracy of generated SQL queries on Arabic natural language questions. Use when the user wants to benchmark on AraSpider, or asks about evaluating this task. Reports Execution accuracy.
This benchmark probes large language models' ability to reason over and understand Arabic tabular data across three core tasks: direct question answering, fact verification, and complex reasoning. It specifically tests whether models can extract, compare, and synthesize information from structured Arabic tables while adhering to linguistic and formatting nuances. Use when the user wants to benchmark on AraTable, or asks about evaluating this task. Reports accuracy.
Evaluates the adversarial robustness of binarized neural networks (BNNs) against white-box and black-box attacks. It measures how well BNNs maintain prediction accuracy under controlled perturbation budgets compared to their clean accuracy, highlighting robustness trends across dataset scales. Use when the user wants to benchmark on CIFAR-10, ImageNet, or asks about evaluating this task. Reports ACC_norm.
Evaluates human action recognition models under varying camera viewpoints and different subjects. It probes the model's ability to generalize across unseen subjects, unseen camera angles, and continuous 360-degree view changes. Use when the user wants to benchmark on Varying-view RGB-D Action Dataset, or asks about evaluating this task. Reports average recognition accuracy.
Measures general fluid intelligence and developer-aware generalization by requiring systems to infer abstract transformation rules from few input-output grid demonstrations and apply them to novel test cases. It explicitly avoids measuring task-specific memorization or crystallized knowledge, focusing instead on abstraction, reasoning, and broad generalization under strict prior constraints. Use when the user wants to benchmark on Abstraction and Reasoning Corpus (ARC), or asks about evaluati...
Evaluates large language models' ability to generate correct Python code for interactive data science notebooks, requiring multi-turn reasoning, grounded understanding of DataFrame schemas, and composition of pandas API calls based on preceding notebook context and natural language intents. Use when the user wants to benchmark on ARCADE, or asks about evaluating this task. Reports pass@k.
Evaluates the ability of LLM/VLM systems to generate high-quality, narrative-coherent presentation slides from academic papers. It probes content coverage, rhetorical structure preservation, textual fluency, and visual layout quality compared to human-authored references. Use when the user wants to benchmark on ArcBench, or asks about evaluating this task. Reports VLM-based Q/A Quiz Accuracy.
Evaluates a multimodal document understanding model across visual QA, multilingual text comprehension, English reading comprehension, and table extraction tasks. It probes the model's ability to process long-context documents, answer questions from images or text, and extract structured tabular data from unstructured layouts. Use when the user wants to benchmark on SQuAD2.0, DocVQA, Arctic-TILT, MLQA, xQuAD, or asks about evaluating this task. Reports ANLS*.
Evaluates the ability of automated black-box testing tools and reinforcement learning agents to explore Android applications effectively. It probes how well algorithms navigate complex UI states, maximize code/activity coverage, and trigger unique application crashes within a fixed time budget. Use when the user wants to benchmark on F-Droid top starred apps, AndroTest, Synthetic FATE models, or asks about evaluating this task. Reports AUC.
This evaluation protocol assesses the safety alignment and general capability of large language models after undergoing an adaptive red-teaming and repair process. It probes the model's ability to refuse harmful or unsafe prompts while maintaining performance on standard knowledge and reasoning benchmarks, and measures the false refusal rate to ensure utility is preserved. Use when the user wants to benchmark on RedTeam, StrongReject, HarmBench, PKU-SafeRLHF, XSTest, MMLU, GSM8K, TruthfulQA, ...
Evaluates the ability of spoof-speech detection models to distinguish between real (bonafide) Arabic speech and synthetic speech generated by various Text-to-Speech (TTS) models across multiple Arabic dialects. It probes robustness against different voice-cloning systems and measures detection accuracy alongside perceptual realism and ASR fidelity. Use when the user wants to benchmark on ArFake, or asks about evaluating this task. Reports Equal Error Rate (EER).
Compute argmaxinc/detailed-wer via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of argmaxinc/detailed-wer.
Evaluates the out-of-distribution generalization capability of trajectory prediction models on unseen HD map geometries. It measures how well models maintain prediction accuracy when transferred from seen to unseen domains without retraining. Use when the user wants to benchmark on argoverse-shift, or asks about evaluating this task. Reports minADE.
Evaluates autonomous driving models on joint trajectory prediction and controllable generation tasks. It probes the model's ability to forecast multi-agent future paths accurately and generate realistic, goal-conditioned trajectories efficiently using diffusion-based sampling. Use when the user wants to benchmark on Argoverse 2, or asks about evaluating this task. Reports avgBrierMinFDE_K.
Evaluates the plausibility, diversity, and kinematic consistency of generated multi-agent traffic scene continuations conditioned on 5-second histories. It probes a model's ability to generate realistic, diverse, and physically plausible future trajectories over a 6-second horizon, including out-of-distribution generalization across different autonomous driving datasets. Use when the user wants to benchmark on Argoverse 2 (A2), Waymo (WO), or asks about evaluating this task. Reports minADE.
Evaluates the ability of dialogue agents to select supportive facts from scientific papers and generate contextually appropriate responses in argumentative scientific dialogues. Probes document-grounded response generation and fact selection under expert-level, opinion-driven interactions. Use when the user wants to benchmark on ArgSciChat, or asks about evaluating this task. Reports Fact-F1.
Evaluates NeRF-based models on their ability to synthesize novel views from egocentric, multimodal sensor data captured in dynamic real-world environments. It probes how well current neural rendering methods handle temporal dynamics, lens distortion, and non-visual cues like IMU and gaze. Use when the user wants to benchmark on Aria-NeRF Dataset, or asks about evaluating this task. Reports PSNR.
Evaluates unsupervised anomaly detection models on credit card transaction time series to identify fraudulent spending deviations. It probes the ability of models to balance precision and recall in highly imbalanced, real-world financial data without relying on labeled fraud examples. Use when the user wants to benchmark on Credit card transaction time series, or asks about evaluating this task. Reports F-Measure.
Evaluates 3D indoor scene understanding by testing object detection (single-frame and whole-scene) and color-guided depth upsampling on real-world mobile RGB-D data captured with consumer LiDAR devices. Use when the user wants to benchmark on ARKitScenes, or asks about evaluating this task. Reports mAP (mean average precision).
Evaluates code embedding models on a semantic code retrieval task where the goal is to find the correct ArkTS function given a natural language docstring or comment. It probes the model's ability to align bilingual documentation with declarative UI and distributed application code semantics. Use when the user wants to benchmark on ArkTS-CodeSearch, or asks about evaluating this task. Reports MRR.
This benchmark evaluates how AI model size, pruning, and quantization affect deployment on bare-metal ARM Cortex-M0+/M4/M7 processors. It specifically probes the trade-offs between energy efficiency, inference latency, and model accuracy across different hardware architectures and application duty cycles. Use when the user wants to benchmark on Embedded AI Use Cases (e.g., Optical Digit Recognition, Visual Wake Words), or asks about evaluating this task. Reports inference cycle energy.
Evaluates the semantic retrieval and similarity capabilities of text embedding models on low-resource Armenian text. It probes cross-lingual alignment, domain coverage, and robustness to noisy synthetic training data by measuring retrieval accuracy and semantic similarity correlation across diverse tasks. Use when the user wants to benchmark on MTEB [hye], Manual Retrieval Dataset, MS MARCO [hye], STS [hye], or asks about evaluating this task. Reports Average.