
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates multi-task learning architectures for post-click conversion rate (CVR) prediction in online advertising, specifically testing how parameter sharing and entire-space training mitigate data sparsity and bias. Use when the user wants to benchmark on Unspecified industry advertising dataset, or asks about evaluating this task. Reports Performance.
Evaluates the robustness of AI-generated text detectors against back-translation manipulations. It probes whether detectors can maintain high true positive rates when AI-generated text is translated to intermediate languages and back-translated to English, preserving semantics while evading detection. Use when the user wants to benchmark on ESPERANTO, or asks about evaluating this task. Reports True Positive Rate (TPR).
Evaluates the quality and linguistic characteristics of argumentative essays generated by different AI models compared to human-written texts. It probes logical structure, vocabulary richness, syntactic complexity, and stylistic markers through expert human annotation. Use when the user wants to benchmark on Student Essay Dataset (90 topics), or asks about evaluating this task. Reports Mean Rating Score.
Evaluates large language models on native Estonian language capabilities across seven tasks covering factual recall, grammar, morphology, vocabulary, summarization, and structured information extraction. The benchmark emphasizes cultural and linguistic authenticity by using human-curated or native-source data without machine translation, assessing both general and domain-specific competencies. Use when the user wants to benchmark on Exams, Trivia, Declension, Words, Grammar, News, Speaker Nam...
Evaluates the psychological and behavioral impact of an AI-generated emotional self-voice intervention compared to text-only and control conditions on goal-related resilience, confidence, motivation, and emotional states. Use when the user wants to benchmark on Custom Human-Subject Intervention Dataset, or asks about evaluating this task. Reports Self-report questionnaire scores.
Evaluates a model's ability to detect phishing addresses on the Ethereum blockchain by analyzing temporal transaction dynamics and graph topology. It probes whether the model can effectively fuse edge-level temporal patterns with node-level structural and statistical features to distinguish malicious accounts from legitimate ones. Use when the user wants to benchmark on $D_1$, $D_2$, $D_3$, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to detect fraudulent Ethereum accounts by analyzing transaction interaction graphs and behavioral patterns. It specifically probes the model's capacity to handle imbalanced label distributions and distinguish between Ponzi schemes and phishing scams using self-supervised feature learning. Use when the user wants to benchmark on Ponzi Scheme Dataset, Phish Scam Dataset, or asks about evaluating this task. Reports Binary-F1.
Evaluates LLMs' ability to classify multiple emotions in low-resource Ethiopian languages (Amharic, Afan Oromo, Somali, Tigrinya) and English. It probes cross-lingual transfer capabilities, the effectiveness of zero-shot and few-shot prompting strategies, and the impact of fine-tuning on multi-label emotion understanding tasks. Use when the user wants to benchmark on EthioEmo, or asks about evaluating this task. Reports Weighted-averaged F1-score.
Machine translation performance across multiple low-resource Ethiopian languages paired with English, evaluating both English-to-Ethiopian and Ethiopian-to-English directions. It probes how model initialization (training from scratch vs. fine-tuning a multilingual model) and available corpus size impact translation quality. Use when the user wants to benchmark on EthioMT, or asks about evaluating this task. Reports spBLEU.
Evaluates the ability of NLP models to detect and classify hate speech in social media comments. It probes both binary hate/non-hate classification and multi-label categorization across specific demographic/identity-based hate categories. The benchmark emphasizes handling overlapping labels and class imbalance typical of real-world user-generated text. Use when the user wants to benchmark on ETHOS, or asks about evaluating this task. Reports F1-score (macro).
Evaluates the ability of event-based motion forecasting models to predict future binary event occurrence maps (ON/OFF channels) from a sequence of past frames. It probes structural preservation of sparse motion traces, temporal consistency, and downstream utility for segmentation and tracking under varying traffic and high-speed motion regimes. Use when the user wants to benchmark on ETram, E-3DTrack, or asks about evaluating this task. Reports aIoU.
This protocol evaluates the transferability and robustness of pre-trained ECG representations by measuring classification performance on downstream cardiac disease datasets. It probes both supervised linear probing capabilities and cross-modal zero-shot generalization to specific cardiac conditions without fine-tuning the encoder. Use when the user wants to benchmark on PTB-XL, CPSC2018, or asks about evaluating this task. Reports AUC.
Evaluates decentralized model aggregation frameworks on edge devices under IID and Non-IID data distributions. It measures classification accuracy and convergence speed to compare hierarchical tree-based learning against centralized federated learning and fully decentralized gossip learning. Use when the user wants to benchmark on HAR Using Smartphones Dataset, Pendigits, or asks about evaluating this task. Reports classification accuracy.
Evaluates and compares encoder-only versus decoder-only language models across classification, retrieval, and generative reasoning benchmarks. It specifically probes architectural strengths, the impact of cross-objective continued pre-training, and performance scaling across parameter sizes (XXS to 1B). Use when the user wants to benchmark on GLUE, MTEB v2, Generative/Reasoning Suite (ARC, HellaSwag, LAMBADA, OBQA, SIQA, TQA, WG, WSC), MS MARCO Dev, or asks about evaluating this task. Reports...
Evaluates LLMs on extracting and predicting EU Taxonomy-compliant Key Performance Indicators from corporate sustainability reports. It probes two capabilities: multi-label classification of economic activities against regulatory definitions, and zero-shot regression of financial KPI percentages (Turnover, CapEx, OpEx) from unstructured text. Use when the user wants to benchmark on EU Taxonomy Sustainability Reports Dataset, or asks about evaluating this task. Reports multi-label text classifi...
This benchmark evaluates whether deep learning models can accurately classify 3D surfaces by their topological genus (Euler characteristic) using point clouds and mesh connectivity graphs. It specifically probes the models' ability to leverage structural and adjacency information rather than relying solely on raw geometric coordinates. Use when the user wants to benchmark on EuLearn, or asks about evaluating this task. Reports F1 score.
This benchmark evaluates long-form, multi- and cross-lingual summarization capabilities in the legal domain. It probes a model's ability to extract or generate concise summaries from lengthy, structurally complex EU legal documents across 24 official EU languages, including cross-lingual transfer scenarios. Use when the user wants to benchmark on EUR-Lex-Sum, or asks about evaluating this task. Reports ROUGE-1.
Granular, capability-level analysis of large foundation models across multimodal reasoning, language understanding, safety, and stability. It dissects performance across fine-grained subcategories (e.g., geometric depth vs. height, single vs. multi-object detection) to reveal persistent failures and complementary strengths across models. Use when the user wants to benchmark on EUREKA-BENCH, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to classify land use and land cover types from multi-spectral satellite imagery. It probes the capability to distinguish between 10 distinct environmental classes using patch-based remote sensing inputs across different spectral band configurations. Use when the user wants to benchmark on EuroSAT, or asks about evaluating this task. Reports classification accuracy (%).
Assesses the utility of the EuroSpeech multilingual corpus for fine-tuning automatic speech recognition (ASR) models. It measures the reduction in word error rate achieved by training on this corpus compared to baseline models across under-resourced European languages. Use when the user wants to benchmark on EuroSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates a multimodal foundation model's ability to predict drug efficacy, classify disease endotypes, and align cross-species transcriptomic and histological data for immunology and inflammation research. Use when the user wants to benchmark on I&I benchmark, IBDome dataset, or asks about evaluating this task. Reports AUROC.
Evaluates multimodal models' ability to detect evasive or deceptive content in e-commerce product listings. It probes fine-grained single-violation detection and long-context, rule-integrated reasoning across multiple overlapping policy categories. Use when the user wants to benchmark on EVADE, or asks about evaluating this task. Reports Full Accuracy.
This protocol evaluates instruction-tuned LLMs across two distinct paradigms: traditional closed-domain NLP benchmarks and open-ended generation quality. It probes whether performance on standard accuracy-based tasks aligns with preference judgments from a large language model judge, highlighting the tension between task-diverse versus style-aligned training data. Use when the user wants to benchmark on MosaicML Eval Gauntlet, LIMA test set, or asks about evaluating this task. Reports accuracy.
Evaluates reference-free LLM prompting strategies as metrics for machine translation and summarization. It measures how well predicted quality scores correlate with human judgments (MQM for MT, human annotations for summarization). Use when the user wants to benchmark on Eval4NLP 2023 Shared Task (MT & Summarization), or asks about evaluating this task. Reports Kendall correlation.
Evaluates text-to-image generation models on two core capabilities: image faithfulness (consistency with real-world commonsense) and text-image alignment (adherence to the conditioning prompt). It probes whether generated images accurately reflect both visual realism and prompt instructions using a fine-grained, human-aligned framework. Use when the user wants to benchmark on EvalAlign, or asks about evaluating this task. Reports EvalAlign_f, EvalAlign_a.
Evaluates the fine-grained alignment and structural fidelity between generated images and their corresponding text prompts. It probes a model's ability to match specific visual elements (e.g., objects, colors, counts) and overall composition against human-annotated ground truth. Use when the user wants to benchmark on EvalMuse-40K, or asks about evaluating this task. Reports SRCC.
This protocol evaluates an LLM's ability to act as a fine-grained text evaluator using custom score rubrics. It tests both absolute grading (assigning a 1–5 score and generating feedback based on a rubric and reference answer) and ranking grading (predicting human preference between two responses). Use when the user wants to benchmark on Feedback Bench, Vicuna Bench, MT Bench, FLASK Eval, MT Bench Human Judgments, HHH Alignment, or asks about evaluating this task. Reports Pearson correlation.
Evaluates whether a two-dimensional risk metric (Evasive Acceleration) can statistically distinguish crash precursors from routine non-crash traffic conflicts at varying lead times before impact. It tests early-warning timeliness and discrimination capability under realistic false-alarm constraints. Use when the user has predictions and gold and needs to compute AUPRC.
Evaluates domain-specific knowledge in Earth Observation and Earth Sciences through multiple-choice QA, hallucination detection, and open-ended QA with and without retrieval context. It also measures the preservation of general capabilities like reasoning, coding, and instruction following after domain adaptation. Use when the user wants to benchmark on EO and Earth Sciences Benchmark, or asks about evaluating this task. Reports Accuracy.
Evaluates vision-language models' ability to recognize evoked emotions from images in a zero-shot setting. It probes their robustness to prompt perturbations and measures sentiment bias in predicting positive vs. negative emotions. Use when the user wants to benchmark on EvE, or asks about evaluating this task. Reports weighted F1 score.
Evaluates an LLM's ability to perform spatial and contextual reasoning for multi-agent planning in 3D scenes. It tests object arrangement, regional context alignment, and scene state tracking to generate plausible character actions and positions. Use when the user wants to benchmark on Event-Driven Storytelling Benchmark, or asks about evaluating this task. Reports success rate.
Evaluates large language models' ability to infer natural language event sequences from real-valued time series data (specifically win probabilities in sports). It probes causal reasoning, temporal context understanding, and the model's capacity to distinguish underlying time series dynamics from linguistic descriptions. Use when the user wants to benchmark on NBA & NFL Event Inference Benchmark, or asks about evaluating this task. Reports accuracy.
EveTAR evaluates Arabic information retrieval systems across four tasks: event detection, ad-hoc search, tweet timeline generation, and real-time summarization. It probes a model's ability to retrieve, cluster, and summarize relevant Arabic tweets in response to event-driven queries or topics. Use when the user wants to benchmark on EveTAR, or asks about evaluating this task. Reports recall.
This benchmark evaluates a model's ability to identify and classify clinical evidence spans within randomized controlled trial (RCT) documents. Specifically, it probes whether a given evidence span supports a significantly decreased, no significant difference, or significantly increased outcome relative to a clinical intervention. Use when the user wants to benchmark on Evidence Inference, or asks about evaluating this task. Reports macro-averaged F1.
This evaluation probes the fidelity of an LLM-assisted pipeline in extracting structured, PICO-style evidence nodes from unstructured full-text biomedical literature, and the precision of subsequent entity normalization against a reference resource. Use when the user wants to benchmark on HCC and CRC PubMed corpus, or asks about evaluating this task. Reports field-level extraction accuracy.
Evaluates the ability of neural radiance field (NeRF) models to accurately reconstruct 3D scenes from 2D images while simultaneously quantifying both aleatoric (data noise) and epistemic (model ignorance) uncertainties. It probes whether uncertainty estimates reliably correlate with actual rendering errors and calibration across varying scene conditions and data sparsity. Use when the user wants to benchmark on Light Field (LF), Local Light Field Fusion (LLFF), RobustNeRF, or asks about evalu...
Evaluates the ability of Evolutive RNNs (EvoRNNs) to model long-range dependencies in sequential data. It compares EvoRNNs against standard RNN baselines on language modeling and sequential recommendation tasks, emphasizing both predictive accuracy and computational efficiency. Use when the user wants to benchmark on LM1B (Billion word dataset), Sequential recommendation dataset, or asks about evaluating this task. Reports MAP@20.
Evaluates dynamic non-IID transfer learning on graphs by measuring how well a model adapts node classification knowledge from a source temporal graph to a target temporal graph with limited labeled samples. It probes the model's ability to handle evolving graph structures, domain discrepancies, and temporal dependencies across heterogeneous datasets. Use when the user wants to benchmark on DBLP-3, DBLP-5, HCP, or asks about evaluating this task. Reports AUC.
Evaluates lifelong language models' ability to update outdated world knowledge while retaining new information and avoiding catastrophic forgetting. It specifically probes temporal adaptation, numerical reasoning, and the model's capacity to forget obsolete facts during continual pretraining. Use when the user wants to benchmark on EvolvingQA, or asks about evaluating this task. Reports Exact Match (EM).
EWMBench evaluates embodied world models on their ability to generate videos that maintain static scene consistency, follow physically plausible motion trajectories, and align semantically with text instructions. It probes whether video generation models can produce action-consistent, task-grounded behaviors for robotic manipulation rather than just visually plausible but static or semantically drifting clips. Use when the user wants to benchmark on Agibot-World, or asks about evaluating this...
Evaluates video-language models' ability to answer questions about expert-level physical skilled activities in long-form videos. It probes fine-grained action recognition, temporal reasoning, and domain-specific generalization across sports, bike repair, cooking, health, music, and dance. Use when the user wants to benchmark on ExAct, or asks about evaluating this task. Reports accuracy.
Compute the ExactMatch metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ExactMatch, or asks how to score with ExactMatch.
Evaluates multilingual and cross-lingual question answering capabilities on high school-level exams across multiple subjects and languages. Probes domain-specific reasoning, knowledge retrieval, and zero-shot transfer between languages with varying linguistic and subject overlaps. Use when the user wants to benchmark on EXAMS, or asks about evaluating this task. Reports accuracy.
Evaluates explainable anomaly detection (AD) and explanation discovery (ED) capabilities on high-dimensional multivariate time series. It probes an algorithm's ability to detect range-based anomalies across four progressive difficulty levels and assesses the quality of generated explanations based on conciseness, consistency, and predictive accuracy. Use when the user wants to benchmark on Exathlon, or asks about evaluating this task. Reports Range-based Precision.
Evaluates a model's ability to recognize and segment specific chart components (e.g., bars, lines, pie slices, legends, axis titles) within chart images using instance segmentation. Use when the user wants to benchmark on ExcelChart400K, or asks about evaluating this task. Reports mAP.
This benchmark evaluates a model's ability to jointly perform Chinese grammatical error correction and generate edit-wise explanations. It probes the model's capacity to identify specific error spans, assign severity levels, and provide natural-language justifications with evidence words, linguistic rules, and revision advice. Use when the user wants to benchmark on EXCGEC, or asks about evaluating this task. Reports CLEME F0.5.
Evaluates foundation models on detecting, forecasting, and monitoring extreme Earth events across seven categories including heatwaves, storms, floods, and wildfires. It probes model generalizability and transferability under data scarcity, distribution shift, and severe class imbalance across heterogeneous geospatial and meteorological modalities. Use when the user wants to benchmark on ExEBench, or asks about evaluating this task. Reports Accuracy (ACC).
Evaluates repository-level code completion capabilities of LLMs across multiple granularities (span, line, expression, statement, function) using executable validation and string similarity metrics. Use when the user wants to benchmark on ExecRepoBench, or asks about evaluating this task. Reports Pass@1.
Measures the wall-clock execution time of four representative Fully Homomorphic Encryption (CKKS) workloads—bootstrapping, logistic regression training, RNN inference, and ResNet-20 inference—across different GPU architectures to evaluate library performance and memory constraints. Use when the user has predictions and gold and needs to compute execution_time.
This benchmark evaluates computer-use agents' ability to correctly judge whether a GUI interaction trajectory succeeds or fails, and to precisely localize the temporal window where the first error occurs. It probes spatiotemporal reasoning, visual redundancy handling, and fine-grained temporal attribution in long video trajectories. Use when the user wants to benchmark on ExeVR-Bench, or asks about evaluating this task. Reports accuracy.