
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates AI-driven log analytics systems on three core tasks: parsing unstructured log messages into event templates, compressing log data efficiently, and detecting system anomalies using supervised or unsupervised models. It probes how well algorithms generalize across diverse, real-world system logs ranging from distributed systems to mobile apps. Use when the user wants to benchmark on Loghub, or asks about evaluating this task. Reports Parsing Accuracy (PA).
This benchmark evaluates an LLM's ability to perform complex multi-step logical reasoning and chain-of-thought decomposition to generate correct SQL queries from natural language questions. It probes capabilities in mathematical deduction, physical knowledge integration, and cross-domain analytical querying. Use when the user wants to benchmark on LogicCat, or asks about evaluating this task. Reports Execution Accuracy (EX).
Evaluates a model's ability to perform multi-hop question answering and information retrieval over extremely long contexts (up to 128K tokens), while also measuring preservation of short-context reasoning and instruction-following capabilities. Use when the user wants to benchmark on LongBench v1, LongBench v2, MMLU, MATH-500, IFEval, Needle in a Haystack, RULER, or asks about evaluating this task. Reports pass@1 accuracy.
Probes a model's ability to accurately retrieve hidden, specific information (needles) embedded within extremely long multimodal sequences (text, video, audio) and assesses its predictive stability over millions of tokens. Use when the user wants to benchmark on Paul Graham Essays (Synthetic), AlphaGo Documentary, VoxPopuli, or asks about evaluating this task. Reports recall.
This evaluation probes a model's ability to perform complex, multi-step reasoning across mathematics, coding, and scientific domains. It specifically measures the capacity to generate long chain-of-thought traces and produce correct final answers or executable code under strict generation constraints. Use when the user wants to benchmark on AIME24, AIME25, GPQA Diamond, LiveCodeBench v5, LiveCodeBench v6, or asks about evaluating this task. Reports average accuracy.
Evaluates the quality of abstractive summaries for long scientific documents by measuring n-gram overlap between generated text and reference abstracts. Use when the user wants to benchmark on arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-1.
Evaluates long-form scientific summarization models on their ability to generate relevant and faithful abstracts across clinical, chemical, and biomedical domains. It probes how calibration set construction and candidate selection strategies affect model performance on standard relevance and faithfulness metrics. Use when the user wants to benchmark on Scientific Summarization Datasets, or asks about evaluating this task. Reports Rouge-1 F1.
Evaluates named entity recognition capabilities, specifically probing a model's ability to handle class imbalance, out-of-vocabulary terms, and long or complex entity names across biomedical and general domain texts. Use when the user wants to benchmark on NCBI-disease, BC5CDR-disease, BC5CDR-chemical, BC4CHEMD, BC2GM, JNLPBA, LINNAEUS, Species-800, CoNLL-2003, WNUT-2017, or asks about evaluating this task. Reports F1.
Evaluates efficient Transformer architectures on long-context sequence modeling tasks spanning text, images, and structured data. Probes capabilities in hierarchical reasoning, spatial navigation, and retrieval over sequences up to 16K tokens. Use when the user wants to benchmark on ListOps, Text Classification, Retrieval, Image Classification, Pathfinder / Path-X, or asks about evaluating this task. Reports accuracy.
Evaluates the capability of sequence modeling architectures to capture long-range dependencies across text, audio, and image modalities, as well as their computational efficiency and compatibility with standard Transformer and CNN backbones. Use when the user wants to benchmark on Long Range Arena (LRA), Speech Commands (SC), WikiText-103, GLUE, ImageNet-1k, or asks about evaluating this task. Reports accuracy.
This evaluation probes a session-based recommendation model's ability to accurately predict the next item in a user's interaction sequence while mitigating popularity bias. It measures both standard ranking accuracy and the model's capacity to recommend long-tail items, ensuring recommendations align with user-specific item distribution preferences rather than just global popularity. Use when the user wants to benchmark on YOOCHOOSE, Last.fm, or asks about evaluating this task. Reports Recall...
Evaluates long-term motion representations derived from dense point-tracking against image-based baselines across five perceptual tasks. It probes temporal generalization, motion representation efficiency, and the ability to capture spatio-temporal dynamics for classification and regression. Use when the user wants to benchmark on SSV2 (Temporal Dataset subset), Jester, VB100, RAVDESS, MITFabric, ADVIO, or asks about evaluating this task. Reports classification accuracy.
Evaluates long-term visual place recognition (VPR) capabilities in dynamic underwater benthic environments. It probes a model's ability to geolocate camera views over multi-year intervals despite habitat changes, varying terrain ruggedness, and sub-decimeter registration errors. Use when the user wants to benchmark on Benthic Reference Sites Dataset, or asks about evaluating this task. Reports Recall@K.
Evaluates long-form video understanding and multimodal reasoning capabilities across multiple-choice question answering tasks. It probes the model's ability to handle extended temporal dependencies, spatial-temporal reasoning, and tool-augmented retrieval in videos ranging from short clips to hour-long content. Use when the user wants to benchmark on LongVideoBench, VideoMME, LVBench, MLVU, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform temporal reasoning and question-answering on long-duration videos (several minutes to over an hour). It probes memory retention, attention allocation across extended sequences, and the capacity to filter irrelevant visual content while preserving critical frames. Use when the user wants to benchmark on LongVideoBench, MLVU, VideoMME (Long), LVBench, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to follow instructions and retrieve information in long-context scenarios (up to 64k tokens), while also measuring their general capabilities and instruction-following performance in short-context settings. Use when the user wants to benchmark on LongBench-Chat, LongBench, MT-Bench, ARC, HellaSwag, TruthfulQA, MMLU, or asks about evaluating this task. Reports GPT-4 rating (1-10).
Evaluates large language models' ability to understand and process long contexts across bilingual (English and Chinese) multitask scenarios, including single/multi-document QA, summarization, few-shot learning, code completion, and synthetic tasks. Use when the user wants to benchmark on LongBench, or asks about evaluating this task. Reports F1.
Evaluates long-context understanding and reasoning capabilities of LLMs across bilingual (English/Chinese) tasks. It probes retrieval, ranking, ordering, multiple-choice, information extraction, and summarization under varying difficulty levels and context lengths. Use when the user wants to benchmark on LongBench Pro, or asks about evaluating this task. Reports LongBench Pro Score.
Evaluates large language models' ability to comprehend and reason over realistic, extremely long contexts (up to 2M words) across six multitask domains. It probes deep understanding rather than shallow extraction by using challenging multiple-choice questions that require extended reasoning and careful reading. Use when the user wants to benchmark on LongBench v2, or asks about evaluating this task. Reports accuracy.
Evaluates an LLM's ability to generate ultra-long, coherent, and high-quality text (up to 20k words) while strictly adhering to explicit length constraints. It probes long-context generation capabilities, structural coherence over extended outputs, and instruction-following for length requirements. Use when the user wants to benchmark on LongBench-Write, or asks about evaluating this task. Reports Sq (Quality Score).
Probes a model's ability to maintain coherent state, plan, and execute multi-step reasoning over long, interdependent chains of thought spanning tens to hundreds of thousands of tokens across domains like chemistry, mathematics, and chess. Use when the user wants to benchmark on LongCoT, or asks about evaluating this task. Reports accuracy.
Evaluates embedding models' ability to retrieve relevant information from long contexts (up to 32k tokens) and compares the effectiveness of various context window extension strategies. It also probes the extrapolation capabilities of Absolute Positional Encoding (APE) versus Rotary Positional Encoding (RoPE) in retrieval tasks. Use when the user wants to benchmark on LongEmbed, or asks about evaluating this task. Reports accuracy (%).
Evaluates instruction-following long-form text generation capabilities of LLMs across diverse in-domain and out-of-domain tasks, including news summarization, recipe and story generation, long-form QA, and multilingual generation, while also measuring general language understanding via MMLU. Use when the user wants to benchmark on LongForm-C, Writing Prompts, ELI5, Recipe Generation, MMLU, MLSUM, or asks about evaluating this task. Reports METEOR.
Evaluates the ability of LLMs to maintain accuracy and logical consistency when generating long-text responses that answer multiple sequential questions from GSM8K or MMLU in a single pass. It specifically probes performance degradation as the number of generated questions increases. Use when the user wants to benchmark on LongGenBench-GSM8K, LongGenBench-MMLU, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to generate personalized long-form text by integrating retrieved user profiles into a retrieval-augmented generation framework. It probes how well models can adapt their writing style and content to specific user attributes across different domains like emails, abstracts, reviews, and topic-based writing. Use when the user wants to benchmark on LongLaMP, or asks about evaluating this task. Reports METEOR.
Evaluates the effectiveness of a question-aware prompt compression framework on long-context LLM tasks. It measures how well compressed prompts preserve key information and answer accuracy across multi-document QA, summarization, and code completion scenarios. Use when the user wants to benchmark on NaturalQuestions (Liu et al., 2023), LongBench (Bai et al., 2023), ZeroSCROLLS (Shaham et al., 2023), or asks about evaluating this task. Reports accuracy.
This evaluation protocol assesses the long-context understanding, instruction-following, and faithfulness capabilities of LLMs. It combines automated AI-judged scoring on long and short-context benchmarks with human preference alignment tests to validate the effectiveness of the LongReward training method. Use when the user wants to benchmark on LongBench, LongBench-Chat, MT-Bench, AlpacaEval2, or asks about evaluating this task. Reports GPT-4o rating.
Evaluates long-form speech processing capabilities across transcription, translation, summarization, and higher-level reasoning tasks. It probes models' ability to maintain semantic consistency, track temporal progression, and extract structured information from ~10-minute audio segments. Use when the user wants to benchmark on LongSpeech, or asks about evaluating this task. Reports WER, BLEU-4, Numeric Accuracy, Strict Accuracy.
This benchmark evaluates long-context video-language understanding by testing a model's ability to retrieve specific moments from lengthy videos and reason over multimodal details. It distinguishes between single-moment visual perception and multi-moment relational reasoning across 17 fine-grained categories. Use when the user wants to benchmark on LongVideoBench, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of language models to comprehend and reason over long documents (up to 32k+ tokens) by testing short and long dependency tasks, including question answering, cloze completion, and summarization. Use when the user wants to benchmark on LooGLE, or asks about evaluating this task. Reports GPT4_score.
Evaluates long-context language models across 22 benchmarks and 140 tasks, probing capabilities like long-form generation, information retrieval, and reasoning over extended contexts. It also assesses the efficiency of inference acceleration and RAG augmentation methods. Use when the user wants to benchmark on LOOMBench, or asks about evaluating this task. Reports task_accuracy.
Probes long-context multi-document question answering by requiring models to synthesize evidence from all provided documents (10K–250K+ tokens) across financial reports and academic papers. It tests information extraction, comparison, clustering, and chain-of-reasoning capabilities in heterogeneous, document-level agentic retrieval settings. Use when the user wants to benchmark on Loong, or asks about evaluating this task. Reports Avg Score.
Evaluates a model's ability to accurately regress one-loop scattering amplitudes across a high-dimensional kinematic phase space. It specifically probes precision in challenging regions and the reliability of uncertainty quantification inherent to Bayesian neural networks. Use when the user wants to benchmark on One-loop gg→γγg(g) amplitudes, or asks about evaluating this task. Reports Δ (relative amplitude error).
Evaluates an adaptive dual-phase LLM inference acceleration system for multi-turn dialogues, probing its ability to maintain generation accuracy and reduce computational overhead across varying query positions in long-context conversations. It specifically tests whether the system can generalize beyond positional heuristics used by static KV cache compression methods. Use when the user wants to benchmark on MFQA-en, 2WikiMQA, Musique, HotpotQA, NrtvQA, Qasper, MultiNews, GovReport, QMSum, TRE...
Evaluates automatic speech recognition (ASR) model performance across varying training data scales and model sizes, measuring generalization to in-domain and out-of-domain English speech benchmarks. Use when the user wants to benchmark on Loquacious Set, Librispeech, Voxpopuli, CommonVoice, or asks about evaluating this task. Reports WER.
Evaluates the effectiveness of transformer-specific dropout methods (e.g., HiddenKey, DropKey, HiddenCut) when combined with LoRA for parameter-efficient fine-tuning. It probes the model's ability to mitigate overfitting in LoRA settings across diverse natural language understanding and generation tasks. Use when the user wants to benchmark on GLUE, E2E, WebNLG, or asks about evaluating this task. Reports Accuracy, BLEU.
Evaluates parameter-efficient fine-tuning (LoRA) for cross-domain few-shot object detection on aerial imagery. It probes the model's ability to generalize to new domains with limited labeled data while mitigating overfitting. Use when the user wants to benchmark on DOTA, DIOR, or asks about evaluating this task. Reports mAP@0.5.
This evaluation probes the effectiveness of LoRA fine-tuning across 31 diverse NLP tasks by comparing base LLMs against their fine-tuned counterparts and proprietary models like GPT-4. It measures how much performance lift fine-tuning provides and whether smaller open-weight models can surpass larger closed-source models after adaptation. Use when the user wants to benchmark on magicoder, mmlu, glue_wnli, arc_combined, wikisql, boolq, customer_support, glue_cola, winogrande, glue_sst2, dbpedi...
This benchmark probes a model's ability to infer the exact number of training images used to fine-tune a Low-Rank Adaptation (LoRA) adapter solely from its learned weight matrices. It evaluates how well spectral and norm-based features of LoRA parameters correlate with and reveal the scale of the underlying training dataset. Use when the user wants to benchmark on LoRA-WiSE, or asks about evaluating this task. Reports MAE.
Evaluates a model's ability to generate long-term (25s–50s) high-fidelity musical waveforms that are rhythmically synchronized with visual cues from diverse video scenarios like dancing and sports. It measures both the temporal alignment of generated beats with ground-truth audio and the overall subjective musical quality. Use when the user wants to benchmark on LORIS, or asks about evaluating this task. Reports F1.
Evaluates monaural speech enhancement models by measuring how effectively they restore clean speech from noisy recordings. The benchmark probes the model's ability to handle diverse acoustic conditions and varying signal-to-noise ratios across two standard speech enhancement datasets. Use when the user wants to benchmark on VCTK+DEMAND, DNS Challenge 2020, or asks about evaluating this task. Reports PESQ.
Compute LottieW/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of LottieW/accents_unplugged_eval.
Evaluates the effectiveness of a diffusion-based data augmentation method on mitigating long-tail bias and improving cross-domain generalization in remote-sensing semantic segmentation. It specifically probes whether synthetic label-image pairs can increase minority-class exposure while preserving domain realism and data distribution. Use when the user wants to benchmark on LoveDA, or asks about evaluating this task. Reports mIoU.
Evaluates a model's ability to retrieve relevant long-form videos or fine-grained clips based on rich, narrative-driven text queries. It probes temporal reasoning, semantic alignment across extended durations, and robustness to long-context inputs and varying frame sampling strategies. Use when the user wants to benchmark on LoVR, or asks about evaluating this task. Reports Recall@K.
Evaluates automatic speech recognition (ASR) performance on low-resource and high-resource languages using synthetic audio generated from text augmentation. It probes the model's ability to generalize to unseen lexical and syntactic variations when trained on limited real speech data. Use when the user wants to benchmark on Vatlongos, Nashta, Kakabe, Shinekhen Buryat, LibriSpeech, or asks about evaluating this task. Reports WER.
Evaluates sentence embedding quality for low-resource languages trained on synthetic triplet data. It probes cross-lingual semantic similarity and information retrieval capabilities by benchmarking against human-annotated and unsupervised baselines without requiring target-language training data. Use when the user wants to benchmark on Ousidhoum STS/STR, MTEB Retrieval (Low-Resource Subset), or asks about evaluating this task. Reports Spearman's correlation.
Evaluates cross-lingual transfer learning for Named Entity Recognition in low-resource Indian languages (Hindi and Marathi) by measuring how well models trained on combined or assisting-language datasets generalize to target language test sets compared to monolingual baselines. Use when the user wants to benchmark on IIT Bombay (Marathi), IJCNLP (Hindi), Wiki ANN (Hindi and Marathi), or asks about evaluating this task. Reports scores.
Evaluates relation extraction models under low-resource conditions (8-shot, 10%, 100% training data) across diverse domains and languages. It probes few-shot learning capabilities, robustness to long-tailed class distributions, and the effectiveness of data augmentation and self-training strategies. Use when the user wants to benchmark on SemEval 2010 Task 8, TACREV, DialogRE, DuIE2.0, Wiki80, ChemProt, SciERC, CMeIE, or asks about evaluating this task. Reports Macro F1.
Evaluates cross-lingual robustness of spoofing countermeasures by measuring spoof rejection rates on a multilingual synthetic-speech corpus at a fixed operating point calibrated on external benchmarks. Probes how language and synthesizer identity independently affect spoof detection performance. Use when the user wants to benchmark on Low-Resource Language Spoofing Corpus, or asks about evaluating this task. Reports spoof rejection rate (SRR).
Evaluates a model's ability to generate accurate radiology reports from single-view X-ray images under varying quality conditions, from standard to severely degraded. It probes robustness to clinical acquisition artifacts and tests whether the model can extract quality-invariant diagnostic features without relying on historical patient data. Use when the user wants to benchmark on MIMIC-CXR LRRG Benchmarks, or asks about evaluating this task. Reports CheXbert F1.