
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates a model's capacity to identify and classify 14 domain-specific legal entities (e.g., Court, Statute, Precedent, Petitioner Name) within unstructured legal documents. This probes fine-grained information extraction capabilities tailored to legal terminology and structure. Use when the user wants to benchmark on LegalEval L-NER Dataset, or asks about evaluating this task. Reports standard F1 score.
Evaluates end-to-end performance of legal retrieval-augmented generation systems by measuring retrieval accuracy, answer correctness, and answer groundedness. It uses a full factorial design across embedding and generative models, and introduces a hierarchical error decomposition to isolate hallucinations, retrieval failures, and reasoning errors. Use when the user wants to benchmark on Legal RAG Bench, or asks about evaluating this task. Reports correctness.
This benchmark assesses models on hierarchical legal text classification tasks (LAP, JP, CP) across three Swiss languages. It tests domain-specific classification accuracy and robustness to long legal documents and multilingual inputs. Use when the user wants to benchmark on Legal Classification (LAP, JP, CP), or asks about evaluating this task. Reports Hierarchical Macro-F1.
This benchmark probes an AI system's ability to detect previously undiscovered legal vulnerabilities within governance frameworks. It tests whether models can identify systemic flaws that could cause immediate disruption without requiring traditional litigation, measuring their capacity for advanced legal reasoning and regulatory logic parsing. Use when the user wants to benchmark on Legal Zero-Days, or asks about evaluating this task. Reports Accuracy.
This benchmark probes large language models' ability to perform diverse, real-world legal reasoning tasks, including rule-recall, issue-spotting, rule-application, interpretation, and rhetorical understanding. It evaluates how well models can apply legal frameworks, classify contractual clauses, and answer questions based on statutory or case law text. Use when the user wants to benchmark on LegalBench, or asks about evaluating this task. Reports accuracy.
Evaluates the retrieval fidelity of RAG systems in the legal domain by measuring how precisely and completely a model retrieves minimal, highly relevant text snippets from legal documents to answer specific queries. Use when the user wants to benchmark on LegalBench-RAG, or asks about evaluating this task. Reports Precision.
Evaluates large language models' ability to reason about and classify Portuguese legal concepts across 31 distinct legal domains. It probes zero-shot question-answering capabilities using multiple-choice, true/false, matching, and case-analysis formats derived from law exam questions. Use when the user wants to benchmark on LegalBench.PT, or asks about evaluating this task. Reports balanced accuracy.
Evaluates a system's ability to retrieve up-to-date legal information from external sources and reason over it to answer multiple-choice legal questions. It probes factual accuracy, uncertainty calibration, and evidence grounding in dynamic legal domains like federal executive orders and tax provisions. Use when the user wants to benchmark on LegalSearchQA, or asks about evaluating this task. Reports Accuracy.
Evaluates a diffusion model's ability to generate egocentric action frames from a pre-action image and a text prompt. It probes the model's capacity to capture action state transitions while preserving contextual information and aligning with natural language instructions in egocentric video domains. Use when the user wants to benchmark on Ego4D, Epic-Kitchens-100, or asks about evaluating this task. Reports EgoVLP score, EgoVLP+ score.
This benchmark evaluates multilingual text-to-speech synthesis and text-based speech editing capabilities. It probes pronunciation stability, cross-lingual generalization, and the perceptual naturalness of localized audio edits across multiple languages. Use when the user wants to benchmark on LEMAS-Dataset, or asks about evaluating this task. Reports WER.
Evaluates the accuracy and robustness of crystal structure fingerprinting and hashing algorithms for de-duplicating quantum chemistry materials databases. It probes sensitivity to structural perturbations (atomic noise, lattice strain, translations) and performance on disordered crystal systems. Use when the user wants to benchmark on LeMat-Bulk, or asks about evaluating this task. Reports success rate.
Evaluates how neural network hyperparameters (activation functions, depth, learning rate) affect output complexity and robustness to input perturbations. Use when the user has predictions and gold and needs to compute Lempel-Ziv Complexity.
Evaluates the ability of multilingual embedding models to retrieve relevant legislative documents given structured metadata queries. It probes cross-lingual semantic alignment and domain-adaptive retrieval performance across varying language resource levels. Use when the user wants to benchmark on LEMUR, or asks about evaluating this task. Reports Acc@k.
Compute leslyarun/fbeta_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of leslyarun/fbeta_score.
Evaluates a video-language model's capacity for structured video understanding across six professional dimensions (subject, aesthetics, camera language, editing, narrative, dissemination) while maintaining general multimodal capabilities. It probes timeline-grounded reasoning, temporal localization, and document/OCR comprehension. Use when the user wants to benchmark on FeedBench, Open Benchmarks (Video-MME, MVBench, MMBench-EN, etc.), or asks about evaluating this task. Reports accuracy.
Evaluates whether a latent world model captures physical structure and dynamics by probing latent representations for physical quantities and measuring predictive surprise under physical versus visual perturbations. Use when the user wants to benchmark on TwoRoom, PushT, OGBench-Cube, Reacher, or asks about evaluating this task. Reports MSE.
Evaluates text-to-image generation models on their ability to accurately render specified text within images, control visual attributes (color, position, font), and maintain aesthetic quality. It measures OCR fidelity, attribute controllability, and human-perceived aesthetics. Use when the user wants to benchmark on LeX-Bench, SimpleBench, CreateBench, AnyText-Benchmark, or asks about evaluating this task. Reports PNED.
Probes large language models' ability to perform structured, multi-step legal reasoning on real-world law exam questions. It evaluates both open-ended legal analysis and multiple-choice selection across diverse jurisdictions and legal domains. Use when the user wants to benchmark on LEXam, or asks about evaluating this task. Reports accuracy.
Evaluates text classification performance and environmental impact (energy, cost, emissions) of NLP models on legal domain datasets. It compares traditional machine learning approaches against transformer-based models across multiple legal benchmarks. Use when the user wants to benchmark on LexGLUE, or asks about evaluating this task. Reports mF1.
Evaluates the robustness and classification accuracy of pre-trained language models when augmented with rule-based lexical simplification as auxiliary inputs. It probes whether lemmatization and rare-word replacement preserve semantic meaning while mitigating lexical diversity effects on downstream NLU tasks. Use when the user wants to benchmark on SST-2, CR, SUBJ, MR, AG, or asks about evaluating this task. Reports accuracy.
Evaluates how well a model learns word meanings and general language modeling performance when trained with lexicon-level contrastive visual grounding. Probes concrete vs. abstract word acquisition, verb relation learning, and next-token prediction accuracy on held-out text. Use when the user wants to benchmark on Word Relatedness, Semantic Feature Prediction, Context Understanding, Lexical Relation Prediction, SimVerb-3500, or asks about evaluating this task. Reports Perplexity.
Probes large language models' ability to extract structured legal relations (relation types and factual arguments) from Chinese civil court judgments. It evaluates both zero-shot prompting and fine-tuning capabilities, while also measuring performance on long-tail relation types and downstream legal reasoning tasks. Use when the user wants to benchmark on LexRel, or asks about evaluating this task. Reports micro-F1.
This benchmark evaluates the ability of sequence-to-sequence models to generate accurate, concise, and faithful summaries of long legal documents across multiple jurisdictions. It probes domain-specific summarization capabilities, testing how well models handle varying input lengths, compression ratios, and the balance between extractive and abstractive generation in legal English. Use when the user wants to benchmark on BillSum, EurLexSum, GovReport, MultiLexSum-Long, MultiLexSum-Short, Mult...
Evaluates multilingual legal NLP models across text classification and named entity recognition tasks. It probes the ability of models to handle domain-specific legal jargon, long-form documents, and cross-lingual generalization across 24 languages. Use when the user wants to benchmark on LEXTREME, or asks about evaluating this task. Reports macro-F1.
Evaluates a malware detection model's ability to adapt to natural concept drift over time using a rolling monthly update setup on real-world Windows malware binaries. Use when the user wants to benchmark on MB-24+, or asks about evaluating this task. Reports accuracy.
Compute LG-Anonym/VerifiableRewardsForScalableLogicalReasoning via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of LG-Anonym/VerifiableRewardsForScalableLogicalReasoning.
Compute lhy/hamming_loss via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of lhy/hamming_loss.
Compute lhy/ranking_loss via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of lhy/ranking_loss.
This evaluation probes a model's ability to perform binary fact-checking on short political claims by mapping multi-class truthfulness labels to positive/negative categories. It specifically tests how well the system handles compositional reasoning and uncertainty, requiring it to output definitive verdicts or abstain. Use when the user wants to benchmark on LIAR, or asks about evaluating this task. Reports accuracy.
Evaluates whether Vision-Language-Action (VLA) models can follow counterfactual language instructions in robotic manipulation tasks. It specifically probes for 'vision shortcuts' where models default to well-learned visual behaviors instead of adhering to the given text commands. Use when the user wants to benchmark on LIBERO-CF, or asks about evaluating this task. Reports grounding rate.
Evaluates a robot policy's ability to sequentially learn multiple manipulation tasks while transferring knowledge and minimizing catastrophic forgetting. It measures forward transfer speed, backward transfer (forgetting), and overall performance across a curriculum of procedurally generated tasks. Use when the user wants to benchmark on LIBERO-LONG, LIBERO-SPATIAL, LIBERO-OBJECT, LIBERO-GOAL, or asks about evaluating this task. Reports FWT.
Evaluates a Vision-Language-Action model's ability to perform precise robotic manipulation and maintain robustness under environmental perturbations. It probes spatial understanding, object manipulation, instruction following, and long-horizon task execution in both standard and perturbed simulation environments. Use when the user wants to benchmark on LIBERO, LIBERO-Plus, or asks about evaluating this task. Reports success_rate.
Evaluates the robustness of vision-language-action (VLA) models under realistic perturbations across seven dimensions (camera, robot, language, light, background, noise, layout). It probes visual shift tolerance, kinematic reasoning, and linguistic robustness by measuring success rates on a curated set of non-trivial tasks. Use when the user wants to benchmark on LIBERO-Plus, or asks about evaluating this task. Reports success rate.
Evaluates the safety and capability of large language models across a broad set of safety tasks, measuring how well models handle direct risky prompts, adversarial attacks, and benign prompts without over-refusal or unsafe generation. Use when the user wants to benchmark on Libra-Eval, or asks about evaluating this task. Reports task_score.
Evaluates the ability of end-to-end speech separation models to isolate target speakers from noisy multi-speaker mixtures. It probes noise-robustness and speaker separation capability under realistic background noise conditions. Use when the user wants to benchmark on Libri2Mix-noisy, Libri3Mix-noisy, or asks about evaluating this task. Reports SI-SNRi (dB).
This benchmark evaluates non-invasive brain-computer interface (BCI) capabilities by testing neural speech decoding from magnetoencephalography (MEG) recordings. It probes a model's ability to detect speech presence, classify phonemes, and identify words from high-fidelity, within-subject neural data aligned with naturalistic audio stimuli. Use when the user wants to benchmark on LibriBrain, or asks about evaluating this task. Reports Balanced Accuracy.
Evaluates continuous speech separation and speaker diarization in reverberant, multi-microphone environments with varying speaker overlap ratios. It probes the system's ability to separate overlapping speech, estimate speaker locations (DOA), cluster them across time blocks, and produce accurate diarization and speech recognition outputs. Use when the user wants to benchmark on LibriCSS, or asks about evaluating this task. Reports DER.
Evaluates automatic speech recognition (ASR) models on long-form audio, measuring accuracy in predicting word and character sequences. It specifically probes the model's ability to handle full-text formatting, including punctuation and casing, and tests performance across different training data scales. Use when the user wants to benchmark on Libriheavy, or asks about evaluating this task. Reports WER.
Probes the ability of zero-shot text-to-speech systems to generate expressive, character-specific utterances while preserving reference speaker timbre. It evaluates cross-sentence generation where a narration clip guides the synthesis of a fictional quotation, testing prosodic variability, emotional expressiveness, and speech intelligibility. Use when the user wants to benchmark on LibriQuote, or asks about evaluating this task. Reports WER.
Evaluates the acoustic quality and prosodic fidelity of generated German-English speech-to-speech translation audio. It measures perceived naturalness via an automated MOS approximation, and quantifies pitch and energy accuracy against ground truth references. Use when the user wants to benchmark on LibriS2S (Frankenstein subset), or asks about evaluating this task. Reports MOSNet score.
Evaluates confidence-based filtering strategies for applying LLMs to post-hoc correction of ASR transcripts, measuring how well the system reduces transcription errors in low-confidence segments while preserving accurate outputs. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports WER.
Evaluates the robustness of speaker recognition models against adversarial audio perturbations by measuring how effectively masked energy attacks disrupt speaker verification while preserving perceptual audio quality. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports EER (%).
Evaluates the robustness of automatic speech recognition (ASR) models under various noise conditions and signal-to-noise ratios (SNRs) using simulated and real-world noisy speech datasets. Use when the user wants to benchmark on LibriSpeech, CHiME-4, or asks about evaluating this task. Reports WER.
Evaluates speech editing models on their ability to accurately modify target words while preserving speaker identity, acoustic quality, and temporal alignment of unedited regions. Use when the user wants to benchmark on LibriSpeech-Edit, or asks about evaluating this task. Reports WER.
Evaluates speech recognition performance under varying amounts of labeled data (1h, 10h, 100h) and different model sizes. It probes the ability of self-supervised speech models to adapt to downstream transcription tasks with limited supervision. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates the ability of end-to-end automatic speech recognition (ASR) models to correctly predict punctuation marks and word capitalization in transcribed speech. It specifically isolates punctuation-specific errors to enable fine-grained comparison between cascade and end-to-end architectures. Use when the user wants to benchmark on LibriSpeech-PC, or asks about evaluating this task. Reports Punctuation Error Rate (PER).
Evaluates the ability of semantic communication systems to transmit speech spectra over noisy wireless channels (AWGN and Rayleigh) and accurately recover text transcriptions, comparing performance against traditional speech and text transceivers. Use when the user wants to benchmark on LibriSpeech, or asks about evaluating this task. Reports Character Error Rate (CER).
Evaluates the word error rate of streaming speech recognition systems using a first-pass RNN-T model followed by a second-pass rescorer. It probes the ability of parallel Transformer rescoring to improve transcription accuracy while maintaining low-latency streaming constraints on-device. Use when the user wants to benchmark on Librispeech, Google Voice Search, or asks about evaluating this task. Reports WER.
Evaluates the ability of a semantic-aware speech-to-text transmission system to accurately reconstruct text from speech signals under noisy communication channels (AWGN and Rayleigh). It probes semantic feature extraction, redundancy removal, and robustness to channel noise. Use when the user wants to benchmark on Librispeech, or asks about evaluating this task. Reports WER.
This protocol evaluates the audio quality and naturalness of a restored multi-speaker TTS corpus (LibriTTS-R) compared to the original LibriTTS dataset. It measures both ground-truth speech fidelity and the downstream impact on multi-speaker TTS model generation quality using human subjective listening tests. Use when the user wants to benchmark on LibriTTS-R, or asks about evaluating this task. Reports MOS.