All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,231 views
Libritts Selfvc Watermark EvalA

Evaluates the robustness of neural audio watermarking systems against self voice conversion attacks and transmission channel distortions, while measuring speaker identity preservation, linguistic content integrity, and perceptual quality. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports bitwise extraction accuracy.

researchpython
0
3
Libritts Ssd EvalA

Evaluates the zero-shot speaker adaptation capability of an autoregressive speech synthesis model on unseen speakers. It measures content accuracy, voice cloning fidelity, and audio quality, while also quantifying inference speedup over standard autoregressive decoding. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Libritts Tts EvalA

Evaluates the naturalness and quality of synthesized speech from text-to-speech models trained on the LibriTTS corpus. It probes how audio sampling rate, text normalization, and sentence-level splitting affect human-perceived speech naturalness compared to the original LibriSpeech dataset. Use when the user wants to benchmark on LibriTTS, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Lifelong Rl EvalA

Evaluates the ability of reinforcement learning agents to sequentially learn multiple tasks while retaining prior knowledge, generalizing to unseen environments, and leveraging forward transfer from previous tasks. It probes parameter isolation, knowledge composition, and robustness across discrete and continuous action spaces with varying reward and input distributions. Use when the user wants to benchmark on ProcGen, CT-graph, Minigrid, Continual World, or asks about evaluating this task. R...

researchpythonperformance
0
3
Lightgcn Rec EvalA

Evaluates the ability of graph-based collaborative filtering models to rank relevant items for users based on sparse user-item interaction graphs. It probes how well neighborhood aggregation and embedding smoothing capture latent preferences without relying on node semantic features. Use when the user wants to benchmark on Gowalla, Yelp2018, Amazon-Book, or asks about evaluating this task. Reports recall@20.

researchpythongo
0
3
Lightning Wildfire Prediction EvalA

Evaluates machine learning models' ability to classify lightning-ignited wildfire occurrences versus non-occurrences using meteorological, vegetation, and spatio-temporal features. It probes the models' generalization across different feature configurations and geographic regions, highlighting the necessity of separate models for lightning versus anthropogenic fires. Use when the user wants to benchmark on Global Lightning-Ignited Wildfire Dataset, or asks about evaluating this task. Reports ...

datapythongo
0
3
Lightweight Action Recognition EvalA

Evaluates the real-world efficiency (training and inference latency, VRAM footprint) of video action recognition models across desktop GPUs and mobile devices, alongside their classification accuracy on standard benchmarks. Use when the user wants to benchmark on EK100, SSV2, K400, or asks about evaluating this task. Reports relative latency.

researchpython
0
3
Lime Mmt47 EvalA

Evaluates the ability of lightweight Mixture of Experts (MoE) parameter-efficient fine-tuning methods to generalize across diverse multimodal tasks. It probes how well shared PEFT modules with expert modulation vectors capture task-specific specialization without learned routing parameters. Use when the user wants to benchmark on MMT-47, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Limitgen EvalA

Evaluates whether LLMs can accurately identify and articulate critical limitations in scientific research papers across methodological, experimental, analytical, and literature-related dimensions. The benchmark probes the model's ability to ground critiques in domain-specific best practices and produce actionable, substantive feedback rather than superficial presentation critiques. Use when the user wants to benchmark on LimitGen, or asks about evaluating this task. Reports Limitation Quality.

researchpythongo
0
3
Linglanmidian EvalA

Evaluates LLMs on Traditional Chinese Medicine (TCM) knowledge recall, multi-hop clinical reasoning, information extraction, and clinical decision-making. It probes synonym-tolerant clinical labeling, robustness on curated hard subsets, and performance across diverse TCM-specific task formats including QA, NER, and dosage prediction. Use when the user wants to benchmark on LingLanMiDian, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Lingo Space Grounding EvalA

Evaluates a model's ability to ground natural language spatial instructions to specific 2D pixel locations in RGB-D tabletop scenes. It probes both single-relation grounding and incremental/compositional grounding where multiple spatial predicates must be satisfied sequentially or simultaneously. Use when the user wants to benchmark on CLIPort Benchmark, ParaGon Benchmark, SREM Benchmark, LINGO-Space Benchmark, Composite Instruction Task, or asks about evaluating this task. Reports success sc...

researchpythongo
0
3
Lingoly EvalA

This benchmark evaluates large language models' ability to perform multi-step linguistic reasoning and deductive puzzle solving in low-resource and extinct languages. It probes out-of-domain grammatical inference and instruction-following under conditions of minimal pre-training exposure, requiring models to extract and apply novel rules from provided context rather than relying on memorized knowledge. Use when the user wants to benchmark on LINGOLY, or asks about evaluating this task. Report...

researchpythongo
0
3
Lingoqa EvalA

Evaluates vision-language models on autonomous driving video question answering, testing their ability to understand temporal visual context, describe scenes, predict actions, and justify answers based on driving scenarios. Use when the user wants to benchmark on LingoQA, or asks about evaluating this task. Reports Ling-Judge.

researchpythongo
0
3
Linguasafe EvalA

Evaluates multilingual safety alignment of LLMs by measuring their ability to reject harmful prompts and accept benign ones across 12 languages and a hierarchical safety taxonomy. It probes both direct safety performance (vulnerability to harmful content) and indirect performance (oversensitivity to benign requests). Use when the user wants to benchmark on LinguaSafe, or asks about evaluating this task. Reports Vulnerability Score.

researchpythongo
0
3
Linguistic DiversityA

Evaluates the lexical, semantic, and structural diversity of natural language instructions across datasets. It quantifies repetition, vocabulary richness, semantic coverage, and syntactic complexity to identify construction biases and limitations in dataset design. Use when the user has predictions and gold and needs to compute ROUGE-L.

researchpythongo
0
3
Linguistic Probing EvalA

Evaluates how fine-tuning on downstream NLP tasks redistributes linguistic knowledge across transformer layers. It probes for part-of-speech tagging, syntactic chunking, and semantic tagging capabilities using linear classifiers on layer-wise hidden states. Use when the user wants to benchmark on Penn TreeBank, CoNLL 2000, Parallel Meaning Bank, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Linguistic Shibboleth Hiring EvalA

Evaluates whether LLMs systematically penalize candidates for using hedging language in professional interview responses, despite identical substantive content. It probes the model's ability to decouple communication style from perceived technical competence and hiring suitability. Use when the user wants to benchmark on Linguistic Shibboleth Hiring Benchmark, or asks about evaluating this task. Reports average_score.

researchpythongo
0
3
Link Prediction EvalA

Evaluates a model's ability to predict missing or future edges in a graph based on its structural topology. It specifically probes whether the model learns meaningful graph patterns or merely exploits implicit degree biases inherent in the standard edge sampling procedure. Use when the user wants to benchmark on Empirical graphs (90 datasets), or asks about evaluating this task. Reports AUC-ROC.

researchpythonnode
0
3
Lip EvalA

Evaluates a model's ability to perform joint human semantic part segmentation and 16-keypoint pose estimation on diverse, unconstrained images with varying appearances, occlusions, and backgrounds. It probes the model's capacity to leverage structural body priors to resolve ambiguities in part boundaries and joint localization. Use when the user wants to benchmark on LIP, PASCAL-Person-Part, MPII Human Pose, ATR, or asks about evaluating this task. Reports mean IoU.

researchpython
0
3
Lip To Speech EvalA

Evaluates a model's ability to synthesize high-fidelity, intelligible speech directly from visual lip movements. It probes perceptual audio quality, content accuracy, and speaker identity preservation in a cross-dataset generalization setting. Use when the user wants to benchmark on LRS3-TED, LRS2-BBC, or asks about evaluating this task. Reports WER.

researchpython
0
3
Lip2wav EvalA

Evaluates a model's ability to synthesize natural, speaker-specific speech from unconstrained lip movements in large-vocabulary settings. It measures how well the model captures individual speaking styles and contextual cues from video frames. Use when the user wants to benchmark on Lip2Wav, GRID, TCD-TIMIT lip speaker corpus, or asks about evaluating this task. Reports mel reconstruction loss.

researchpythongo
0
3
Liputan6 EvalA

Evaluates the ability of models to generate concise, accurate summaries of Indonesian news articles. It probes both extractive and abstractive summarization capabilities, highlighting performance gaps on more abstract summaries and the limitations of n-gram overlap metrics. Use when the user wants to benchmark on Liputan6, or asks about evaluating this task. Reports ROUGE F-1 (R1, R2, RL).

researchpythongit
0
3
LipvertexerrorA

Compute the LipVertexError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute LipVertexError, or asks how to score with LipVertexError.

documentationpython
0
3
Literaryqa EvalA

Evaluates long-context language models on their ability to answer complex, abstractive questions about entire literary works. It probes narrative event understanding and semantic correctness, measuring how well generated answers align with human judgment rather than just matching reference strings. Use when the user wants to benchmark on LiteraryQA, or asks about evaluating this task. Reports ROUGE-L.

ai-agentspythongit
0
3
Lithology Classification EvalA

Evaluates a model's ability to perform multi-class sequence labeling on multi-channel well-log time series data for lithology classification. It tests the model's capacity to handle diverse geological settings, manage distribution shifts, and produce stratigraphically plausible predictions. Use when the user wants to benchmark on SEAM, Facies, FORCE, GeoLink, or asks about evaluating this task. Reports Weighted F1.

businesspythongit
0
3
Litqa2 EvalA

Evaluates an agentic system's ability to retrieve relevant scientific literature, re-rank passages, and answer multiple-choice questions based on the retrieved text. It probes retrieval coverage, passage localization, and question-answering accuracy under information loss constraints. Use when the user wants to benchmark on LitQA2, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Litxbench EvalA

Evaluates an LLM's ability to extract structured material science experimental data from scientific literature text. It specifically probes the model's capacity to link extracted measurements to material processing lineages and adhere to a predefined schema. Use when the user wants to benchmark on LitXAlloy, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Live Meta Mcg EvalA

Evaluates the ability of objective video quality assessment models to predict human-perceived quality of mobile cloud gaming videos distorted by compression and resizing artifacts. It benchmarks both general-purpose and gaming-specific no-reference models against human subjective ratings. Use when the user wants to benchmark on LIVE-Meta Mobile Cloud Gaming (LIVE-Meta MCG), or asks about evaluating this task. Reports SROCC.

researchpythongo
0
3
Liveaopsbench EvalB

Evaluates large language models' mathematical reasoning capabilities on Olympiad-level competition problems. It specifically probes whether models possess genuine problem-solving skills or merely rely on memorized pre-training data by using a continuously updated, timestamped benchmark to measure contamination-resistant accuracy. Use when the user wants to benchmark on AoPS24, Math, OlympiadBench, OmniMath, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Livebench EvalA

Evaluates large language models across 18 tasks spanning math, coding, reasoning, language, instruction following, and data analysis. It uses dynamically updated, objectively scored questions from recent real-world sources to minimize test-set contamination and avoid LLM-judging biases. Use when the user wants to benchmark on LiveBench, or asks about evaluating this task. Reports LiveBench score.

researchpythongo
0
3
Liveclin EvalA

LiveClin probes real-world clinical reasoning and longitudinal case management by evaluating models on dynamic, multimodal patient scenarios. It tests the ability to maintain context across sequential diagnostic, treatment, and follow-up questions while resisting data contamination from static training corpora. Use when the user wants to benchmark on LiveClin, or asks about evaluating this task. Reports Case Accuracy.

researchpythongo
0
3
Liveclktb EvalA

Evaluates multilingual LLMs on their ability to transfer factual knowledge across languages using time-sensitive, real-world events that occur after the model's training cutoff. It measures both in-language factual recall and cross-lingual generalization performance across domains like music, movies, and sports. Use when the user wants to benchmark on LiveCLKTBench, or asks about evaluating this task. Reports Transfer Score.

researchpythongo
0
3
Livecodebench EvalA

Evaluates large language models' ability to generate, repair, execute, and predict outputs for code across multiple algorithmic and competitive programming problem types. It provides a holistic, contamination-free assessment by continuously updating with new problems and testing models across four distinct coding scenarios. Use when the user wants to benchmark on LiveCodeBench, or asks about evaluating this task. Reports PASS@1.

developmentpythongo
0
3
Livefact EvalA

Evaluates LLMs on fake news detection under dynamic, time-evolving evidence streams. It probes both binary classification capability (Real/Fake) and reasoning capability (handling ambiguity when evidence is incomplete), while explicitly measuring benchmark data contamination and epistemic humility. Use when the user wants to benchmark on LiveFact November 2025 dataset, or asks about evaluating this task. Reports average_score.

researchpythongo
0
3
Livemcpbench EvalA

Evaluates LLM agents' capability to dynamically discover, select, and chain Model Context Protocol (MCP) tools to complete complex, multi-step real-world tasks. It probes meta-tool learning and multi-tool collaboration in large-scale, time-varying tool ecosystems. Use when the user wants to benchmark on LiveMCPBench, or asks about evaluating this task. Reports task success rate.

ai-agentspythongo
0
3
Liveweb Ie EvalA

Evaluates web information extraction systems on live, dynamically evolving websites by testing their ability to identify target attributes and extract corresponding values from natural language queries. It probes the robustness of extraction pipelines against real-time web dynamics and complex layouts that break static HTML parsing. Use when the user wants to benchmark on LiveWeb-IE, or asks about evaluating this task. Reports F1 score.

researchpythontesting
0
3
Livexiv EvalA

This benchmark evaluates the multi-modal reasoning capabilities of Large Multimodal Models (LMMs) on scientific content scraped from ArXiv papers. It specifically probes visual question answering (VQA) on figures and table question answering (TQA) using multiple-choice formats derived from real-time academic publications. Use when the user wants to benchmark on LiveXiv, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Livs T2i Alignment EvalA

Evaluates how well text-to-image models align with pluralistic, intersectional community preferences for urban public space design. Probes whether multi-criteria preference optimization (DPO) improves alignment over a baseline, and how prompt origin and annotator demographics influence preference consistency and rating distributions. Use when the user wants to benchmark on LIVS, or asks about evaluating this task. Reports preference_rate.

researchpython
0
3
Livvie Accents Unplugged EvalA

Compute livvie/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of livvie/accents_unplugged_eval.

developmentpython
0
3
Llama Berry EvalA

Evaluates LLMs on complex mathematical reasoning using search-based inference (SR-MCTS) rather than direct generation. It measures success rates across varying difficulty levels, from grade-school math to Olympiad-level problems, by testing both majority-vote and best-of-k strategies. Use when the user wants to benchmark on AIME24, AMC23, Math Odyssey, GPQA Diamond, OlympiadBench, College Math, MMLU STEM, GSM8K, GSMHard, MATH500, or asks about evaluating this task. Reports major@k.

researchpythongo
0
3
Llama Energy Latency EvalA

This benchmark evaluates the inference latency and energy consumption of LLaMA models (7B-65B) across different GPU hardware (V100, A100) and sharding configurations. It probes the trade-offs between computational throughput, power usage, and hardware efficiency during text generation. Use when the user wants to benchmark on Alpaca, GSM8K, or asks about evaluating this task. Reports energy per second (Watts).

researchpythonperformance
0
3
Llama Vits Tts EvalA

Evaluates the naturalness, intelligibility, and emotional expressiveness of a non-autoregressive TTS model enhanced with LLM-derived semantic embeddings. Probes how well semantic tokens from Llama2 versus BERT improve acoustic quality and emotion similarity compared to baselines. Use when the user wants to benchmark on LJSpeech, 1-hour LJSpeech, EmoV_DB_bea_sem, or asks about evaluating this task. Reports ESMOS.

researchpythonexpress
0
3
Llama4 Benchmark EvalA

Evaluates the language understanding, reasoning, coding, multilingual, and multimodal capabilities of Llama 4 models across a standardized suite of academic and industry benchmarks. It measures performance on text-only, code, and vision-language tasks using few-shot or zero-shot prompting protocols. Use when the user wants to benchmark on MMLU, MMLU-Pro, MATH, MBPP, LiveCodeBench, GPQA Diamond, ChartQA, DocVQA, MMMU, MTOB, or asks about evaluating this task. Reports macro_avg/acc.

researchpythongo
0
3
Llamarec EvalA

This evaluation probes a model's ability to rank candidate items based on a user's sequential interaction history and item metadata. It measures how effectively the system retrieves and re-ranks relevant products or movies against a large candidate pool using standard recommendation metrics. Use when the user wants to benchmark on ML-100k, Beauty, Games, or asks about evaluating this task. Reports NDCG@k.

ai-agentspythontesting
0
3
Llara EvalA

Evaluates a model's ability to predict the next item in a user's sequential interaction history by leveraging both semantic item metadata and behavioral embeddings. It also measures the model's instruction-following capability in generating valid recommendations from a candidate set. Use when the user wants to benchmark on MovieLens100K, Steam, or asks about evaluating this task. Reports HitRatio@1.

ai-agentspythongo
0
3
Llase G1 EvalA

Evaluates a LLaMA-based generative speech enhancement model's ability to perform multiple audio restoration tasks (noise suppression, packet loss concealment, target speaker extraction, acoustic echo cancellation, and speech separation) in a task-agnostic manner. It probes the model's capacity to preserve acoustic fidelity and semantic content while generalizing across different acoustic conditions and device types. Use when the user wants to benchmark on DNS Challenge blind test set (Intersp...

researchpythongo
0
3
Llava Bench EvalA

Assesses multimodal chatbot capabilities, including conversation, detailed description, and complex visual reasoning. It measures how well a model follows instructions and understands novel or challenging visual inputs compared to a strong text-only baseline. Use when the user wants to benchmark on LLaVA-Bench, or asks about evaluating this task. Reports relative_score.

researchpythonperformance
0
3
Llava Cot EvalA

Evaluates the effectiveness of structured chain-of-thought prompting and test-time scaling algorithms on multimodal reasoning tasks. It probes whether enforcing a specific reasoning order (summary, caption, reasoning, conclusion) and selecting among multiple generated candidates improves answer accuracy over baseline prompting or dense supervision. Use when the user wants to benchmark on Unspecified multimodal reasoning benchmarks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Llava Le EvalA

Evaluates multimodal instruction-following and complex geological reasoning on lunar surface imagery. It probes the model's ability to interpret crater morphology, degradation states, and inferred geological processes beyond simple visual description. Use when the user wants to benchmark on LUCID (held-out eval set), or asks about evaluating this task. Reports average_overall_score.

researchpython
0
3
Llava Onevision 1.5 EvalA

This evaluation probes the multimodal reasoning, visual question answering, OCR, and chart understanding capabilities of large multimodal models. It tests the model's ability to process high-resolution images, extract fine-grained text, and perform complex reasoning across diverse visual domains. Use when the user wants to benchmark on MMStar, MMEBench, MME-RealWorld, SeedBench, CV-Bench, RealWorldQA, MathVista, WeMath, MathVision, MMMU, MMMU-Pro, ChartQA, CharXiv, DocVQA, OCRBench, AI2D, Inf...

researchpythongo
0
3