All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,297 views
Speaker Independent Voice Conv EvalA

Evaluates a model's ability to convert speech emotion (neutral to angry) while preserving speaker identity across both seen and unseen speakers. It measures spectral and prosody conversion quality objectively, and assesses perceived speech quality, emotion similarity, and speaker similarity subjectively. Use when the user wants to benchmark on English emotional speech corpus, EmoV-DB, JL-Corpus, or asks about evaluating this task. Reports MCD, LSD, PCC.

researchpythonexpress
0
3
Speakersleuth EvalA

This benchmark evaluates Large Audio-Language Models (LALMs) on their ability to detect, localize, and discriminate speaker inconsistencies in multi-turn dialogues. It specifically probes whether models rely on acoustic cues or are biased toward textual coherence when judging speaker consistency. Use when the user wants to benchmark on SpeakerSleuth, or asks about evaluating this task. Reports detection_accuracy, discrimination_accuracy, localization_f1.

researchpythongo
0
3
Speakstream Streaming Tts EvalA

Evaluates the quality and latency of streaming text-to-speech models that generate audio incrementally from interleaved text and speech inputs. It probes the model's ability to maintain speech accuracy and naturalness while minimizing first-token latency under streaming constraints. Use when the user wants to benchmark on LJSpeech, LibriSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythonperformance
0
3
Spearman CorrelationA

Measures the rank-order agreement between an NLG evaluator's predicted scores and human reference judgments. It probes the evaluator's ability to capture task-specific quality dimensions such as fluency, coherence, consistency, and groundedness across summarization and dialogue generation. Use when the user has predictions and gold and needs to compute Spearman correlation.

ai-agentspythongo
0
3
SpearmancorrcoefA

Compute the SpearmanCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpearmanCorrCoef, or asks how to score with SpearmanCorrCoef.

documentationpython
0
3
SpearmanrA

Compute the spearmanr metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute spearmanr, or asks how to score with spearmanr.

documentationpython
0
3
Spec Compliance EvalA

Evaluates whether frontier LLM responses adhere to their published model specifications when faced with generated value tradeoff scenarios. It also measures the consistency of model-based judges in detecting specification violations and identifies specification flaws like contradictions and ambiguities. Use when the user wants to benchmark on Generated Value Tradeoff Scenarios, or asks about evaluating this task. Reports compliance.

researchpythonaws
0
3
Spec Cpu2017 Speed EvalA

Evaluates CPU energy efficiency and performance under varying RAPL power caps and core counts using standard SPEC CPU 2017 benchmarks. Probes how power capping influences the trade-off between energy consumption and execution latency across memory-intensive, balanced, and compute-intensive workloads. Use when the user wants to benchmark on SPEC CPU 2017 Speed suite, or asks about evaluating this task. Reports normalized energy usage.

researchpythonperformance
0
3
Spec Narrowing EvalA

Evaluates two solver-based algorithms for synthesizing minimal test suites to distinguish between candidate formal specifications (Alloy models). It measures how execution time and test suite size scale with the number of candidate specifications and the domain scope. Use when the user wants to benchmark on Alloy4Fun, or asks about evaluating this task. Reports execution_time.

researchpythongo
0
3
SpecificityA

Compute the Specificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute Specificity, or asks how to score with Specificity.

documentationpythondocumentation
0
3
SpecificityatsensitivityA

Compute the SpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpecificityAtSensitivity, or asks how to score with SpecificityAtSensitivity.

documentationpythondocumentation
0
3
Specpower EvalA

Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying JVM-based web service workloads. It measures how power consumption scales with workload intensity and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECpower_ssj2008, or asks about evaluating this task. Reports watts.

researchpython
0
3
SpectralanglemapperA

Compute the SpectralAngleMapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpectralAngleMapper, or asks how to score with SpectralAngleMapper.

documentationpython
0
3
SpectraldistortionindexA

Compute the SpectralDistortionIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpectralDistortionIndex, or asks how to score with SpectralDistortionIndex.

documentationpython
0
3
Spectrumbench EvalA

Evaluates multimodal large language models on spectroscopy tasks spanning signal processing, perception, semantic understanding, and molecular generation. It probes the models' ability to align cross-modal data, reason over spectral patterns, and generate accurate chemical structures or spectra. Use when the user wants to benchmark on SpectrumBench, or asks about evaluating this task. Reports accuracy (%).

researchpythongo
0
3
Specweb EvalA

Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying PHP-based e-commerce web workloads. It measures how power consumption scales with session load and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECweb2009, or asks about evaluating this task. Reports watts.

researchpythonphp
0
3
Speech Commands EvalA

This benchmark evaluates limited-vocabulary keyword spotting models for on-device speech recognition. It probes a model's ability to correctly identify isolated spoken words from a fixed set of 10 commands, while also handling background silence and unrecognized speech in both aligned and continuous streaming audio contexts. Use when the user wants to benchmark on Speech Commands, or asks about evaluating this task. Reports Top-One Error.

researchpythongo
0
3
Speech Continuation Bias EvalA

This benchmark probes voice-based and gendered biases in speech continuation models by evaluating how well generated continuations preserve semantic coherence, sentiment, agency, emotional framing, and avoid objectification across different voice qualities (breathy, creaky, end creak) and speaker genders. Use when the user wants to benchmark on SS_set, NOP_set, or asks about evaluating this task. Reports Semantic Coherence.

researchpython
0
3
Speech Deepfake Detection EvalA

Evaluates a model's ability to distinguish between authentic human speech and synthetically generated or manipulated speech (deepfakes), with a specific focus on robustness against expressive and emotional synthesis attacks. Use when the user wants to benchmark on LibriSpeech, ASVspoof 2019 LA, ASVspoof 2021 LA, ASVspoof 2024, EmoFake, EmoSpoof-TTS, or asks about evaluating this task. Reports EER (Equal Error Rate).

researchpythonexpress
0
3
Speech Df Arena EvalA

This benchmark evaluates the robustness and cross-domain generalization of speech deepfake detection models across diverse synthetic speech generation techniques, including TTS, voice conversion, neural codecs, and real-world social media leaks. It measures how well models maintain performance when faced with unseen attack types, languages, and distribution shifts. Use when the user wants to benchmark on ASVspoof 2019, ASVspoof 2021, ASVspoof 2024, ADD 2022, ADD 2023, CodecFake, LibriSeVoc, S...

researchpythongo
0
3
Speech Drame Archetype EvalA

Evaluates a speech foundation model's ability to generate role-play responses that align with top-down, stereotype-driven character archetypes within specific scene contexts. It measures how well the model captures broad, accessible scoring criteria and general impressions of character consistency without relying on fine-grained prosodic nuance. Use when the user wants to benchmark on DRAME-RoleBench (Archetype), or asks about evaluating this task. Reports archetype_score.

researchpythongit
0
3
Speech Drame Realism EvalA

Evaluates a speech foundation model's ability to generate role-play responses grounded in bottom-up, human-grounded realism. It measures fine-grained aspects such as prosodic dynamics, emotional fidelity, character consistency, and contextual fit based on realistic dialogue and media sources. Use when the user wants to benchmark on DRAME-RoleBench (Realism), or asks about evaluating this task. Reports realism_score.

researchpythongit
0
3
Speech Enhancement EvalA

Evaluates speech enhancement models on their ability to restore audio degraded by multiple distortion types (noise, reverberation, packet loss, clipping, bandwidth limitation, codec artifacts) across varying sampling rates, while preserving speaker identity, phonetic content, and perceptual quality. Use when the user wants to benchmark on DNS 2020 test set, PLC 2024 validation set, VoiceFixer GSR test set, URGENT 2025 non-blind test set, or asks about evaluating this task. Reports DNSMOS.

researchpythongo
0
3
Speech Enhancement Ms EvalA

Evaluates semi-supervised speech enhancement algorithms by measuring how well they recover clean speech from noisy mixtures across varying SNRs and noise types. It probes both perceptual quality and speech intelligibility preservation. Use when the user wants to benchmark on IEEE Speech + Environmental/Industrial Noise Database, or asks about evaluating this task. Reports HASQI.

researchpythongo
0
3
Speech Likelihood EvalA

Evaluates the ability of generative latent variable models and autoregressive baselines to model speech audio distributions at varying temporal resolutions. It measures how well models capture intra-frame and inter-frame correlations in audio waveforms by optimizing likelihood objectives. Use when the user wants to benchmark on TIMIT, LibriSpeech, or asks about evaluating this task. Reports bits per frame (bpf).

researchpythongo
0
3
Speech Rep EvalA

Evaluates the quality of compressed semantic speech representations across automatic speech recognition, speech-to-text translation, and voice conversion tasks. It probes how well adaptive entropy-based token aggregation preserves linguistic and acoustic information under varying compression ratios. Use when the user wants to benchmark on LibriSpeech, CVSS-C, or asks about evaluating this task. Reports WER.

researchpythonexpress
0
3
Speech Separation EvalA

Evaluates speech separation models' ability to isolate individual speaker signals from multi-speaker mixtures under various acoustic conditions, including moving sources, environmental noise, and musical noise. It measures both objective signal quality and subjective perceptual metrics to assess generalization from synthetic to real-world dynamic scenarios. Use when the user wants to benchmark on SonicSet, RealSEP, HumanSEP, LRS2-2Mix, Libri2Mix, or asks about evaluating this task. Reports SI...

researchpythonperformance
0
3
Speechfake EvalA

This benchmark evaluates speech deepfake detection models on their ability to generalize across diverse synthesis methods (TTS, voice conversion, neural vocoders), multiple languages, and unseen speaker identities. It probes whether models learn inherent spoofing artifacts or merely memorize specific speaker voices or generation techniques. Use when the user wants to benchmark on SpeechFake, or asks about evaluating this task. Reports EER.

researchpythongit
0
3
Speechglue EvalA

Evaluates whether self-supervised speech models capture linguistic knowledge by probing them on a speech-adapted version of the GLUE benchmark. Tasks include grammaticality judgment, sentence similarity, paraphrase identification, and natural language inference, all converted from text to speech via TTS. Use when the user wants to benchmark on SpeechGLUE, or asks about evaluating this task. Reports Accuracy (Acc).

researchpythongo
0
3
Speechinstructbench EvalA

Evaluates speech instruction-following capabilities across closed-ended, open-ended, and adjustment tasks under varying acoustic conditions (background noise, accents, disfluencies) in English and Chinese. Use when the user wants to benchmark on SpeechInstructBench, or asks about evaluating this task. Reports instruction-level accuracy (I).

researchpythongo
0
3
Speechjudge EvalA

This benchmark evaluates how well automated metrics and large audio-language models can judge speech naturalness and detect deepfakes compared to human preferences. It probes the alignment of computational quality scores with human perceptual judgments across multiple languages and TTS models. Use when the user wants to benchmark on SpeechJudge-Eval, or asks about evaluating this task. Reports AudioLLM Pairwise Accuracy.

researchpythonapi
0
3
Speechllm EvalA

Evaluates a unified speech-language model's ability to perform end-to-end automatic speech recognition (ASR), named entity recognition (NER), and sentiment analysis (SA) on low-resource speech datasets. It tests parameter-efficient adapter-based alignment of speech encoder features to a language model, along with classifier regularization and LoRA fine-tuning. Use when the user wants to benchmark on LibriSpeech, SLUE-VoxPopuli, SLUE-VoxCeleb, or asks about evaluating this task. Reports WER, F...

researchpythonperformance
0
3
Speechmedbench EvalA

Evaluates speech language models on medical consultation tasks, covering single-turn medical knowledge Q&A, multi-turn diagnostic conversations, real-world clinical robustness, and speech output quality. Use when the user wants to benchmark on SpeechMedBench, CMB, CME, MedDG, AIHospital, MedSafetyBench, Wild, or asks about evaluating this task. Reports CMB, CME, MedDG, AIHospital.

researchpythongo
0
3
Speechmentalmanip EvalA

Binary classification of spoken multi-speaker dialogues to detect the presence of mental manipulation tactics. It probes an audio-language model's ability to identify subtle manipulative cues in synthetic speech without relying on text transcripts. Use when the user wants to benchmark on SpeechMentalManip, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Speechparaling Bench EvalA

Evaluates large audio-language models (LALMs) on their ability to generate speech with fine-grained paralinguistic features, including dynamic intra-utterance variation and context-aware adaptation. It probes how well models interpret and modulate tone, pitch, emotion, and non-linguistic vocalizations in response to textual instructions and contextual cues. Use when the user wants to benchmark on SpeechParaling-Bench, or asks about evaluating this task. Reports Judge Score (0-100).

researchpython
0
3
Speechr EvalA

Probes speech reasoning capabilities in large audio-language models across factual, procedural, and normative dimensions. It tests whether models can perform multi-step inference, maintain logical coherence, and make normative judgments when processing spoken input under varying prosodic and emotional conditions. Use when the user wants to benchmark on SpeechR, or asks about evaluating this task. Reports Accuracy.

researchpythongit
0
3
Speed Bench EvalA

Evaluates the accuracy and throughput of speculative decoding methods across diverse semantic domains and varying input sequence lengths. It probes how draft length, batch size, vocabulary pruning, and inference frameworks impact real-world serving efficiency compared to baseline autoregressive generation. Use when the user wants to benchmark on SPEED-Bench, or asks about evaluating this task. Reports AL.

researchpythongo
0
3
Spert EvalA

Evaluates a model's ability to jointly identify named entity spans with their types and extract relational tuples between them from unstructured text. It probes span-based representation learning, localized context modeling, and joint classification without relying on sequential tagging schemes like BIO. Use when the user wants to benchmark on CoNLL04, SciERC, ADE, or asks about evaluating this task. Reports F1 score (micro/macro-averaged).

researchpythongo
0
3
Spgispeech 2.0 EvalA

Evaluates end-to-end speaker-tagged automatic speech recognition (ASR) and speaker diarization on financial domain audio. It probes a model's ability to accurately transcribe speech while correctly assigning speaker identities to utterance segments in multi-speaker conversations. Use when the user wants to benchmark on SPGISpeech 2.0, or asks about evaluating this task. Reports cpWER.

researchpythonapi
0
3
Spgispeech EvalA

Evaluates end-to-end speech-to-text models on financial domain audio, specifically testing their ability to produce fully formatted orthographic transcriptions including punctuation, capitalization, number denormalization, and disfluency handling. The benchmark measures how well acoustic architectures can learn text formatting directly from audio signals without relying on post-processing pipelines. Use when the user wants to benchmark on SPGISpeech, or asks about evaluating this task. Report...

researchpythontesting
0
3
Sphere EvalA

Evaluates spatial reasoning and visual understanding in vision-language models across a hierarchy of tasks, including single-skill perception (position, counting, distance, size), multi-skill integration, and complex physical-world reasoning (object occlusion and manipulation). It specifically probes egocentric vs. allocentric perspective taking and susceptibility to object hallucination. Use when the user wants to benchmark on SPHERE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sphere Exoplanet Detection EvalA

This evaluation probes an algorithm's ability to detect faint exoplanet signals buried in structured stellar speckle noise and accurately characterize their physical properties in direct imaging observations. It measures detection sensitivity across varying false alarm rates and quantifies regression accuracy for astrophysical parameters like flux and sub-pixel position. Use when the user wants to benchmark on SPHERE, or asks about evaluating this task. Reports ARE.

researchpythongo
0
3
Spherical Voronoi Radiance EvalA

This evaluation probes a differentiable spherical Voronoi partition for modeling view-dependent appearance and reflections in 3D Gaussian Splatting. It measures novel-view synthesis reconstruction fidelity and rendering efficiency against established radiance field baselines across synthetic and real-world scenes. Use when the user wants to benchmark on Mip-NeRF360, DeepBlending, Tanks&Temples, NeRF-Synthetic, Ref-NeRF, GlossySynthetic, Ref-Real, or asks about evaluating this task. Reports PSNR.

researchpython
0
3
Sphinx EvalA

Probes visual perception and reasoning capabilities of vision-language models across 25 distinct task types, including symmetry, spatial transformations, chart interpretation, and sequence prediction. Uses a synthetic environment with verifiable ground truth to measure model accuracy against human baselines. Use when the user wants to benchmark on Sphinx, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
SpiNNaker2 Benchmarks EvalA

Evaluates the energy efficiency and computational capability of the SpiNNaker2 processing element architecture across a suite of neuromorphic and deep learning workloads, including classical spiking neural networks, hybrid SNN/DNN frameworks, and standard DNN layers. Use when the user wants to benchmark on SpiNNaker2 Benchmark Suite, or asks about evaluating this task. Reports energy_efficiency.

researchpythongit
0
3
Spice Reasoning EvalA

This evaluation probes a model's ability to solve challenging mathematical and general reasoning tasks, both from standard benchmarks and document-grounded self-play generated questions. It measures how well the model can extract information, perform multi-step logical deduction, and produce verifiable answers across diverse academic and competition-level datasets. Use when the user wants to benchmark on MATH-500, OlympiadBench, Minerva Math, GSM8K, AMC, AIME'24, AIME'25, SuperGPQA, GPQA-Diam...

researchpythongo
0
3
SpiceA

Evaluates the semantic propositional content of image captions by transforming them into scene graphs that encode objects, attributes, and relations. It computes an F-score over these logical propositions to measure how well a generated caption captures the underlying meaning of an image compared to human references. Use when the user has predictions and gold and needs to compute SPICE.

researchpythongo
0
3
Spider 2.0 EvalA

Evaluates language models' ability to perform real-world enterprise text-to-SQL workflows. It probes agentic reasoning, multi-step data transformation, database schema navigation, and SQL dialect adaptation across complex, long-context tasks. Use when the user wants to benchmark on Spider 2.0, Spider 2.0-lite, Spider 2.0-snow, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Spider Cosql EvalA

Evaluates the text-to-SQL and dialog state tracking capabilities of language models by measuring how accurately they generate syntactically and semantically valid SQL queries from natural language questions, with and without constrained auto-regressive decoding. Use when the user wants to benchmark on Spider, CoSQL, or asks about evaluating this task. Reports exact-set-match accuracy.

researchpythongo
0
3
Spider EvalA

Evaluates a model's ability to translate natural language questions into correct SQL queries across diverse database domains. It probes schema linking, lexical matching, and complex query synthesis including joins, aggregations, and subqueries. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports exact matching accuracy.

researchpythongo
0
3