All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs1,992 views
Atbench EvalA

Evaluates vision-language models on five assistive technology tasks for people with visual impairments: panoptic segmentation, depth estimation, optical character recognition, image captioning, and visual question answering. It probes the model's ability to unify multiple visual understanding and generation tasks within a single parameter set using task-specific prompts. Use when the user wants to benchmark on ADE-150, NYU-V2, OCR (IC13, IC15, SVT, IIIT5K, SVTP, CUTE), VizWiz_Cap, VizWiz_VQA,...

researchpythongo
0
3
Atc Asr Csd EvalA

Evaluates automatic speech recognition (ASR) and call sign detection (CSD) capabilities on real-world air traffic control audio. It probes the model's ability to transcribe accented, noisy pilot and controller speech and accurately identify aircraft call signs in domain-specific phraseology. Use when the user wants to benchmark on Airbus ATC Speech Recognition 2018 Challenge Corpus, or asks about evaluating this task. Reports WER.

researchpythonapi
0
3
Atc Asr Domain Shift EvalA

This benchmark evaluates the robustness of self-supervised speech recognition models under domain shift in air traffic control communications. It probes few-shot fine-tuning capabilities, sensitivity to audio quality and accents, and potential gender bias in transcription performance. Use when the user wants to benchmark on NATS, ISAVIA, LiveATC-Test, ATCO2-Test, LDC-ATCC, UWB-ATCC, ATCOSIM, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Atg Pvd EvalA

Evaluates a drone-based suspect-and-investigate system for detecting and classifying illegally parked cars, moving cars, and legally parked cars from aerial imagery. Use when the user wants to benchmark on ATG-PVD, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Athletics Anomaly Detection EvalA

Evaluates performance-based anomaly detection methods to identify athletes with confirmed anti-doping rule violations. It measures how well different algorithms surface sanctioned athletes while balancing precision and recall, accounting for environmental factors like wind and altitude. Use when the user wants to benchmark on 100 m Sprint, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Atlas Benchmark EvalA

This benchmark evaluates human motion prediction algorithms by testing their ability to forecast future trajectories given past observations and environmental context. It systematically probes robustness to perception noise, generalization across different social and cultural environments, and performance under varying observation and prediction horizons. Use when the user wants to benchmark on ETH, ATC, THÖR, or asks about evaluating this task. Reports ADE.

researchpythongo
0
3
Atlas Chat EvalA

Evaluates large language models on Moroccan Arabic (Darija) across multiple-choice reasoning, instruction following, translation, summarization, and sentiment analysis. It probes dialect-specific linguistic features, script variability (Arabic vs. Arabizi), and real-world instruction-following capabilities in a low-resource setting. Use when the user wants to benchmark on DarijaMMLU, DarijaHellaSwag, Belebele_Ary, DarijaBench, DarijaAlpacaEval, or asks about evaluating this task. Reports Accu...

researchpythongo
0
3
Atlas EvalA

This benchmark probes frontier scientific reasoning across multiple disciplines (e.g., physics, chemistry, biology, computer science, mathematics) using original, multi-step problems. It evaluates a model's ability to generate complex, open-ended, LaTeX-formatted answers and assesses both solution accuracy and inference stability across multiple sampling runs. Use when the user wants to benchmark on ATLAS, or asks about evaluating this task. Reports Accuracy.

researchpythongit
0
3
Atm Bench EvalA

Evaluates long-term personalized referential memory QA by testing a model's ability to retrieve and reason over multi-source, multimodal personal data spanning years. It probes conflict-aware aggregation, temporal-visual grounding, and accurate reference resolution across different question types. Use when the user wants to benchmark on ATM-Bench, or asks about evaluating this task. Reports QS.

ai-agentspythongo
0
3
Atmospheric Gap Imputation EvalA

Evaluates the ability of machine learning models to reconstruct missing multivariate atmospheric data across time and altitude. It probes spatiotemporal continuity, physical gradient preservation, and performance under varying gap lengths (short, medium, long). Use when the user wants to benchmark on SD-WACCM-X synthetic atmospheric data, or asks about evaluating this task. Reports Pearson correlation R.

datapythonperformance
0
3
Atmospheric Subgrid EvalA

Evaluates neural network parameterizations for predicting subgrid atmospheric processes (e.g., microphysical tendencies, momentum fluxes) using single-column versus non-local (3x3 grid) inputs. It probes the model's ability to capture mesoscale convective dynamics and frontal systems, and assesses performance across different atmospheric stability regimes. Use when the user wants to benchmark on SAM (Simple Atmospheric Model) simulation, or asks about evaluating this task. Reports R^2.

researchpythontesting
0
3
Atrbench EvalA

Evaluates SAR automatic target recognition (ATR) capabilities under realistic, wild conditions. It probes fine-grained vehicle classification and detection robustness across varying imaging geometries, scene complexities, and domain shifts (SOC vs EOC settings). Use when the user wants to benchmark on ATRBench, or asks about evaluating this task. Reports overall accuracy (%).

researchpythongo
0
3
Atrw Reid EvalA

Evaluates Amur tiger re-identification in the wild by measuring how well models can match tiger identities across different camera views and detection/pose conditions. It probes robustness to non-rigid body deformation, extreme pose variation, and domain shifts between controlled (plain) and uncontrolled (wild) environments. Use when the user wants to benchmark on ATRW, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Attackviz EvalA

This benchmark probes the robustness of multimodal large language models (MLLMs) to maliciously manipulated chart visualizations. It measures how well models answer chart-based questions when presented with data-consistent but misleading charts compared to correct charts, and quantifies the rate at which models switch from correct to incorrect answers due to visual misleaders. Use when the user wants to benchmark on AttackViz, or asks about evaluating this task. Reports relaxed accuracy.

datapythongo
0
3
Atten Mixer EvalA

Evaluates a session-based recommendation model's ability to predict the next item in a user's browsing session. It probes the model's capacity to capture multi-level user intent and item semantics through attention mechanisms while handling session-specific inductive biases. Use when the user wants to benchmark on Diginetica, Gowalla, Last.fm, or asks about evaluating this task. Reports HR@K, MRR@K.

researchpythongo
0
3
Attention Pruning EvalA

Evaluates the effectiveness of attention head pruning methods in reducing gender bias in large language models while preserving language model utility. It probes the trade-off between fairness mitigation and general language modeling performance across models of varying sizes. Use when the user wants to benchmark on HolisticBias, WikiText-2, or asks about evaluating this task. Reports HolisticBias.

researchpythongo
0
3
Attentionddi EvalA

Predicts whether a pair of drugs interacts based on multi-modal similarity features (targets, pathways, side effects, chemical structure, etc.). Evaluates performance on highly imbalanced drug-drug interaction datasets using precision-recall metrics. Use when the user wants to benchmark on DS1, DS2, DS3 (CYP), DS3 (NCYP), or asks about evaluating this task. Reports AUPR.

researchpythonperformance
0
3
Attentionspan EvalA

Evaluates algorithmic reasoning and out-of-distribution generalization in Transformers by measuring prediction accuracy and attention pattern alignment against ground-truth reference masks on synthetic tasks. Use when the user wants to benchmark on AttentionSpan, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Attestable Audits EvalA

Evaluates the feasibility and performance of running standard AI safety benchmarks inside Trusted Execution Environments (TEEs) using quantized models. It probes zero-shot reasoning accuracy, toxicity refusal capabilities, and text generation quality under hardware and cryptographic constraints. Use when the user wants to benchmark on MMLU, ToxicChat, Summarization, or asks about evaluating this task. Reports MMLU Accuracy (%).

researchpythonrust
0
3
Attrgau EvalA

Evaluates session-based recommendation models enhanced with the AttrGAU framework on their ability to predict the next item in a user session. It probes robustness to data sparsity and noisy interactions, and measures the model-agnostic performance gain over vanilla backbones. Use when the user wants to benchmark on Dressipi, Diginetica, Retailrocket, or asks about evaluating this task. Reports HR@N, MRR@N.

researchpythonperformance
0
3
Attribution EvalA

Evaluates the faithfulness and robustness of perturbation-based image attribution methods by measuring how well generated heatmaps localize objects, predict probability drops upon feature removal, and maintain consistency under hyperparameter variations. Use when the user wants to benchmark on ImageNet, Places365, or asks about evaluating this task. Reports deletion metric.

researchpythonperformance
0
3
Attrseg Ovss EvalA

This evaluation probes a model's ability to perform open-vocabulary semantic segmentation using either direct class names or decomposed attribute descriptions. It specifically tests robustness to textual ambiguity, neologisms, and unnameable categories by measuring pixel-level alignment across standard and novel datasets. Use when the user wants to benchmark on PASCAL-5i, COCO-20i, PASCAL VOC, PASCAL Context, Fantastic Beasts, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Audeeter EvalA

This benchmark probes the ability of deepfake audio detectors to generalise to novel synthesis systems and diverse human voice corpora in open-world settings. It specifically evaluates robustness against domain shifts in both speech synthesis methods and real audio sources, revealing how well models handle unseen acoustic patterns and distribution shifts. Use when the user wants to benchmark on AUDETER, or asks about evaluating this task. Reports Equal Error Rate (EER).

researchpythongit
0
3
Audio Compositionality EvalA

Probes whether audio encoders preserve algebraic consistency when identical sources are added to different base scenes (A-COAT), and whether representations can be accurately reconstructed from discrete attribute-level primitives like timbre, pitch, rate, and amplitude (A-TRE). Use when the user wants to benchmark on Synthetic Audio Scenes, or asks about evaluating this task. Reports A-COAT.

researchpythongit
0
3
Audio Crowd Mt EvalA

This protocol evaluates machine translation quality by comparing crowd-sourced human judgments of text-only outputs versus multimodal (text + audio) outputs. It probes whether audio-based assessments improve inter-rater consistency and reveal system-level differences through prosodic and expressive features unavailable in text. Use when the user wants to benchmark on WMT German-English, or asks about evaluating this task. Reports standardized score.

researchpythonexpress
0
3
Audio Deepfake Detection EvalA

Evaluates pretrained audio deepfake detectors across 28 diverse datasets to measure robustness against different manipulation types, generation methods, and real-world conditions like in-the-wild noise and perturbations. The protocol standardizes audio preprocessing and label formats to enable fair cross-dataset comparison and highlights generalization gaps when lab-trained models face advanced generation techniques. Use when the user wants to benchmark on ASVspoof2019_LA, MLAAD-v5, In-the-wi...

researchpython
0
3
Audio Deepfake Generalization EvalA

This benchmark evaluates the generalization capability of audio deepfake detection models by testing their performance on controlled, studio-recorded spoofing data versus real-world, uncontrolled in-the-wild audio. It probes whether models trained on standard lab benchmarks can robustly distinguish real from synthetic speech in practical deployment scenarios. Use when the user wants to benchmark on ASVspoof 2019 LA, In-the-Wild Data, or asks about evaluating this task. Reports EER.

researchpythontesting
0
3
Audio Loop Gen EvalA

Evaluates the quality, diversity, and realism of generated audio drum loops. It probes a model's ability to capture spectral-temporal patterns, genre characteristics, and seamless looping properties in fixed-length music generation. Use when the user wants to benchmark on FreeSound Loop Dataset (FSLD), or asks about evaluating this task. Reports IS.

researchpythongit
0
3
Audio Source Separation EvalA

Evaluates an audio-only model's ability to disentangle and reconstruct individual sound sources from mixed audio using semantic category guidance, without relying on visual cues. Use when the user wants to benchmark on MUSIC, FUSS, MUSDB18, VGG-Sound, or asks about evaluating this task. Reports SDR.

researchpythongo
0
3
Audio Spoof Detection EvalA

Evaluates the robustness of audio spoof detection models against real-world audio degradation and manipulation attacks (laundering), including reverberation, additive noise, and re-compression. Use when the user wants to benchmark on ASVspoof 2019 LA, ASVspoof Laundered Database, or asks about evaluating this task. Reports EER.

researchpythondatabase
0
3
Audio To Image EvalA

Evaluates an audio-to-image generative model's ability to synthesize semantically aligned images from audio prompts. It measures cross-modal alignment, perceptual image quality, and distributional similarity against ground-truth visuals. Use when the user wants to benchmark on Greatest Hits, Landscapes, Into The Wild (ITW), VEGAS, VGGSound, or asks about evaluating this task. Reports Fréchet Inception Distance (FID).

researchpythongo
0
3
Audio Turing Test EvalA

Evaluates the human-likeness of Chinese text-to-speech systems using a Turing-test-inspired protocol where human listeners classify audio as human, unclear, or machine. It also benchmarks an automatic LLM-based evaluator against human judgments and traditional MOS prediction models to measure alignment and trap-item detection capability. Use when the user wants to benchmark on ATT-Corpus, or asks about evaluating this task. Reports HLS.

researchpythongo
0
3
Audio Visual Gfsl EvalA

Evaluates audio-visual few-shot video classification by measuring how well models generalize to novel classes with limited training examples (1, 5, 10-shot). It reports both generalised few-shot learning (HM) and standard few-shot learning (FSL) accuracy to assess bias towards base classes. Use when the user wants to benchmark on VGGSound-FSL, UCF-FSL, ActivityNet-FSL, or asks about evaluating this task. Reports HM.

researchpythongo
0
3
Audio Visual Navigation EvalA

Probes an embodied agent's ability to navigate unmapped 3D environments using fused audio-visual observations to locate both static and moving sound sources. It tests generalization to unseen environments and unheard audio distributions under clean and noisy conditions. Use when the user wants to benchmark on Replica, Matterport3D, or asks about evaluating this task. Reports Success rate (SR).

researchpythongo
0
3
Audio Visual Separation EvalA

This protocol evaluates an audio-visual model's ability to isolate and separate target instrument sounds from mixed multi-source video audio using visual object cues. It quantifies separation accuracy and artifact suppression across held-out test clips and synthetically mixed pairs. The evaluation also probes generalization to unseen object combinations and visually-guided denoising on real-world videos. Use when the user wants to benchmark on MUSIC, AudioSet-Unlabeled, AudioSet-SingleSource,...

researchpython
0
3
Audiocaps Sep EvalA

Evaluates zero-shot language-queried audio source separation using natural language captions rather than fixed labels. The benchmark tests the model's ability to separate a target sound described by human-annotated captions from a mixed audio mixture. Use when the user wants to benchmark on AudioCaps, or asks about evaluating this task. Reports SDRi.

researchpythongo
0
3
Audiocrag EvalA

Evaluates speech-in-speech-out dialogue systems on their ability to accurately answer spoken queries using external tools, measuring both answer correctness and system latency under streaming versus open-book settings. Use when the user wants to benchmark on AudioCRAG, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Audiomarathon EvalA

Evaluates long-context audio understanding and inference efficiency across speech, sound, and music domains. It probes temporal dependency modeling, multi-hop reasoning, and memory/token-pruning scalability in Large Audio Language Models. Use when the user wants to benchmark on AudioMarathon, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Audiomarkbench EvalA

Evaluates the robustness of audio watermarking detectors against watermark removal and forgery under no-box, black-box, and white-box threat models. It also measures the perceptual and signal quality of perturbed audio to ensure attacks do not excessively degrade the host signal. Use when the user wants to benchmark on AudioMarkData, LibriSpeech, or asks about evaluating this task. Reports FNR, FPR.

researchpythongo
0
3
Audiomnist EvalA

Evaluates audio classification performance on spoken digits and speaker sex using raw waveforms and spectrograms, serving as a benchmark for explainable AI (XAI) methods in the audio domain. Use when the user wants to benchmark on AudioMNIST, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Audiomotionbench EvalA

Evaluates large audio-language models' ability to perceive and reason about spatial motion in binaural audio. It probes whether models can correctly infer motion direction and trajectories from interaural cues, rather than relying on linguistic or spectral heuristics. Use when the user wants to benchmark on AudioMotionBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Audiorole EvalA

Evaluates a model's ability to generate audio responses that simultaneously maintain contextual semantic appropriateness and preserve the specific acoustic identity (timbre, prosody) of a target character given reference audio. Use when the user wants to benchmark on AudioRole-Demo, or asks about evaluating this task. Reports Acoustic Personalization (AP).

ai-agentspythonperformance
0
3
Audiosafetybench EvalA

Evaluates audio safety guardrails on detecting both audio-native risks (e.g., harmful sound events, voice attributes) and semantic content risks (e.g., jailbreaks, policy violations). It measures joint accuracy where both risk types must be correctly classified, alongside end-to-end inference latency. Use when the user wants to benchmark on AudioSafetyBench, Jailbreak-AudioBench, Nemotron-Content-Safety-Audio, Omni-SafetyBench, AdvWave, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Audioset Sep EvalA

Evaluates zero-shot language-queried audio source separation on a diverse set of environmental and acoustic sound classes. The benchmark tests the model's ability to isolate a target sound described by a text label from a synthetically mixed audio mixture. Use when the user wants to benchmark on AudioSet, or asks about evaluating this task. Reports SDRi.

researchpythongo
0
3
Aunp Mechanism EvalA

Evaluates whether large language models can correctly reason about physicochemical mechanisms in gold nanoparticle synthesis using multiple-choice questions. It probes both factual recall and the depth of mechanistic understanding by measuring prediction accuracy and model confidence derived from output logits. Use when the user wants to benchmark on AuNP Synthesis Mechanism Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
AurocA

Compute the AUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AUROC, or asks how to score with AUROC.

documentationpythonperformance
0
3
Aurora Weather Extremes EvalA

Evaluates the predictability and forecast skill of the Aurora AI weather model across selected extreme weather events, including tropical cyclones, winter freezes, and heatwaves. It probes the model's ability to maintain deterministic track accuracy, temperature amplitude, and spatial pattern fidelity across short-range (1–7 day) to subseasonal (14–21 day) lead times. Use when the user wants to benchmark on Selected Weather Extremes Case Studies, or asks about evaluating this task. Reports Tr...

researchpythongo
0
3
Auslaw Citation EvalA

Evaluates an LLM's ability to predict the correct legal citation for a given query text. It probes the model's capacity to retrieve or generate accurate references from a large Australian legal corpus, testing both retrieval and generation capabilities in a domain-specific setting. Use when the user wants to benchmark on AusLaw Citation Benchmark, or asks about evaluating this task. Reports ACC@1.

researchpythongo
0
3
Authenhallu EvalA

Evaluates an LLM's capability to detect and categorize hallucinations in authentic, real-world human-LLM dialogues. It specifically probes whether models can identify input-conflicting, context-conflicting, and fact-conflicting errors in query-response pairs. Use when the user wants to benchmark on AuthenHallu, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Authfix Oidc Repair EvalA

Evaluates an automated program repair system's ability to fix real-world security and logic bugs in OpenID Connect implementations. It measures both the rate of successfully generating correct patches and the semantic quality of those patches compared to human-written developer fixes. Use when the user wants to benchmark on OpenID Connect Bug Dataset, or asks about evaluating this task. Reports fix_accuracy.

researchpythongit
0
3