All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,119 views
Visual Wetlandbirds EvalA

This benchmark evaluates deep learning models on fine-grained bird species classification and spatio-temporal behavior recognition in ecological video footage. It probes the model's ability to localize birds, identify their species, and classify their actions across video frames in real-world wetland environments. Use when the user wants to benchmark on Visual WetlandBirds Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
VisualinformationfidelityA

Compute the VisualInformationFidelity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute VisualInformationFidelity, or asks how to score with VisualInformationFidelity.

documentationpython
0
3
Visualoverload EvalA

This benchmark probes fine-grained visual understanding of Vision-Language Models in densely populated, high-resolution scenes. It evaluates capabilities across six core tasks including activity recognition, attribute recognition, counting, OCR, visual reasoning, and global scene classification. Use when the user wants to benchmark on VisualOverload, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visualwebbench EvalA

Evaluates multimodal LLMs' ability to understand web pages and ground UI elements. It probes capabilities across seven subtasks including image captioning, web question answering, OCR, element/action grounding, and action prediction. Use when the user wants to benchmark on VisualWebBench, or asks about evaluating this task. Reports Average Score.

researchpythongo
0
3
Visulogic EvalA

VisuLogic probes vision-centric reasoning in multimodal large language models by presenting problems that require retaining critical visual cues during image description. It eliminates text-based reasoning shortcuts, forcing models to perform genuine visual inference across categories like spatial relations, quantitative shifts, and stylistic details. Use when the user wants to benchmark on VisuLogic, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visuriddles EvalA

Evaluates multimodal large language models' ability to perform abstract visual reasoning across five fine-grained perceptual dimensions (numerosity, attributes, style, position, spatial relations) and two high-level reasoning tasks (analogical pattern matching and constraint-based logic). Use when the user wants to benchmark on VisuRiddles, or asks about evaluating this task. Reports exact match.

researchpythongo
0
3
Vit Robustness EvalA

Evaluates the robustness of Vision Transformer models against input perturbations including adversarial attacks (FGSM/PGD), spatial transformations, and restricted attention. It probes whether ViTs maintain classification performance under distribution shifts and targeted attacks compared to standard CNNs. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Vit Zero Shot Clustering EvalA

Evaluates the ability of Vision Transformer models combined with dimensionality reduction and clustering algorithms to perform zero-shot species-level clustering of animal images. It probes how well unsupervised pipelines can recover ground-truth taxonomic labels and capture intra-specific variation without manual annotation. Use when the user wants to benchmark on Animal Images (Birds & Mammals), or asks about evaluating this task. Reports V-measure.

researchpythongo
0
3
Vitalbench EvalA

Evaluates long-term multivariate time-series forecasting of intraoperative vital signs under three clinically realistic conditions: complete data, variable missingness, and cross-center generalization. It probes a model's ability to handle heterogeneous clinical data, adapt to missing sensor inputs, and generalize across different hospital centers. Use when the user wants to benchmark on VitalDB, MOVER-SIS, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3
Vivd 10m EvalA

Evaluates video editing models on local, entity-level modifications (addition, modification, deletion) by measuring background preservation, text alignment, temporal consistency, and visual quality. Use when the user wants to benchmark on VIVID-10M-Eval, or asks about evaluating this task. Reports Text Alignment (TA).

researchpythonaws
0
3
Vl Compositionality EvalA

Evaluates vision-language models on compositional reasoning capabilities, specifically testing their ability to correctly bind attributes, understand semantic relations, and parse word order in image-text pairs. It also measures systematic generalization to unseen concept combinations and zero-shot classification and retrieval performance. Use when the user wants to benchmark on ARO, CREPE, SVO, VL-Checklist, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vl Jepa Zero Shot BenchmarksA

Evaluates zero-shot video understanding capabilities on action recognition and text-to-video retrieval tasks across diverse benchmarks. Use when the user wants to benchmark on Something-something-v2 (SSv2), EPIC-KITCHENS-100 (EK-100), EgoExo4D Keysteps, Kinetics-400, COIN, CrossTask, MSR-VTT, ActivityNet, DiDeMo, MSVD, YouCook2, PVD-Bench, Dream-1K, VDC-1K, or asks about evaluating this task. Reports top-1 accuracy, recall@1.

researchpythongo
0
3
Vl Rethinker EvalA

Evaluates the multimodal reasoning and self-reflection capabilities of vision-language models across math, multi-discipline, and real-world benchmarks. It probes whether models can correctly interpret visual-textual inputs and produce accurate final answers under greedy decoding. Use when the user wants to benchmark on MathVista, MathVerse, MathVision, MMMU-Pro, MMMU, EMMA, MegaBench, or asks about evaluating this task. Reports Pass@1 accuracy.

researchpythongo
0
3
Vl Rewardbench EvalA

Evaluates vision-language generative reward models (VL-GenRMs) on their ability to judge multimodal response preferences. It specifically probes visual perception, reasoning, and hallucination detection by presenting models with image-text queries and paired candidate responses. Use when the user wants to benchmark on VL-RewardBench, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Vla Cross Embodiment EvalA

Evaluates a vision-language-action model's ability to generalize across diverse robotic embodiments, simulation environments, and real-world platforms. It probes cross-embodiment adaptation, parameter-efficient fine-tuning capabilities, and dexterous manipulation performance. Use when the user wants to benchmark on Libero, Simpler, Calvin, VLABench, RoboTwin-2.0, NAVSIM, BridgeData-v2, Soft-Fold, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Vla Generality Benchmark EvalA

Evaluates the cross-domain generality of vision-language-action models across vision-language understanding, discrete multi-agent control, and continuous robot manipulation. Probes strict format compliance, semantic alignment, action prediction accuracy, and failure modes such as output collapse or modality misalignment. Use when the user wants to benchmark on PIQA, SQA3D, RoboVQA, ODINW, BFCL, Overcooked, Open-X, or asks about evaluating this task. Reports EMR.

ai-agentspythongo
0
3
Vlabench EvalA

Evaluates the generalization, long-horizon reasoning, and language-conditioned manipulation capabilities of Vision-Language-Action (VLA) models, workflow frameworks, and Vision-Language Models (VLMs) in simulated robotic environments. It probes performance across seen/unseen objects, semantic instruction understanding, and composite task decomposition. Use when the user wants to benchmark on VLABench, or asks about evaluating this task. Reports task_progress_score.

researchpythongo
0
3
Vlaser Embodied Reasoning EvalA

This evaluation probes a model's embodied reasoning capabilities, including spatial understanding, visual grounding, task planning, and closed-loop robotic control. It measures how well vision-language models transfer general multimodal knowledge to robot-specific manipulation tasks and identifies the domain gap between internet-scale pretraining and real-world embodiment. Use when the user wants to benchmark on ERQA, Ego-Plan2, Where2place, Pointarena, Paco-Lavis, Pixmo-Points, VSI-Bench, Re...

researchpythongo
0
3
Vlasta Pr AucA

Compute Vlasta/pr_auc via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Vlasta/pr_auc.

developmentpython
0
3
Vlegal Bench EvalA

Evaluates large language models on Vietnamese legal reasoning within a civil law framework. It probes capabilities ranging from statutory recall and hierarchical navigation to multi-step conflict detection, penalty estimation, and ethical bias analysis. Use when the user wants to benchmark on VLegal-Bench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Vlguard EvalA

Evaluates the safety alignment and helpfulness of vision-language models (VLLMs) by measuring their ability to reject harmful image-text prompts while maintaining performance on benign queries. Use when the user wants to benchmark on VLGuard, or asks about evaluating this task. Reports ASR.

researchpythongit
0
3
Vllm EvalA

Evaluates Vietnamese large language models on contextual reasoning, academic knowledge, general trivia, and long-form reading comprehension. Probes both language modeling capability (perplexity) and factual/reasoning accuracy across culturally and linguistically specific tasks. Use when the user wants to benchmark on LAMBADA Vietnamese, Exam Vietnamese, General Knowledge, Comprehension QA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vlm Benchmarks EvalA

Evaluates vision-language models on instruction-following and multimodal reasoning tasks across multiple established benchmarks. Probes capabilities in general VQA, mathematical reasoning, scientific understanding, hallucination detection, and multilingual comprehension. Use when the user wants to benchmark on MMBench, MME, MathVista, HallusionBench, SEEDBench, LLaVABench, ScienceQA, or asks about evaluating this task. Reports evaluation metric.

researchpythongo
0
3
Vlm Deflection Bench EvalA

This benchmark evaluates the ability of large vision-language models to correctly answer knowledge-based visual questions while properly deferring when evidence is missing or hallucinating when faced with noisy or conflicting retrieval contexts. It disentangles parametric memorization from retrieval robustness across four controlled scenarios. Use when the user wants to benchmark on VLM-DeflectionBench, or asks about evaluating this task. Reports Deflection Rate.

ai-agentspythongo
0
3
Vlm Gaussian Noise Robustness EvalA

Evaluates the robustness of Vision-Language Models against Gaussian noise perturbations on input images, measuring both capability degradation (helpfulness, OCR, knowledge) and safety alignment (toxicity, attack success rate) under noisy conditions. It probes whether noise-augmented fine-tuning preserves model utility while mitigating vulnerability to adversarial or distribution-shifted visual inputs. Use when the user wants to benchmark on MM-Vet, RealToxicityPrompts, or asks about evaluatin...

researchpythongo
0
3
Vlm Interaction Reasoning EvalA

Evaluates vision-language models on general visual understanding, spatial/relational reasoning, and specifically interactional reasoning in dynamic scenes using a suite of standard VQA and scene understanding benchmarks. Use when the user wants to benchmark on VQAv2, VizWiz, TextVQA, GQA, VSR, RealWorldQA, MMT-Bench, SEEDBench, A-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vlm Safety EvalA

Evaluates the safety and generalization capabilities of vision-language models by measuring their ability to redirect unsafe content to safe alternatives, maintain zero-shot classification accuracy, and generate safe text/images from unsafe prompts or inputs. Use when the user wants to benchmark on ViSU, NSFWCaps, I2P, NudeNet/SMID/NSFW URLs, Zero-shot Benchmarks (ImageNet variants, Caltech101, Oxford Pets, Flowers102, Stanford Cars, UCF101, DTD), or asks about evaluating this task. Reports %...

researchpythongo
0
3
Vlm Subtlebench EvalA

This benchmark evaluates vision-language models' ability to perform subtle comparative reasoning between pairs of images. It probes capabilities across ten fine-grained difference types, including spatial, temporal, viewpoint, attribute, and existence changes, requiring models to detect and explain nuanced visual discrepancies that are often missed by standard prompting or simple image concatenation. Use when the user wants to benchmark on VLM-SubtleBench, or asks about evaluating this task. ...

ai-agentspythongo
0
3
Vlmbench EvalA

This benchmark evaluates a robot agent's ability to execute 6D manipulation tasks guided by natural language instructions and visual observations. It probes compositional reasoning, object localization, and precise pose estimation in both seen and unseen object settings. Use when the user wants to benchmark on VLMbench, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Vln Ce EvalA

Evaluates an agent's ability to follow natural language instructions to navigate to a target location in a continuous 3D environment. It probes low-level action control, obstacle avoidance, and spatial reasoning without relying on a pre-defined graph topology or oracle localization. Use when the user wants to benchmark on VLN-CE, or asks about evaluating this task. Reports SR, SPL.

researchpythongo
0
3
Vln Task Planning EvalA

Evaluates an agent's ability to decompose coarse-grained natural language navigation instructions into executable subtasks and navigate through simulated environments to reach target locations or interact with objects. It probes task planning, visual-language grounding, and dynamic error recovery in continuous or discrete navigation spaces. Use when the user wants to benchmark on R2R, REVERIE, ALFRED, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Vlsbench EvalA

Evaluates the safety alignment of multimodal large language models (MLLMs) by testing their ability to correctly identify and appropriately respond to unsafe image-text pairs. It specifically probes how well models handle Visual Safety Information Leakage (VSIL), where harmful content might be implicitly revealed in the textual query rather than the image. Use when the user wants to benchmark on VLSBench, or asks about evaluating this task. Reports safety rate (%).

researchpythongo
0
3
Vlue EvalA

Evaluates Vietnamese natural language understanding across five diverse tasks including machine reading comprehension, natural language inference, emotion recognition, hate speech detection, and part-of-speech tagging. It assesses a model's ability to comprehend text, reason over sentence pairs, classify emotions and hate speech, and perform syntactic analysis in Vietnamese. Use when the user wants to benchmark on UIT-ViQuAD 2.0, ViNLI, VSMEC, ViHOS, NIIVTB POS, or asks about evaluating this ...

researchpythongo
0
3
Vmbench EvalA

Evaluates how well text-to-video models generate videos that align with human perceptual preferences across five motion dimensions: object integrity, motion smoothness, commonsense adherence, perceptible amplitude, and temporal coherence. Use when the user wants to benchmark on MMPG-set, or asks about evaluating this task. Reports Spearman correlation.

researchpythongo
0
3
VmeasurescoreA

Compute the VMeasureScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute VMeasureScore, or asks how to score with VMeasureScore.

documentationpython
0
3
Vn Mteb EvalA

Evaluates the quality of Vietnamese text embeddings across six standard information retrieval and NLP tasks. It probes a model's ability to capture semantic similarity, perform document retrieval, classify text, cluster documents, and rank relevant passages in Vietnamese. Use when the user wants to benchmark on VN-MTEB, or asks about evaluating this task. Reports Average Task Score.

researchpythongo
0
3
Vocalbench EvalA

This benchmark evaluates end-to-end speech interaction models across semantic understanding, acoustic quality, conversational fluency, and robustness to noisy or diverse inputs. It probes the model's ability to generate natural, emotionally expressive speech while accurately following instructions and maintaining safety alignment. The evaluation also measures computational efficiency and latency to assess real-time usability. Use when the user wants to benchmark on VocalBench, or asks about e...

researchpythongo
0
3
Vocalbench Zh EvalA

Evaluates Mandarin speech-to-speech conversational agents across semantic understanding, acoustic quality, dialogue management, and robustness. It probes capabilities like cultural context adaptation, emotional empathy, instruction following, and handling of code-switching or noisy inputs. Use when the user wants to benchmark on VocalBench-zh, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Vocalbridge EvalA

Evaluates a latent diffusion purification model's ability to remove adversarial perturbations from voiceprint defenses while preserving speaker identity and perceptual quality. It measures how effectively the purifier restores speaker verification scores and maintains speech naturalness and intelligibility for downstream voice cloning tasks. Use when the user wants to benchmark on LibriSpeech, VCTK, or asks about evaluating this task. Reports ARR.

researchpythongo
0
3
Vocbench EvalA

Evaluates the audio synthesis quality, computational efficiency, and speaker generalization of neural vocoders across autoregressive, GAN-based, and diffusion-based architectures. It probes how well models preserve waveform fidelity and spectrogram structure while balancing inference speed and training complexity. Use when the user wants to benchmark on LJ Speech, LibriTTS, VCTK, or asks about evaluating this task. Reports MOS.

researchpython
0
3
Vocsim EvalA

Evaluates the intrinsic geometric alignment and zero-shot content identity of frozen audio embeddings across diverse single-source audio corpora. It measures how well models can retrieve semantically similar audio clips without task-specific fine-tuning, highlighting generalization gaps on low-resource or out-of-distribution speech. Use when the user wants to benchmark on VocSim, or asks about evaluating this task. Reports GSR.

researchpythonbash
0
3
Voice Accompaniment Separation EvalA

Evaluates a model's ability to separate vocal and accompaniment tracks from mixed music audio. It probes long-term dependency modeling and pattern repetition exploitation in audio source separation. Use when the user wants to benchmark on DSD100, MedleyDB, CCMixer, or asks about evaluating this task. Reports SDR.

researchpython
0
3
Voice Ai Platform EvalA

This benchmark evaluates the quality of commercial voice AI testing platforms across two independent dimensions: simulation quality (how realistically platforms generate test conversations) and evaluation accuracy (how accurately platforms assess conversation quality against human ground truth). It probes whether automated testing systems can reliably replace human quality assurance in high-stakes voice AI deployments. Use when the user wants to benchmark on Custom Voice AI Testing Benchmark,...

researchpythongo
0
3
Voice Cloning Accent EvalA

Evaluates how commercial voice cloning systems preserve speaker identity and speech intelligibility for standard versus accented Mandarin speakers. It probes the alignment between acoustic embedding distances and human perceptual judgments of similarity and intelligibility across accent conditions. Use when the user wants to benchmark on Standard and Accented Mandarin Speech Dataset, or asks about evaluating this task. Reports Intelligibility gain score.

researchpythonapi
0
3
Voice Conversion EvalA

Evaluates non-parallel voice conversion quality by measuring spectral distortion, pitch accuracy, voicing correctness, duration modification capability, and subjective naturalness/speaker similarity. Use when the user wants to benchmark on VCTK, CMU ARCTIC, or asks about evaluating this task. Reports MCD.

researchpython
0
3
Voice Indistinguishability EvalA

Evaluates the privacy guarantee (voice-indistinguishability) and utility of perturbed speech data. It measures how effectively a sanitization framework hides speaker identity while preserving speech recognition accuracy and perceptual naturalness. Use when the user has predictions and gold and needs to compute ACC (Speaker Verification Accuracy).

researchpythongo
0
3
Voice Morph Threshold EvalA

This evaluation probes auditory self-recognition boundaries by measuring how much AI voice morphing a participant can tolerate before they stop recognizing their own voice. It assesses perceptual thresholds, decision latency, and the influence of acoustic embedding distances and demographic factors on voice identity perception. Use when the user wants to benchmark on VoiceMorph Experimental Dataset, or asks about evaluating this task. Reports lowess_T.

researchpython
0
3
Voice Of India EvalA

Evaluates automatic speech recognition (ASR) systems on real-world, unscripted telephonic conversations across 15 Indian languages. It probes geographic, demographic, and audio quality disparities in model performance, particularly focusing on code-mixed speech and natural orthographic variations. Use when the user wants to benchmark on Voice of India, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythonperformance
0
3
Voice Search Wer EvalA

Evaluates speech recognition accuracy and latency trade-offs for streaming vs. non-streaming decoding on a proprietary voice-search dataset. It probes the model's ability to maintain low word error rate while minimizing output delay and computational overhead. Use when the user wants to benchmark on Voice Search, or asks about evaluating this task. Reports WER.

researchpython
0
3
Voiceagentbench EvalA

Evaluates speech language models and ASR-LLM pipelines on agentic speech tasks. It probes single/multi-tool orchestration, multi-turn dialogue, and safety refusal capabilities across multiple languages, including English, Hindi, and five Indic languages. Use when the user wants to benchmark on VoiceAgentBench, or asks about evaluating this task. Reports PF (Parameter Filling).

ai-agentspythongo
0
3