All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,003 views
Vane Bench EvalA

Evaluates the ability of video-language models to detect subtle, rapid, and contextually nuanced anomalies in both real-world surveillance footage and high-fidelity AI-generated videos. It probes fine-grained temporal reasoning, visual grounding, and robustness to synthetic artifacts through a multiple-choice question-answering format. Use when the user wants to benchmark on VANE-Bench, or asks about evaluating this task. Reports MC-Video QA accuracy.

researchpythongo
0
3
Varex EvalA

Evaluates multi-modal structured data extraction from documents, testing a model's ability to parse visual or textual layouts, adhere to a provided JSON schema, and generate compliant structured outputs. It specifically probes schema compliance, layout understanding, and cross-modal robustness across plain text, spatial text, and image inputs. Use when the user wants to benchmark on VAREX, or asks about evaluating this task. Reports exact match (EM).

researchpythongo
0
3
Varta Headline EvalA

Evaluates abstractive headline generation across 15 Indic languages and English. It probes cross-lingual transfer, script normalization effects, and the impact of language-family-specific pretraining on low-resource generation. Use when the user wants to benchmark on Varta, or asks about evaluating this task. Reports ROUGE-L.

researchpythongit
0
3
Vbench++ EvalA

Evaluates the quality and trustworthiness of text-to-video and image-to-video generative models across 16 fine-grained dimensions, including spatial consistency, temporal dynamics, and subject identity. It measures how well automated scores align with human preferences and compares frame-wise generation capabilities against text-to-image baselines. Use when the user wants to benchmark on VBench++, or asks about evaluating this task. Reports VBench score.

researchpythonrust
0
3
Vbench EvalA

Evaluates the visual quality, semantic alignment, temporal consistency, and aesthetic fidelity of text-to-video generation models. It probes the model's ability to produce coherent, high-fidelity videos that match textual prompts across multiple perceptual and technical dimensions. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Total Score.

researchpython
0
3
Vbench T2vA

Evaluates text-to-video generation quality and semantic alignment across dimensions like human action, scene composition, object consistency, and aesthetic quality. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Overall.

researchpythonperformance
0
3
Vbenchcomp EvalA

Evaluates video language models by disentangling question types into LLM-Answerable, Semantic, Temporal, and Others. It isolates true temporal and spatial understanding from language priors and static visual cues by computing accuracy exclusively on the Semantic and Temporal subsets. Use when the user wants to benchmark on LongVideoBench, Egoschema, NextQA, VideoMME, MLVU, LVBench, PerceptionTest, or asks about evaluating this task. Reports VBenchComp score.

researchpythongo
0
3
Vbvr Bench EvalA

Systematic video reasoning capabilities grounded in five cognitive faculties: perception, transformation, spatiality, abstraction, and knowledge. It probes spatiotemporal reasoning, mental manipulation, and rule-based problem solving on video sequences. Use when the user wants to benchmark on VBVR-Dataset, or asks about evaluating this task. Reports rule-based scorer.

researchpythongo
0
3
Vc Ifeval EvalA

Evaluates multimodal large language models' ability to follow instructions containing vision-dependent constraints, such as spatial, stylistic, and structural requirements. It isolates the contribution of visual input to instruction adherence and assesses generalization on standard visual reasoning tasks. Use when the user wants to benchmark on VC-IFEval, MM-IFEval, IFEval, or asks about evaluating this task. Reports instruction-following accuracy.

researchpythongo
0
3
Vc InspectorA

Evaluates the factual accuracy and overall quality of video captions in a reference-free setting. It measures how well a model's predicted quality scores and explanations align with human judgments across diverse video and image domains. Use when the user has predictions and gold and needs to compute Kendall's correlation ($\tau_b$).

researchpythongo
0
3
Vc Mos EvalA

Evaluates one-shot voice conversion quality by measuring how naturally the converted speech sounds and how closely it matches the target speaker's voice compared to human baselines. Use when the user wants to benchmark on VCTK, LibriTTS, or asks about evaluating this task. Reports MOS (Naturalness & Similarity).

researchpython
0
3
Vcb Bench EvalA

VCB Bench evaluates audio-grounded large language models on instruction following with speech-level controls, knowledge reasoning, and robustness under real-world acoustic perturbations. It probes how well models understand and generate spoken responses in Chinese and English using authentic human speech rather than synthetic data. Use when the user wants to benchmark on VCB Bench, or asks about evaluating this task. Reports 1-5 scale score.

researchpythongo
0
3
Vcbench EvalA

Evaluates multimodal mathematical reasoning capabilities of vision-language models, specifically focusing on vision-centric elementary math problems that require explicit visual dependencies across multiple images. It probes spatial, temporal, geometric, logical, and pattern recognition skills to measure how well models integrate cross-modal information for compositional reasoning. Use when the user wants to benchmark on VCBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vcc18 Spoofing EvalA

Evaluates voice conversion systems for processing artifacts by repurposing spoofing countermeasures from automatic speaker verification. It measures how easily a detector can distinguish real speech from converted speech, using Equal Error Rate (EER) as a proxy for artifact quality. Use when the user wants to benchmark on VCC'18, or asks about evaluating this task. Reports Equal Error Rate (EER).

researchpythonperformance
0
3
Vcc2018 EvalA

Evaluates voice conversion systems on speech naturalness and target speaker similarity using crowdsourced perceptual tests, and assesses linguistic consistency via automatic speech recognition word error rates. It covers both parallel (Hub) and non-parallel (Spoke) conversion tasks. Use when the user wants to benchmark on VCC2018, or asks about evaluating this task. Reports Naturalness (MOS).

researchpythonperformance
0
3
Vcc2018 Spoke EvalA

Evaluates non-parallel voice conversion by measuring how effectively a system transfers a source speaker's identity to a target speaker while preserving linguistic content. It probes intra-lingual conversion fidelity and cross-lingual adaptation using subjective human ratings. Use when the user wants to benchmark on VCC2018 (Voice Conversion Challenge 2018), VCTK, or asks about evaluating this task. Reports Quality.

researchpythongo
0
3
Vcc2018 Vc EvalA

Evaluates non-parallel voice conversion by measuring how well a model transforms a source speaker's speech into a target speaker's voice while preserving linguistic content. It assesses both acoustic fidelity using spectral distortion metrics and perceptual quality through human listening tests for naturalness and speaker similarity. Use when the user wants to benchmark on VCC 2018, or asks about evaluating this task. Reports MCD.

researchpythongo
0
3
Vcc2020 Cross Lingual Vc EvalA

Evaluates a model's ability to convert speech across different languages without parallel data, disentangling speaker characteristics from linguistic content. The benchmark probes cross-lingual generalization and non-parallel training capabilities when target speakers only record in foreign languages. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).

researchpython
0
3
Vcc2020 Intra Lingual Vc EvalA

Evaluates a model's ability to convert speech from a source speaker to a target speaker within the same language, leveraging limited parallel data alongside a larger non-parallel corpus. The benchmark probes how well systems can disentangle speaker identity from linguistic content when only a small set of aligned sentences is available for training. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).

researchpython
0
3
Vclimb EvalA

This benchmark evaluates video class incremental learning capabilities, testing a model's ability to sequentially learn new action categories while retaining knowledge of previous tasks using limited episodic memory. It specifically probes how well models handle temporal consistency, frame-level memory selection, and classification on both trimmed and untrimmed video data without catastrophic forgetting. Use when the user wants to benchmark on UCF101, Kinetics, ActivityNet-Trim, ActivityNet-U...

researchpythongo
0
3
Vdialogue EvalA

Evaluates visually-grounded dialogue systems across five core tasks: multi-modal intent prediction, dialog retrieval (text-to-image and image-to-text), dialog state tracking, and response generation. It provides a unified hierarchical scoring metric to compare cross-task performance and generalization. Use when the user wants to benchmark on VisDial, PhotoChat, MMDialog, Image-Chat, or asks about evaluating this task. Reports VDscore.

researchpythongo
0
3
Vdr Compression EvalA

Evaluates multi-vector visual document retrieval (VDR) models under varying compression ratios. It measures how well pruning and merging strategies maintain retrieval accuracy while reducing storage and computational overhead. Use when the user wants to benchmark on ViDoRe-V1, ViDoRe-V2, JinaVDR-Bench, REAL-MM-RAG, ViDoSeek, MMLongBench-Doc, or asks about evaluating this task. Reports nDCG@5.

researchpythongit
0
3
Vector EvalA

Evaluates a model's ability to understand and reason about the temporal order of multiple events in long-form videos. It probes whether the model can correctly sequence events, identify relative ordering, and detect pattern anomalies across varying sequence lengths and difficulties. Use when the user wants to benchmark on VECTOR, or asks about evaluating this task. Reports EM (Exact Match).

researchpythongo
0
3
Vefx Bench EvalA

Evaluates video editing reward models and VLM judges on their ability to align with human preferences across three dimensions: instruction following, rendering quality, and edit exclusivity. It tests both global score correlation and local pairwise preference consistency within candidate sets. Use when the user wants to benchmark on VEFX-Bench, or asks about evaluating this task. Reports SRCC.

researchpythongo
0
3
Vega Hardware BenchA

Measures the performance, energy efficiency, and latency of the Vega SoC on floating-point near-sensor analytic applications (NSAA) and deep neural network (DNN) inference workloads. Use when the user has predictions and gold and needs to compute Energy Efficiency.

researchpythongo
0
3
Velocity Dealiasing EvalA

Evaluates a U-Net model's ability to predict velocity fold numbers and produce dealiased radar velocity fields from folded inputs. It measures both classification accuracy for fold detection and reconstruction fidelity via velocity error metrics. Use when the user wants to benchmark on WSR-88D Level-II/III Radar Data, or asks about evaluating this task. Reports velocity RMSE.

researchpythongo
0
3
Venkatasg GleuA

Compute venkatasg/gleu via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of venkatasg/gleu.

developmentpython
0
3
Venusbench Gd EvalA

This benchmark evaluates GUI grounding capabilities across a hierarchical taxonomy of basic (element, visual, spatial) and advanced (functional, reasoning, refusal) tasks. It probes a model's ability to accurately locate UI elements in screenshots and handle complex, domain-specific, or unanswerable instructions across web, mobile, and desktop platforms. Use when the user wants to benchmark on VenusBench-GD, or asks about evaluating this task. Reports accuracy.

researchpython
0
3
Vera EvalA

Evaluates the reasoning capabilities of voice and multimodal models under real-time streaming constraints, quantifying the performance gap between text and voice modalities on tasks with well-defined ground truth. Use when the user wants to benchmark on VERA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Verafi Financial Qa EvalA

Probes an agentic RAG system's ability to retrieve relevant SEC filings and generate factually correct, complete financial answers. It specifically tests the impact of neurosymbolic policy validation on suppressing hallucinations and mathematical errors in high-stakes financial domains. Use when the user wants to benchmark on FinanceBench-style Financial QA Dataset, or asks about evaluating this task. Reports Factual Correctness.

ai-agentspythondatabase
0
3
Veriequivbench EvalA

Evaluates an LLM's ability to generate formally verifiable code that aligns with natural language problem descriptions and passes unit tests. It probes complex algorithmic reasoning and code-specification alignment without requiring manual ground-truth specifications. Use when the user wants to benchmark on VeriEquivBench, or asks about evaluating this task. Reports equivalence_score.

documentationpythongo
0
3
Verifact EvalA

Evaluates the factual correctness of long-form LLM-generated responses by decomposing them into atomic facts, detecting and refining incomplete or missing information, and verifying each fact against external web evidence. Use when the user wants to benchmark on Long-form LLM responses, or asks about evaluating this task. Reports Supported/Contradicted/Undecided classification accuracy.

researchpythongo
0
3
Verisoftbench EvalA

This benchmark probes an AI system's ability to perform repository-scale formal verification in Lean 4. It specifically tests context-aware proof automation, measuring how well models handle project-specific abstractions and transitive dependency closures beyond standard mathematical libraries. Use when the user wants to benchmark on VeriSoftBench-Full, or asks about evaluating this task. Reports solve_rate.

researchpythongo
0
3
Verite EvalA

Evaluates multimodal misinformation detection models on real-world and synthetic image-caption pairs, specifically probing their ability to distinguish truthful content from out-of-context (OOC) and miscaptioned (MC) misinformation while measuring susceptibility to unimodal bias. Use when the user wants to benchmark on VERITE, COSMOS, VMU-Twitter, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Veroeval EvalA

Evaluates general visual reasoning capabilities across a diverse set of 30 benchmarks spanning six task categories, including chart/OCR, STEM, spatial/action, knowledge/recognition, grounding, and captioning/instruction following. Use when the user wants to benchmark on VeroEval, or asks about evaluating this task. Reports overall averages.

researchpythongo
0
3
Versebench EvalA

Evaluates a unified multimodal model's ability to generate synchronized audio and video from text, phonemes, and reference media. It probes zero-shot voice cloning fidelity, lip-sync accuracy, acoustic quality, and cross-modal temporal alignment. Use when the user wants to benchmark on VerseBench, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Vertaix VendiscoreA

Compute Vertaix/vendiscore via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Vertaix/vendiscore.

developmentpython
0
3
Veu Bench EvalA

VEU-Bench evaluates a model's ability to understand video editing by probing 19 fine-grained tasks across 10 dimensions (e.g., shot size, cut types, transitions) and three cognitive stages: recognition, reasoning, and judging. It tests whether models can identify editing components, infer their functions, and judge their effects in video content. Use when the user wants to benchmark on VEU-Bench, or asks about evaluating this task. Reports Scoreall.

researchpythongo
0
3
Vevo Voice Imitation EvalA

This benchmark evaluates a model's ability to perform zero-shot voice imitation by disentangling linguistic content, speaker timbre, and vocal style (accent/emotion). It probes the model's capacity to generate high-intelligibility speech that accurately transfers the target speaker's identity and stylistic attributes from a reference clip without task-specific fine-tuning. Use when the user wants to benchmark on Vevo Evaluation Set (AB, CV, ACCENT, EMOTION), or asks about evaluating this task...

researchpythongo
0
3
Vg Cot EvalA

Evaluates the visual reasoning and grounding capabilities of Large Vision-Language Models (LVLMs) by measuring the quality of their step-by-step rationales, the accuracy of their final answers, and the alignment between the generated reasoning and the prediction. Use when the user wants to benchmark on VG-CoT, or asks about evaluating this task. Reports Rationale Quality (RQ), Answer Accuracy (AA), Reasoning-Answer Alignment (RAA).

researchpythonrust
0
3
Vga Bench EvalA

Evaluates text-to-video generation models across three dimensions: aesthetic quality, aesthetic tagging, and generation fidelity. It measures how well automated neural assessors align with human judgments and ranks models based on normalized scores across multiple visual and formal dimensions. Use when the user wants to benchmark on VGA-Bench, or asks about evaluating this task. Reports five-class accuracy.

researchpythongo
0
3
Vga Gui Comprehension EvalA

Evaluates a vision-language model's ability to understand graphical user interfaces (GUIs) and answer user questions based on visual content. It specifically probes the model's capacity to avoid hallucinations by grounding responses in actual GUI elements rather than relying solely on textual priors. Use when the user wants to benchmark on GUI Comprehension Bench, or asks about evaluating this task. Reports GPT evaluation score.

researchpythongo
0
3
Vggsound Continual EvalA

Evaluates a model's ability to perform continual audio-visual classification across sequential tasks without catastrophic forgetting. It measures how well the model retains performance on previously learned categories while learning new ones, across audio, visual, and cross-modal fusion settings. Use when the user wants to benchmark on VGGSound-Instruments, VGGSound-100, VGG-Sound Source, or asks about evaluating this task. Reports Average accuracy.

researchpythongo
0
3
Vggsound Sep EvalA

Evaluates zero-shot language-queried audio source separation on human actions, sound-emitting objects, and human-object interactions. The benchmark tests isolation of a target sound from a mixed audio mixture using text labels. Use when the user wants to benchmark on VGGSound, or asks about evaluating this task. Reports SDRi.

researchpythongo
0
3
Vggsounder EvalA

Evaluates audio-visual foundation and embedding models on multi-label video classification, probing their ability to recognize sound and visual events across different input modalities. It specifically measures modality alignment, unimodal versus multimodal performance, and susceptibility to distraction from irrelevant background audio or static visuals. Use when the user wants to benchmark on VGGSounder, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Vgphrasecut EvalA

This benchmark evaluates language-based image segmentation by requiring models to ground natural language phrases into precise image regions. It probes a model's ability to handle long-tail categories, attributes, relationships, and varying object sizes in open-vocabulary settings. Use when the user wants to benchmark on VGPhraseCut, or asks about evaluating this task. Reports mean-IoU.

researchpythongo
0
3
Vhd11k EvalA

Evaluates multimodal models' ability to detect harmful content in images and videos across ten specific harmful categories and a general unharmful class. It probes binary classification robustness against dataset imbalance and multi-class reasoning capabilities under varying prompt conditions. Use when the user wants to benchmark on VHD11K, SMID, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vhelm EvalA

Holistic evaluation of vision-language models across multiple dimensions including visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. Use when the user wants to benchmark on VHELM Scenarios, or asks about evaluating this task. Reports scenario_score.

researchpythongit
0
3
Vib Probe Hallucination EvalA

Evaluates the ability of Vision-Language Models to avoid generating unfaithful or non-existent visual details in both closed-set discriminative QA and open-ended generative captioning. It also measures how well an auxiliary probing framework can detect these hallucinations via attention dynamics and mitigate them at inference time without retraining. Use when the user wants to benchmark on POPE, AMBER, M-HalDetect, COCO-Caption, or asks about evaluating this task. Reports AUPRC.

researchpythongo
0
3
Vibe Eval EvalA

Probes multimodal reasoning and visual understanding on real-world images. Specifically designed with a 'hard' subset of prompts that are unsolvable by current frontier models to measure genuine performance gaps and contamination-free generalization. Use when the user wants to benchmark on Vibe-Eval, or asks about evaluating this task. Reports Vibe-Eval Score.

researchpythongit
0
3