All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,093 views
Vibepass EvalA

Evaluates LLMs on fault-targeted test generation and fault-targeted program repair. It probes discriminative fault detection, fault hypothesis generation, and the ability to debug subtle semantic bugs under diagnostic guidance. Use when the user wants to benchmark on VIBEPASS, or asks about evaluating this task. Reports D_{IO}.

researchpythondebugging
0
3
Vickyage Accents Unplugged EvalA

Compute Vickyage/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Vickyage/accents_unplugged_eval.

developmentpython
0
3
Vicuna Benchmark EvalA

This benchmark assesses general language model capabilities and safety across diverse tasks like Fermi problems, roleplay, and coding. It evaluates how well models balance helpfulness, accuracy, and safety on non-safety-specific queries. Use when the user wants to benchmark on Vicuna_Benchmark, or asks about evaluating this task. Reports Net Win Rate.

researchpythonsecurity
0
3
Vid Ad EvalA

Probes image-level logical anomaly detection under vision-induced distractions such as background changes, blur, and low light. It tests whether models can identify violations of logical constraints (e.g., quantity, length, type, placement) by reasoning over textual descriptions rather than relying on brittle low-level visual features. Use when the user wants to benchmark on VID-AD, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Vidas EvalA

Evaluates a model's ability to assess danger levels in videos by identifying risk elements, understanding context, and assigning severity scores. It probes multimodal perception and risk reasoning capabilities. Use when the user wants to benchmark on ViDAS, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3
Videgothink EvalA

Egocentric video understanding for embodied AI, probing capabilities in video question-answering, hierarchical task planning, visual grounding, and reward modeling. It evaluates how well multimodal models comprehend first-person, action-oriented video contexts required for robotic interaction. Use when the user wants to benchmark on VidEgoThink, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Video Action Classification EvalA

Evaluates a model's ability to recognize and classify human actions in video clips by predicting action categories from sampled frames. It probes temporal dynamics modeling and spatial feature extraction capabilities in video understanding tasks. Use when the user wants to benchmark on Kinetics-400, Something-Something-V2, Epic-Kitchens-100, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Video Analytics EvalA

Evaluates long-video understanding and agentic retrieval capabilities by testing models on temporal grounding, summarization, reasoning, and event causality across ultra-long video streams. Use when the user wants to benchmark on LVBench, VideoMME-Long, Ava-100, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Video Bench EvalA

Evaluates video generation models across two core dimensions: video-condition alignment (how well the generated video matches the text prompt in terms of objects, actions, colors, scenes, and overall consistency) and video quality (technical fidelity, aesthetics, temporal consistency, and motion quality). Use when the user wants to benchmark on Video-Bench, or asks about evaluating this task. Reports Video-text Consistency.

researchpythonaws
0
3
Video Captioning EvalA

Evaluates fine-grained audiovisual captioning quality, attribute-level instruction following, and downstream reasoning capabilities like QA and temporal grounding. Use when the user wants to benchmark on video-SALMONN-2, UGC-VideoCap, VDC, VidCapBench-AE, Daily-Omni, World-Sense, Charades-STA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Video Class Agnostic EvalA

Evaluates a model's ability to segment moving and unknown objects in autonomous driving videos without relying on a closed set of known classes. It probes open-set and motion-based instance segmentation capabilities under varying data distributions and synthetic scenarios. Use when the user wants to benchmark on Cityscapes-VPS, KITTI-MOTS, Carla, or asks about evaluating this task. Reports CAQ, CA-IoU.

researchpythontesting
0
3
Video Mme EvalA

Evaluates long-horizon video understanding and reasoning capabilities of multimodal models on extended video sequences. Use when the user wants to benchmark on Video-MME, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Video Oasis EvalA

This protocol audits video understanding benchmarks to measure genuine spatio-temporal reasoning versus shortcut reliance. It filters out samples solvable without video context and evaluates models under diagnostic conditions (e.g., blind, audio-only, center-frame) to quantify performance degradation and dependency on actual video content. Use when the user wants to benchmark on EgoSchema, ImplicitQA, VSI-Bench, TVBench, VCR-Bench, RTV-Bench, Video-Holmes, MINERVA, MMR-V, VideoMME, MVBench, L...

researchpythongo
0
3
Video Outpainting EvalA

Evaluates a model's ability to generate spatially and temporally consistent video content outside the original frame boundaries (video outpainting), while preserving source structure and visual realism. Use when the user wants to benchmark on DAVIS 2017, YouTube-VOS, or asks about evaluating this task. Reports FVD.

researchpython
0
3
Video Panels EvalA

Evaluates the ability of vision-language models to understand long videos using a training-free visual prompting strategy that combines consecutive frames into multi-frame 'panels'. It probes temporal reasoning, needle-in-a-haystack retrieval, and question-answering capabilities under varying context window constraints. Use when the user wants to benchmark on VideoMME, TimeScope, MLVU, MF2, VNBench, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Video Phy 2 EvalA

Evaluates the ability of text-to-video models to generate physically plausible content by testing adherence to real-world action-centric physical rules, such as conservation of mass/momentum and gravity. It probes whether models understand fundamental physical commonsense beyond superficial motion or visual aesthetics. Use when the user wants to benchmark on VideoPhy-2, or asks about evaluating this task. Reports joint_performance.

researchpythongo
0
3
Video Prediction EvalA

Evaluates the ability of generative models to perform long-horizon open-loop video prediction. It probes how well models maintain temporal consistency, preserve object identities, and adapt to varying scene dynamics across diverse visual domains. Use when the user wants to benchmark on MineRL Navigate, KTH Action, GQN Mazes, Moving MNIST, or asks about evaluating this task. Reports FVD.

researchpython
0
3
Video Qa EvalA

Evaluates video large multimodal models on question-answering tasks across multiple benchmark datasets. Probes the model's ability to understand video content, generate factually accurate long-form responses, and align with language model-derived preferences using direct preference optimization. Use when the user wants to benchmark on MSVD-QA, MSRVTT-QA, TGIF-QA, ActivityNet-QA, VIDAL-QA, WebVid-QA, SSV2-QA, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Video Reality Test EvalA

This benchmark probes the ability of video-language models and humans to distinguish real ASMR videos from AI-generated ones, evaluating perceptual realism and audio-visual consistency. It also measures how effectively video generation models can deceive video understanding models by producing indistinguishable synthetic content. Use when the user wants to benchmark on Video Reality Test, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Video Star EvalA

Evaluates open-vocabulary action recognition by testing a model's ability to generalize to unseen action categories and cross-dataset distributions. It probes fine-grained video understanding and cross-modal reasoning capabilities under base-to-novel and cross-dataset generalization settings. Use when the user wants to benchmark on UCF-101, HMDB-51, Kinetics-400, Kinetics-600, Something-Something V2, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Video Thinking Test EvalA

Evaluates video large language models on their ability to understand complex visual narratives and answer questions correctly. It specifically probes robustness by testing model performance on naturally adversarial or misleading variations of the same video question. Use when the user wants to benchmark on Video Thinking Test, or asks about evaluating this task. Reports Correctness score (accuracy).

researchpythongo
0
3
Video To 4d Mesh EvalA

Evaluates a model's ability to generate temporally consistent, animated 3D meshes from input videos. It probes per-frame geometric reconstruction accuracy, overall 4D sequence fidelity, and motion transfer quality while maintaining topology consistency across frames. Use when the user wants to benchmark on Objaverse, Consistent4D, DAVIS, or asks about evaluating this task. Reports CD-3D.

researchpython
0
3
Video To C EvalA

Evaluates fine-grained video understanding, spatio-temporal reasoning, and hallucination mitigation in multimodal large language models by testing their ability to locate key visual cues and answer questions across diverse video benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, MVBench, TempCompass, VideoMME, VideoHallucer, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Video To Music EvalA

This evaluation protocol assesses a model's ability to generate high-fidelity, diverse instrumental music that is semantically and temporally aligned with a given 10-second video and optional fine-grained text prompt. It probes audio quality, distributional fidelity, generative diversity, and cross-modal alignment using both automated perceptual metrics and human/LLM preference judgments. Use when the user wants to benchmark on ReelBench, LORIS, V2MBench, or asks about evaluating this task. R...

researchpythongo
0
3
Videoaesbench EvalA

Evaluates large multimodal models' ability to perceive and judge video aesthetics across visual form, style, and affectiveness dimensions. It tests performance on diverse video sources (UGC, AIGC, RGC, compression, gaming) using multiple-choice, true/false, and open-ended questions. Use when the user wants to benchmark on VideoAesBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videoconviction EvalA

Evaluates whether LLMs and MLLMs can accurately extract stock tickers, identify explicit investment actions, and quantify human conviction levels from financial influencer videos and transcripts. It probes multimodal reasoning, financial domain understanding, and the ability to filter out noisy or promotional content. Use when the user wants to benchmark on VideoConviction, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Videocraftbench Calvin EvalA

Evaluates a model's ability to learn transferable, long-horizon action dynamics from real-world videos and generate coherent task execution sequences across different environments and robotic setups. Use when the user wants to benchmark on Video-CraftBench, CALVIN, or asks about evaluating this task. Reports Sequential Success Rate (%).

researchpythongo
0
3
Videocube EvalA

Evaluates a model's ability to track arbitrary visual instances across complex, unstructured real-world videos without assuming motion continuity or fixed categories. It measures both local search accuracy and global robustness against challenges like occlusion, fast motion, and scene transitions. Use when the user wants to benchmark on VideoCube, or asks about evaluating this task. Reports PRE.

researchpythongo
0
3
Videodpo EvalA

Evaluates text-to-video diffusion models on visual quality and semantic alignment with input prompts. It measures intra-frame fidelity, aesthetic appeal, and inter-frame temporal consistency using automated benchmarks and human-preference predictors. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench.

researchpython
0
3
Videogamebunny EvalA

Probes vision-language models' ability to understand video game contexts from screenshots, including recognizing actions, characters, UI elements, spatial relationships, and game mechanics. It evaluates how instruction-tuning on game-specific data improves performance compared to larger general-purpose models. Use when the user wants to benchmark on VideoGameBunny Dataset, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Videogameqa Bench EvalA

Evaluates vision-language models on video game quality assurance tasks, including glitch detection, temporal reasoning, and bug reporting. It probes the model's ability to process sampled video frames, identify visual anomalies, and generate structured or descriptive reports about game glitches. Use when the user wants to benchmark on VideoGameQA-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videograin EvalA

Evaluates a video editing model's ability to perform multi-grained edits (class, instance, and part levels) on video-text pairs. It measures semantic alignment, temporal consistency, and pixel-level fidelity of the generated edits. Use when the user wants to benchmark on VideoGrain Evaluation Set, or asks about evaluating this task. Reports CLIP-T.

researchpython
0
3
Videohallucer EvalA

This benchmark evaluates large video-language models for intrinsic and extrinsic hallucinations by presenting paired basic and adversarially modified Yes/No questions about video content. It probes whether models can correctly identify factual content while resisting fabricated or unverifiable details, and measures susceptibility to language bias. Use when the user wants to benchmark on VideoHallucer, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Videollama3 EvalA

Evaluates multimodal foundation models on image and video understanding across multiple dimensions, including document and chart text recognition, mathematical reasoning, multi-image comprehension, general knowledge QA, long-form video comprehension, and temporal reasoning. Use when the user wants to benchmark on ChartQA, DocVQA, MathVista, VideoMME, Charades-STA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videop2r EvalA

Evaluates large video language models on their ability to perceive visual details and perform multi-step reasoning over video content. It measures how well models decompose video understanding into distinct perception and reasoning stages across multiple benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, VCR, MV, TempCom, VideoMME, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videoscore EvalA

Evaluates how well automatic video quality metrics correlate with human ratings across multiple dimensions such as visual quality, temporal consistency, and text alignment. It also measures pairwise preference accuracy to simulate human choice between generated videos. Use when the user wants to benchmark on VideoFeedback-test, GenAI-Bench, VBench, EvalCrafter, or asks about evaluating this task. Reports Spearman's ρ.

researchpythongo
0
3
Videoscore2 EvalA

This evaluation probes a model's ability to assess generative videos across three key dimensions: visual quality, text-to-video alignment, and physical/common-sense consistency. It measures how well automated scoring models align with human judgments on both in-domain and out-of-domain video benchmarks. Use when the user wants to benchmark on VideoGenReward Bench, T2VQA-DB, MJ-Bench-Video, VideoPhy2-test, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vidmuse Video To Music EvalA

Evaluates a model's ability to generate high-fidelity, semantically aligned music conditioned on video input. It probes audio quality, diversity, and cross-modal alignment using statistical distance metrics, beat alignment scores, and subjective human preference tests. Use when the user wants to benchmark on V2M, AIST++, LORIS, TikTok, or asks about evaluating this task. Reports FAD.

researchpythongo
0
3
Vidore EvalA

Evaluates page-level document retrieval on visually rich documents across diverse domains and languages. It probes the model's ability to leverage visual cues, layout, and text within document images without relying on traditional OCR or layout parsing pipelines. Use when the user wants to benchmark on ViDoRe, or asks about evaluating this task. Reports nDCG@5.

researchpython
0
3
Vidore V3 EvalA

This benchmark evaluates end-to-end Retrieval Augmented Generation (RAG) systems on visually rich, real-world documents across multiple professional domains. It probes a model's ability to retrieve relevant pages, generate accurate answers to complex open-ended and multi-hop queries, and precisely ground those answers with bounding boxes in multimodal content. Use when the user wants to benchmark on ViDoRe V3, or asks about evaluating this task. Reports F1 score (Dice coefficient).

researchpythongo
0
3
Vidoseek EvalA

Evaluates a multi-agent RAG framework's ability to retrieve relevant pages from visually rich documents and generate accurate answers through iterative reasoning. It probes hybrid visual-textual retrieval and dynamic token allocation for document comprehension. Use when the user wants to benchmark on ViDoSeek, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Vietmeagent EvalA

This benchmark evaluates a model's ability to generate culturally accurate and coherent explanations for Vietnamese visual question answering. It probes both linguistic fluency and the model's capacity to ground visual evidence in domain-specific cultural knowledge through structured, stepwise reasoning. Use when the user wants to benchmark on Vietnamese VQA dataset, or asks about evaluating this task. Reports Cultural Accuracy.

researchpythongo
0
3
Vietmed Asr EvalA

Evaluates automatic speech recognition (ASR) performance on Vietnamese medical domain audio. It measures how well models transcribe speech containing medical terminology and regional accents, assessing cross-domain transfer capabilities. Use when the user wants to benchmark on VietMed, or asks about evaluating this task. Reports WER.

researchpythongit
0
3
Vietmed Ner EvalA

Evaluates the ability of NER models to identify and classify medically defined entity spans in Vietnamese spoken text. It specifically probes robustness to ASR-generated noise and compares monolingual vs. multilingual, encoder vs. seq2seq architectures. Use when the user wants to benchmark on VietMed-NER, or asks about evaluating this task. Reports micro F1 score.

researchpythongo
0
3
Vietmed Sum EvalA

Evaluates abstractive summarization capabilities on real-world and simulated medical conversations in Vietnamese, testing both human-transcribed and ASR-generated noisy transcripts. Use when the user wants to benchmark on VietMed-Sum, or asks about evaluating this task. Reports ROUGE.

researchpythontesting
0
3
Vietnamese Abusive Span Detection EvalA

Identifies and categorizes abusive content spans within long-form Vietnamese narrative texts. It probes a model's ability to perform sequence labeling for both span detection and fine-grained abuse classification across six distinct categories. Use when the user wants to benchmark on Vietnamese Narrative Abusive Span Dataset, or asks about evaluating this task. Reports Strict F-score.

researchpythongo
0
3
Viewdelta Scd EvalA

This evaluation probes a model's ability to perform scene change detection conditioned on natural language prompts. It measures how well the model distinguishes relevant semantic changes from nuisance variations across diverse domains (street-view, satellite, indoor) and handles viewpoint misalignments. Use when the user wants to benchmark on CSeg, PSCD, SYSU-CD, VL-CMU-CD, or asks about evaluating this task. Reports IoU.

researchpythongo
0
3
Vigil Cognitive Bias EvalA

Evaluates a real-time browser extension's ability to detect and mitigate cognitive bias triggers in online text. It probes span-level identification of persuasive rhetoric and bias patterns, alongside system latency and mitigation quality. Use when the user wants to benchmark on SemEval-2020 Task 11, Moralization Corpus, or asks about evaluating this task. Reports micro-F1.

content-marketingpythongo
0
3
Vihos EvalA

Evaluates a model's ability to detect hate and offensive text spans within Vietnamese social media comments. It probes sequence tagging capabilities, specifically requiring precise boundary identification of offensive content in noisy, informal text. Use when the user wants to benchmark on ViHOS, or asks about evaluating this task. Reports macro-average F1-score.

researchpythongo
0
3
Vikl Mammography EvalA

Evaluates a model's ability to extract robust visual features from single-channel mammography images for binary classification of malignant versus benign breast lumps. It tests cross-dataset generalization and the effectiveness of multimodal contrastive pretraining on pathological classification tasks. Use when the user wants to benchmark on MVKL, CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3