All authors
feiyang-k avatar

Claude Skills by feiyang-k

github.com/feiyang-k
258 skillsA× 257B× 10 installs89 views
Pali A Jointly Scaled Multilingual Language Image Model Arxiv 2209 06794v4A

Use this skill when you want to train a multilingual vision-language model using web-scale multilingual image-text data. Avoid it when you only need English language VLM capabilities.

devopsgoperformance
0
9
Perception Test A Diagnostic Benchmark For Multimodal Video Models Arxiv 2305 13520v3A

Use this skill when you need a diagnostic benchmark testing fine-grained perception in video models across memory, abstraction, physics, and semantics. Avoid it when standard video QA benchmarks are sufficient.

testinggotesting
0
9
Probing Multimodal Llms As World Models For Driving Arxiv 2405 05956v2A

Use this skill when you need a benchmark testing VLMs as world models for autonomous driving understanding. Avoid it when driving is not your domain.

testinggotesting
0
9
Redcaps Web Curated Image Text Data Created By The People For The People Arxiv 2111 11431v1A

Use this skill when you want to curate image-caption data from Reddit with transparent provenance and community-driven quality. Avoid it when you need professionally written captions or cannot comply with Reddit's terms of use.

datagoapi
0
9
Referring Expression Comprehension A Survey Of Methods And Datasets Arxiv 2007 09554v2A

Use this skill when you need an overview of referring expression datasets (RefCOCO, RefCOCO+, RefCOCOg) for VLM grounding training. Avoid it when you already know the referring expression dataset landscape.

code-qualitygoexpress
0
9
Sa 1b Segment Anything 1 Billion Masks Dataset Arxiv Sa1b 2023A

Use this skill when you need the largest segmentation mask dataset (1.1B masks) for training or augmenting VLM grounding data. Avoid it when existing segmentation datasets are sufficient.

testinggo
0
9
Sbu Captions Dataset Crossref Nips 2011 SbuA

Use this skill when you want to mine naturally occurring image-caption pairs from Flickr for vision-language pretraining. Avoid it when you need larger scale data or more descriptive captions than Flickr provides.

devopsgoapi
0
9
Scaling Up Visual And Vision Language Representation Learning With Noisy Text Supervision Arxiv 2102 05918v2A

Use this skill when you have a massive noisy image-text dataset and want to train a dual-encoder model that is robust to caption noise. Avoid it when you need clean, curated data or your dataset is too small for noisy supervision to work.

ai-agentsgoperformance
0
9
Scannet Richly Annotated 3d Reconstructions Of Indoor Scenes Arxiv 1702 04405v2A

Use this skill when you need a richly annotated 3D indoor scene dataset for 3D vision-language understanding research. Avoid it when you do not work with 3D scene understanding.

researchgoapi
0
9
Scienceqa Learn To Explain Multimodal Reasoning Via Thought Chains For Science Question Answering Arxiv 2209 09513v2A

Use this skill when you need a multimodal science QA dataset with detailed explanations and chain-of-thought reasoning for training VLMs. Avoid it when you do not need science reasoning or chain-of-thought data.

code-qualitygoperformance
0
9
Seed Bench Benchmarking Multimodal Llms With Generative Understanding Arxiv 2307 16125v2A

Use this skill when you need a large-scale benchmark testing VLM generative understanding across 12 evaluation dimensions. Avoid it when you only need perception evaluation or have simpler benchmarks.

testinggotesting
0
9
Segment Anything Arxiv 2304 02643v1A

Use this skill when you need to build a billion-scale segmentation mask dataset using a model-in-the-loop interactive annotation approach. Avoid it when you need text-grounded segmentation rather than prompt-based mask generation.

researchgo
0
9
Snli Ve Visual Entailment Dataset Arxiv 1901 06706v1A

Use this skill when you need a visual entailment dataset where the model must determine if a text hypothesis is entailed, contradicted, or neutral with respect to an image premise. Avoid it when you do not need visual entailment or NLI-style evaluation.

researchgoperformance
0
9
Something Something V2 A Large Scale Video Understanding Dataset Arxiv 1706 04261v2A

Use this skill when you need a video dataset requiring temporal understanding of human actions for training video-language models. Avoid it when you do not need temporal action understanding.

testinggotesting
0
9
Textvqa Towards Reasoning About Text In Images Arxiv 1904 08920v2A

Use this skill when you need a VQA dataset requiring models to read and reason about text visible in images. Avoid it when your VQA task does not involve reading text in images.

testinggoperformance
0
9
The Pile An 800gb Dataset Of Diverse Text For Language Modeling Arxiv 2101 00027v1A

Use this skill when you want to construct a large diverse text corpus from 22 curated sources for LLM pretraining that is also used as the text component in multimodal training. Avoid it when you only need image-text data without a standalone text corpus.

datagogit
0
9
Vatex A Large Scale High Quality Multilingual Dataset For Video And Language Research Arxiv 1904 03493v6A

Use this skill when you need a multilingual video captioning dataset with English and Chinese captions for training video-language models. Avoid it when English-only video captions are sufficient.

researchgo
0
9
Visual Genome Connecting Language And Vision Using Crowdsourced Dense Image Annotations Arxiv 1602 07332v1A

Use this skill when you need a densely annotated dataset with region descriptions, objects, attributes, relationships, and QA pairs for training grounding-capable VLMs. Avoid it when you only need image-level captions or cannot use crowdsourced annotations.

testinggoapi
0
9
Vizwiz Grand Challenge Answering Visual Questions From Blind People Arxiv 1802 08218v4A

Use this skill when you need VQA data from real blind users asking questions about photos they took, testing VLM accessibility applications. Avoid it when you do not need accessibility-focused VQA evaluation.

testinggotesting
0
9
Vqav2 Making The V In Vqa Matter Arxiv 1612 00837v3A

Use this skill when you need a large-scale balanced VQA dataset where each question has complementary image pairs to reduce language bias. Avoid it when you do not need VQA data or simple VQA is sufficient.

devopsgoperformance
0
9
Webgpt Browser Assisted Question Answering With Human Feedback Arxiv 2112 09332v3A

Use this skill when you want to create training data for LLMs to use web browsing for answering questions with citations. Avoid it when you do not need web-augmented question answering.

researchrustgo
0
9
Webvid 10m A Large Scale Video Text Dataset Arxiv 2104 00650v1A

Use this skill when you want to collect video-text pairs from the web at 10M scale using stock video alt-text as captions. Avoid it when you need high-quality descriptions rather than stock video alt-text.

researchgoapi
0
9
Wildvision Evaluating Vision Language Models In The Wild With Human Preferences Arxiv 2406 11069v2A

Use this skill when you want to evaluate VLMs using real user interactions and preferences collected in the wild. Avoid it when you only need standard benchmark evaluation.

documentationgo
0
9
Wit Wikipedia Based Image Text Dataset For Multimodal Multilingual Machine Learning Arxiv 2103 01913v2A

Use this skill when you need a multilingual image-text dataset mined from Wikipedia covering 108 languages with curated metadata. Avoid it when you only need English data or web-crawled alt-text.

datagodatabase
0
9
Wukong A 100 Million Large Scale Chinese Cross Modal Pre Training Benchmark Arxiv 2202 06767v2A

Use this skill when you need a large-scale Chinese image-text dataset for training multilingual or Chinese-specific VLMs. Avoid it when you only need English image-text data.

researchgoperformance
0
9
An Empirical Study Of Training Self Supervised Vision Transformers Arxiv 2104 02057v2A

Use this skill when you want to understand empirical best practices for training self-supervised ViTs (MoCo v3) including data, augmentation, and training stability. Avoid it when you are not training self-supervised vision models.

researchgo
0
9
Anyres Plug And Play Dynamic Resolution For Vision Language Models Arxiv Anyres 2024A

Use this skill when you want to process images at any resolution by tiling them into patches that fit the vision encoder's native resolution. Avoid it when fixed-resolution processing is sufficient.

businessgo
0
9
Blip Bootstrapping Language Image Pre Training For Unified Vision Language Understanding And Generation Arxiv 2201 12086v1A

Use this skill when you want to bootstrap better captions from noisy web data using a captioner-filter loop. Avoid it when you have clean gold-standard captions and do not need bootstrapping.

developmentgoperformance
0
9
Cambrian 1 A Fully Open Vision Centric Exploration Of Multimodal Llms Arxiv 2406 16860v1A

Use this skill when you want a data-centric approach to VLM design that systematically evaluates vision encoders, connectors, and instruction data. Avoid it when you want a simple plug-and-play VLM without extensive ablation studies.

code-qualitygoperformance
0
9
Chameleon Mixed Modal Early Fusion Foundation Models Arxiv 2405 09818v1A

Use this skill when you want to train a mixed-modal model using early fusion of image and text tokens for both understanding and generation. Avoid it when you only need understanding (not generation) or cannot tokenize images.

code-qualitygo
0
9
Coca Contrastive Captioners Are Image Text Foundation Models Arxiv 2205 01917v2A

Use this skill when you want to combine contrastive and captioning objectives on the same image-text data for a unified vision-language foundation model. Avoid it when you only need one objective (contrastive or captioning) and not both.

code-qualitygo
0
9
Cogvlm Visual Expert For Pretrained Language Models Arxiv 2311 03079v2A

Use this skill when you want to add visual expert modules to an LLM with a two-stage data pipeline covering alignment and multi-task training. Avoid it when you are using a simpler linear projection approach.

code-qualitygoperformance
0
9
Convnext V2 Co Designing And Scaling Convnets With Masked Autoencoders Arxiv 2301 00808v2A

Use this skill when you want to pre-train ConvNet vision encoders using masked autoencoder self-supervised learning for VLM pipelines. Avoid it when ViT-based encoders are preferred.

code-qualitygo
0
9
D4 Improving Llm Pretraining Via Document De Duplication And Diversification Arxiv 2308 12284v2A

Use this skill when you want to both deduplicate and diversify training data using embedding-based clustering and selection. Avoid it when simple deduplication is sufficient or you cannot compute embeddings for all data.

testinggo
0
9
Data Centric Artificial Intelligence A Survey Arxiv 2303 10158v3A

Use this skill when you need a comprehensive overview of data-centric AI techniques including data quality, augmentation, selection, and engineering. Avoid it when you are looking for a specific method rather than a broad overview.

researchgo
0
9
Data Juicer A One Stop Data Processing System For Large Language Models Arxiv 2309 02033v5A

Use this skill when you need a configurable data processing pipeline with 50+ operators for cleaning, filtering, and preparing LLM/VLM training data. Avoid it when you have a simple dataset that does not need complex multi-stage processing.

code-qualitygodebugging
0
9
Datacomp In Search Of The Next Generation Of Multimodal Datasets Arxiv 2304 14108v2A

Use this skill when you need a systematic benchmark-driven approach to evaluate and compare data filtering strategies for CLIP-style vision-language pretraining. Avoid it when you already have a fixed, curated dataset and do not intend to experiment with filtering.

researchgoapi
0
9
Dataperf Benchmarks For Data Centric Ai Development Arxiv 2207 10062v2A

Use this skill when you need benchmarks for evaluating data-centric AI tasks including data selection, labeling, and slice discovery. Avoid it when you need model-centric benchmarks rather than data-centric ones.

researchgo
0
9
Dclm Datacomp For Language Models Arxiv 2406 11794v3A

Use this skill when you want a benchmark-driven approach to evaluate text data filtering strategies for LLM training, analogous to DataComp for CLIP. Avoid it when you are not filtering text data for LLM/VLM training.

researchgoperformance
0
9
Deduplicating Training Data Makes Language Models Better Arxiv 2107 06499v2A

Use this skill when you want to apply exact and near-duplicate removal to LLM training data using suffix arrays and MinHash. Avoid it when your dataset is already deduplicated or too small for dedup to matter.

datago
0
9
Demystifying Clip Data Arxiv 2309 16671v4A

Use this skill when you want to curate CLIP training data using metadata from an existing high-quality dataset rather than model-based filtering. Avoid it when you do not have a reference metadata set or prefer model-based scoring over metadata matching.

businessgoperformance
0
9
Dino Emerging Properties In Self Supervised Vision Transformers Arxiv 2104 14294v2A

Use this skill when you want to train a self-supervised ViT that learns rich visual features with emerging object segmentation properties. Avoid it when CLIP-supervised features are sufficient for your VLM.

code-qualitygo
0
9
Dinov2 For Dense Prediction Tasks Arxiv Dinov2 Dense 2024A

Use this skill when you want to use DINOv2 self-supervised features for dense prediction tasks like segmentation and depth in VLM pipelines. Avoid it when CLIP features are sufficient for your VLM.

researchgoperformance
0
9
Dinov2 Learning Robust Visual Features Without Supervision Arxiv 2304 07193v2A

Use this skill when you want to curate and deduplicate large-scale image data for self-supervised vision pretraining using retrieval-based curation. Avoid it when you are not doing self-supervised pretraining or have curated image data.

ai-agentsgo
0
9
Direct Preference Optimization Your Language Model Is Secretly A Reward Model Arxiv 2305 18290v2A

Use this skill when you want to align a language/vision-language model using preference data without training a separate reward model. Avoid it when you prefer RLHF with explicit reward models or do not have preference data.

code-qualitygo
0
9
Emu2 Generative Multimodal Models Are In Context Learners Arxiv 2312 13286v1A

Use this skill when you want to train a generative multimodal model that can both understand and generate images with in-context learning from interleaved data. Avoid it when you only need understanding without generation.

devopsgoperformance
0
9
Eva Clip Improved Training Techniques For Clip At Scale Arxiv 2303 15389v1A

Use this skill when you want to improve CLIP training efficiency through better data, larger models, and optimized training techniques. Avoid it when you are training a small CLIP model and do not need advanced optimization.

researchgoperformance
0
9
Filtering Distillation And Hard Negatives For Vision Language Pre Training Arxiv 2301 02280v2A

Use this skill when you want to filter web-crawled image-text data using a learned data filtering network (DFN) trained to predict which pairs improve downstream performance. Avoid it when you do not have labeled data for training a filtering network or prefer simple threshold-based filtering.

researchgoapi
0
9
Fineweb Decanting The Web For The Finest Text Data At Scale Arxiv 2406 17557v1A

Use this skill when you want to understand and apply state-of-the-art web text filtering techniques including quality classifiers trained on curated data. Avoid it when you are not processing web text or have sufficient text data already.

researchgo
0
9
Gemini A Family Of Highly Capable Multimodal Models Arxiv 2312 11805v3A

Use this skill when you want to understand Google's industrial-scale multimodal training data strategy for frontier VLMs. Avoid it when you need open, reproducible data strategies.

researchgoperformance
0
9