Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 15,601–15,624 of 23,891 skills
Use this skill when you want to understand Google's industrial-scale multimodal training data strategy for frontier VLMs. Avoid it when you need open, reproducible data strategies.
Use this skill when you want to understand and apply state-of-the-art web text filtering techniques including quality classifiers trained on curated data. Avoid it when you are not processing web text or have sufficient text data already.
Use this skill when you want to filter web-crawled image-text data using a learned data filtering network (DFN) trained to predict which pairs improve downstream performance. Avoid it when you do not have labeled data for training a filtering network or prefer simple threshold-based filtering.
Use this skill when you want to improve CLIP training efficiency through better data, larger models, and optimized training techniques. Avoid it when you are training a small CLIP model and do not need advanced optimization.
Use this skill when you want to use DINOv2 self-supervised features for dense prediction tasks like segmentation and depth in VLM pipelines. Avoid it when CLIP features are sufficient for your VLM.
Use this skill when you want a benchmark-driven approach to evaluate text data filtering strategies for LLM training, analogous to DataComp for CLIP. Avoid it when you are not filtering text data for LLM/VLM training.
Use this skill when you need benchmarks for evaluating data-centric AI tasks including data selection, labeling, and slice discovery. Avoid it when you need model-centric benchmarks rather than data-centric ones.
Use this skill when you need a systematic benchmark-driven approach to evaluate and compare data filtering strategies for CLIP-style vision-language pretraining. Avoid it when you already have a fixed, curated dataset and do not intend to experiment with filtering.
Use this skill when you need a comprehensive overview of data-centric AI techniques including data quality, augmentation, selection, and engineering. Avoid it when you are looking for a specific method rather than a broad overview.
Use this skill when you want to understand empirical best practices for training self-supervised ViTs (MoCo v3) including data, augmentation, and training stability. Avoid it when you are not training self-supervised vision models.
Use this skill when you need a large-scale Chinese image-text dataset for training multilingual or Chinese-specific VLMs. Avoid it when you only need English image-text data.
Use this skill when you want to collect video-text pairs from the web at 10M scale using stock video alt-text as captions. Avoid it when you need high-quality descriptions rather than stock video alt-text.
Use this skill when you want to create training data for LLMs to use web browsing for answering questions with citations. Avoid it when you do not need web-augmented question answering.
Use this skill when you need a multilingual video captioning dataset with English and Chinese captions for training video-language models. Avoid it when English-only video captions are sufficient.
Use this skill when you need a visual entailment dataset where the model must determine if a text hypothesis is entailed, contradicted, or neutral with respect to an image premise. Avoid it when you do not need visual entailment or NLI-style evaluation.
Use this skill when you need to build a billion-scale segmentation mask dataset using a model-in-the-loop interactive annotation approach. Avoid it when you need text-grounded segmentation rather than prompt-based mask generation.
Use this skill when you need a richly annotated 3D indoor scene dataset for 3D vision-language understanding research. Avoid it when you do not work with 3D scene understanding.
Use this skill when you need a comprehensive multimodal autonomous driving dataset with cameras, LiDAR, radar, and 3D annotations for driving VLMs. Avoid it when driving data is not relevant to your work.
Use this skill when you need a visual reasoning benchmark where the model must determine if a statement is true for a pair of images. Avoid it when single-image reasoning evaluation is sufficient.
Use this skill when you need a video description dataset with 10K web clips and 200K sentences for training and evaluating video-language models. Avoid it when you have sufficient video captioning data.
Use this skill when you need a comprehensive VLM benchmark testing both perception and cognition abilities with yes/no questions. Avoid it when you need open-ended evaluation or have specific benchmark needs.
Use this skill when you need a large-scale medical image-report dataset for training biomedical VLMs on chest X-ray understanding. Avoid it when you do not work with medical imaging.
Use this skill when you need a unified evaluation framework for consistently evaluating VLMs across dozens of benchmarks. Avoid it when you only evaluate on 1-2 benchmarks.
Use this skill when you need a diagnostic benchmark testing VLM vulnerability to both language hallucination and visual illusion. Avoid it when you only need standard hallucination evaluation (use POPE instead).