Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,891
skills in category
996
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 15,601–15,624 of 23,891 skills

Gemini A Family Of Highly Capable Multimodal Models Arxiv 2312 11805v3A

Use this skill when you want to understand Google's industrial-scale multimodal training data strategy for frontier VLMs. Avoid it when you need open, reproducible data strategies.

researchgoperformance
0
9
Fineweb Decanting The Web For The Finest Text Data At Scale Arxiv 2406 17557v1A

Use this skill when you want to understand and apply state-of-the-art web text filtering techniques including quality classifiers trained on curated data. Avoid it when you are not processing web text or have sufficient text data already.

researchgo
0
9
Filtering Distillation And Hard Negatives For Vision Language Pre Training Arxiv 2301 02280v2A

Use this skill when you want to filter web-crawled image-text data using a learned data filtering network (DFN) trained to predict which pairs improve downstream performance. Avoid it when you do not have labeled data for training a filtering network or prefer simple threshold-based filtering.

researchgoapi
0
9
Eva Clip Improved Training Techniques For Clip At Scale Arxiv 2303 15389v1A

Use this skill when you want to improve CLIP training efficiency through better data, larger models, and optimized training techniques. Avoid it when you are training a small CLIP model and do not need advanced optimization.

researchgoperformance
0
9
Dinov2 For Dense Prediction Tasks Arxiv Dinov2 Dense 2024A

Use this skill when you want to use DINOv2 self-supervised features for dense prediction tasks like segmentation and depth in VLM pipelines. Avoid it when CLIP features are sufficient for your VLM.

researchgoperformance
0
9
Dclm Datacomp For Language Models Arxiv 2406 11794v3A

Use this skill when you want a benchmark-driven approach to evaluate text data filtering strategies for LLM training, analogous to DataComp for CLIP. Avoid it when you are not filtering text data for LLM/VLM training.

researchgoperformance
0
9
Dataperf Benchmarks For Data Centric Ai Development Arxiv 2207 10062v2A

Use this skill when you need benchmarks for evaluating data-centric AI tasks including data selection, labeling, and slice discovery. Avoid it when you need model-centric benchmarks rather than data-centric ones.

researchgo
0
9
Datacomp In Search Of The Next Generation Of Multimodal Datasets Arxiv 2304 14108v2A

Use this skill when you need a systematic benchmark-driven approach to evaluate and compare data filtering strategies for CLIP-style vision-language pretraining. Avoid it when you already have a fixed, curated dataset and do not intend to experiment with filtering.

researchgoapi
0
9
Data Centric Artificial Intelligence A Survey Arxiv 2303 10158v3A

Use this skill when you need a comprehensive overview of data-centric AI techniques including data quality, augmentation, selection, and engineering. Avoid it when you are looking for a specific method rather than a broad overview.

researchgo
0
9
An Empirical Study Of Training Self Supervised Vision Transformers Arxiv 2104 02057v2A

Use this skill when you want to understand empirical best practices for training self-supervised ViTs (MoCo v3) including data, augmentation, and training stability. Avoid it when you are not training self-supervised vision models.

researchgo
0
9
Wukong A 100 Million Large Scale Chinese Cross Modal Pre Training Benchmark Arxiv 2202 06767v2A

Use this skill when you need a large-scale Chinese image-text dataset for training multilingual or Chinese-specific VLMs. Avoid it when you only need English image-text data.

researchgoperformance
0
9
Webvid 10m A Large Scale Video Text Dataset Arxiv 2104 00650v1A

Use this skill when you want to collect video-text pairs from the web at 10M scale using stock video alt-text as captions. Avoid it when you need high-quality descriptions rather than stock video alt-text.

researchgoapi
0
9
Webgpt Browser Assisted Question Answering With Human Feedback Arxiv 2112 09332v3A

Use this skill when you want to create training data for LLMs to use web browsing for answering questions with citations. Avoid it when you do not need web-augmented question answering.

researchrustgo
0
9
Vatex A Large Scale High Quality Multilingual Dataset For Video And Language Research Arxiv 1904 03493v6A

Use this skill when you need a multilingual video captioning dataset with English and Chinese captions for training video-language models. Avoid it when English-only video captions are sufficient.

researchgo
0
9
Snli Ve Visual Entailment Dataset Arxiv 1901 06706v1A

Use this skill when you need a visual entailment dataset where the model must determine if a text hypothesis is entailed, contradicted, or neutral with respect to an image premise. Avoid it when you do not need visual entailment or NLI-style evaluation.

researchgoperformance
0
9
Segment Anything Arxiv 2304 02643v1A

Use this skill when you need to build a billion-scale segmentation mask dataset using a model-in-the-loop interactive annotation approach. Avoid it when you need text-grounded segmentation rather than prompt-based mask generation.

researchgo
0
9
Scannet Richly Annotated 3d Reconstructions Of Indoor Scenes Arxiv 1702 04405v2A

Use this skill when you need a richly annotated 3D indoor scene dataset for 3D vision-language understanding research. Avoid it when you do not work with 3D scene understanding.

researchgoapi
0
9
Nuscenes A Multimodal Dataset For Autonomous Driving Arxiv 1903 11027v5A

Use this skill when you need a comprehensive multimodal autonomous driving dataset with cameras, LiDAR, radar, and 3D annotations for driving VLMs. Avoid it when driving data is not relevant to your work.

researchgo
0
9
Nlvr2 A Visual Reasoning Benchmark For Natural Language Arxiv 1811 00491v2A

Use this skill when you need a visual reasoning benchmark where the model must determine if a statement is true for a pair of images. Avoid it when single-image reasoning evaluation is sufficient.

researchgo
0
9
Msr Vtt A Large Video Description Dataset For Bridging Video And Language Arxiv Msrvtt 2016A

Use this skill when you need a video description dataset with 10K web clips and 200K sentences for training and evaluating video-language models. Avoid it when you have sufficient video captioning data.

researchgo
0
9
Mme A Comprehensive Evaluation Benchmark For Multimodal Large Language Models Arxiv 2306 13394v2A

Use this skill when you need a comprehensive VLM benchmark testing both perception and cognition abilities with yes/no questions. Avoid it when you need open-ended evaluation or have specific benchmark needs.

researchgotesting
0
9
Mimic Cxr A De Identified Publicly Available Database Of Chest Radiographs With Free Text Reports Arxiv Mimic Cxr 2019A

Use this skill when you need a large-scale medical image-report dataset for training biomedical VLMs on chest X-ray understanding. Avoid it when you do not work with medical imaging.

researchgodatabase
0
9
Lmms Eval Reality Check On The Evaluation Of Large Multimodal Models Arxiv 2407 12772v3A

Use this skill when you need a unified evaluation framework for consistently evaluating VLMs across dozens of benchmarks. Avoid it when you only evaluate on 1-2 benchmarks.

researchgoapi
0
9
Hallusionbench An Advanced Diagnostic Suite For Entangled Language Hallucination And Visual Illusion In Large Vision Lan Arxiv 2310 14566v3A

Use this skill when you need a diagnostic benchmark testing VLM vulnerability to both language hallucination and visual illusion. Avoid it when you only need standard hallucination evaluation (use POPE instead).

researchgotesting
0
9