Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 15,577–15,600 of 23,891 skills
Use this skill when you want to synthesize high-quality instruction data using GPT-4V on diverse image sources for training small VLMs. Avoid it when you cannot afford GPT-4V API calls or are training a large-scale model.
Use this skill when you want to know the compute-optimal ratio of model size to training data based on Chinchilla scaling laws. Avoid it when you are not making model/data scaling decisions.
Use this skill when you want to understand power-law relationships between model size, data size, and compute for optimal resource allocation. Avoid it when you are not making scaling decisions.
Use this skill when you want to understand how to optimally train when data is limited and must be repeated, and how to trade off epochs vs model size. Avoid it when you have unlimited unique data and do not need to repeat data.
Use this skill when you want to estimate the influence of individual training examples on model predictions for data debugging and selection. Avoid it when you cannot afford the compute for influence estimation or have too many training examples.
Use this skill when you want to prune training data using early-epoch error norms (EL2N scores) to select the most informative examples. Avoid it when you cannot compute early training predictions or all data is equally important.
Use this skill when you want to predict LLM performance from data mixing ratios and optimize the mix without expensive full-scale training. Avoid it when you have a single data source or cannot run ablation experiments.
Use this skill when you want to beat standard scaling laws by pruning low-quality or redundant data points using perplexity or EL2N scores. Avoid it when you do not have compute for data quality scoring or your data is already curated.
Use this skill when you want to select diverse, uncertain samples for annotation using gradient embeddings that capture both uncertainty and diversity. Avoid it when simple uncertainty sampling is sufficient or you cannot compute gradients.
Use this skill when you want to understand which visual tokenizer properties matter most for VLM performance. Avoid it when you have already chosen your visual encoder.
Use this skill when you want systematic ablation insights on VLM pretraining choices: frozen vs unfrozen LLM, interleaved vs paired data, text mixing. Avoid it when you have a well-established pretraining recipe.
Use this skill when you want to understand the C4 text cleaning methodology and its impact on LLM pretraining quality. Avoid it when you are not cleaning web text or have a different cleaning pipeline.
Use this skill when you want to understand how to scale instruction fine-tuning across number of tasks, model size, and chain-of-thought data for maximum benefit. Avoid it when you are doing single-task fine-tuning.
Use this skill when you want insights on training a smaller but stronger VLM through improved data quality, SigLIP encoder, and efficient training recipe. Avoid it when you are not optimizing VLM training efficiency.
Use this skill when you want to train CLIP models on open datasets (LAION) with reproducible training recipes and systematic data ablations. Avoid it when you are using pre-trained CLIP models and do not need to train your own.
Use this skill when you want to train an object detector using image captions as weak supervision instead of bounding box annotations. Avoid it when you have abundant bounding box annotations.
Use this skill when you want to understand how CLIP learns multimodal neurons that respond to both visual and textual concepts for debugging data quality. Avoid it when you do not need interpretability or representation analysis.
Use this skill when you need a comprehensive survey of multimodal learning with transformers covering fusion strategies, pretraining, and applications. Avoid it when you need a specific method rather than a broad overview.
Use this skill when you want to understand optimal data mixing recipes for multimodal LLM pretraining across interleaved, image-text, and text-only data. Avoid it when you only have one data type or cannot run ablation studies to tune the mix.
Use this skill when you want a unified training recipe for single-image, multi-image, and video understanding with curated data at each stage. Avoid it when you only need single-image understanding.
Use this skill when you want to align an LLM with just 1,000 carefully curated instruction examples, demonstrating that quality trumps quantity. Avoid it when you have abundant alignment data or need maximum diversity.
Use this skill when you want to understand how label noise in training data affects optimization and generalization, particularly for noisy web data. Avoid it when you do not deal with noisy labels.
Use this skill when you want to curate a high-quality academic-task-oriented instruction dataset for VLM fine-tuning. Avoid it when you need only synthetic conversation data or cannot access the academic VQA datasets.
Use this skill when you want an efficient 8B VLM training recipe with curated data combining web interleaved, paired, and instruction data. Avoid it when you need a larger model or have a different data strategy.