LLMs develop rich visual priors despite text-only training. We reveal that visual priors are composed of separable perception and reasoning priors with unique scaling trends and origins. Visual reasoning ability is predominantly developed by pre-training on reasoning-centric data (code, math, academia). We propose a data-centric recipe for pre-training vision-aware LLMs verified in 1T token scale pre-training across 100+ controlled experiments consuming 500,000 GPU-hours.
Scanned 10/3/2026
npx -y skills add hiyenwong/ai_collection --skill arxiv-2509-26625-learning-to-see-before-seeing-demystifying-llm-visual-priors --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Arxiv 2509 26625 Learning To See Before Seeing Demystifying Llm Visual Priors?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-arxiv-2509-26625-learning-to-see-before-seeing-dem)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
title: "Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training"
authors: "Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos"
arxiv_id: "2509.26625"
categories: "cs.LG; cs.AI; cs.CV; cs.MM"
utility: 0.9
date_added: "2026-09-29"
category: "vision-generative"
---
# Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
## Abstract
LLMs develop rich visual priors despite text-only training. We reveal that visual priors are composed of separable perception and reasoning priors with unique scaling trends and origins. Visual reasoning ability is predominantly developed by pre-training on reasoning-centric data (code, math, academia). We propose a data-centric recipe for pre-training vision-aware LLMs verified in 1T token scale pre-training across 100+ controlled experiments consuming 500,000 GPU-hours.
## Key Contributions
- Novel approach in vision generative domain
- Utility score: 0.9
- Published on arXiv: 2509.26625
## Potential Applications
- Research reference for vision generative
- Building block for related systems
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!