Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

Check out our website: LFG! Arxiv here

Abstract:

LFG is a label-free, teacher-guided framework for learning autonomous driving representations directly from unposed videos. A pretrained encoder and a causal autoregressive module predict point maps, camera poses, semantic layouts, confidence maps, and motion masks over a short horizon under multiple teacher-guided supervision. LFG achieves state-of-the-art planning performance from a single front camera, is data efficient, and transfers across semantic, geometric, and motion tasks.