Google Deepmind argues video generators already contain the world models computer vision has been missing
Google Deepmind researchers have introduced GenCeption, a framework that repurposes video generators to perform computer vision tasks like depth estimation and object segmentation. By training almost entirely on synthetic video data, the system achieves performance levels comparable to specialized models while requiring significantly less information. This development suggests that video generators inherently contain the complex spatial understanding required for vision tasks, potentially changing how models learn to interpret physical environments.
Covered by 1 source
- TThe Decoder↗Jonathan Kemper2d ago