VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
Apple researchers have introduced VideoFlexTok, a new method for video tokenization that utilizes a coarse-to-fine approach to convert raw pixels into compressed data representations. By dynamically adjusting the level of detail processed, this technique aims to improve how visual information is organized and preserved for downstream machine learning tasks.