LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
Apple researchers have introduced LensVLM, a method that allows vision language models to process text as images rather than traditional token sequences. By utilizing selective context expansion, this approach maintains efficiency by adjusting rendering resolution, which potentially improves how models handle long-form text while bypassing standard tokenization limitations.