H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
H Company has released NeoMME, a family of multimodal encoders available in 260 million and 800 million parameter sizes. These models function as bidirectional encoders that process text and raw image patches within a single Transformer architecture, intentionally omitting the pretrained vision towers and causal decoders found in typical systems.
Covered by 1 source
- MMarkTechPost↗Asif RazzaqSep 6