H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
Most visible doc retrievers in manufacturing at this time are hand-me-downs. ColPali and the fashions that adopted it take a generative vision-language mannequin and repurpose it as an encoder. The consequence nonetheless carries a individually pretrained imaginative and prescient tower and a causal decoder that by no means generates a token. That is parameter and…
