Tag: multimodal

H Company’s NeoMME encoders skip the vision tower entirely

Tiny single-tower models aim to make visual document retrieval faster and cheaper.

Gemini Omni 1.1 Flash extends video scenes to 40 seconds

Google's updated video model reads ten seconds of context, pins start and end frames, and renders drafts in…

MiniMax H3 blends audio and video into one generation model

MiniMax's H3 model generates 2K video with native stereo audio from mixed text, image and sound prompts.