Tiny single-tower models aim to make visual document retrieval faster and cheaper.
Google's updated video model reads ten seconds of context, pins start and end frames, and renders drafts in…
MiniMax's H3 model generates 2K video with native stereo audio from mixed text, image and sound prompts.