Tag: multimodal AI

Alibaba builds its first agent model to read video and audio

Qwen3.8-Omni-Flash handles audio and video together and calls tools on its own.

Tencent splits a voice agent into a brain and a cerebellum

Gander keeps a conversation alive in one model while a second model does the work.

Gemini Flash models now pick which video frames to watch

Google's agentic video processing cuts tokens by up to 88 percent on long recordings.

Meta’s Glimmer release puts desktop agents in reach

The Apache-licensed 30B model brings agentic AI to local machines.

Black Forest Labs debuts Flux 3 with video, audio, and robotic vision

The multimodal model generates images, 20-second video clips with sound, and even predicts physical actions from a single…