Qwen3.8-Omni-Flash handles audio and video together and calls tools on its own.
Gander keeps a conversation alive in one model while a second model does the work.
Google's agentic video processing cuts tokens by up to 88 percent on long recordings.
The Apache-licensed 30B model brings agentic AI to local machines.
The multimodal model generates images, 20-second video clips with sound, and even predicts physical actions from a single…