Agent workloads burn tokens in ways chat never did, and Alibaba has priced its newest multimodal model accordingly. Qwen3.8-Omni-Flash lists at $0.15 per million input tokens and $0.47 per million output tokens.
Google’s Gemini 3.8 Flash asks $0.75 for input and $3.75 for output at its introductory rate. Those figures double on January 1, 2027.
The model itself is Qwen’s first multimodal release built for agents rather than conversation. It reads audio and video together, reaches conclusions, and calls tools without a separate orchestrator. Editing vlogs, translating short videos and summarizing films are the representative jobs. Context runs to one million tokens, and Qwen says audio-video performance lands close to Gemini 3.8 Flash.
Audio input costs under $0.01 an hour by Qwen’s estimate. A 720p stream carrying audio at one frame per second runs roughly $0.20 before response costs.
Much of the practical weight sits with two open source companions. Qwen-MM-Plugins is the first. It hands coding agents extra abilities, among them editing video, telling speakers apart, and turning PDFs into notes. Whole workflows can be reused across sessions. Claude Code, Gemini CLI and Qwen Code all accept it. The second, Qwen-Live Harness, wires a camera and a microphone into a live session. Access runs through Qwen Studio, Qwen Cloud or the API.
Undercutting a rival on price buys developer adoption while capability converges, and Alibaba is spending on exactly that.