A robot that pauses to think is a robot that drops things, which is the problem Black Forest Labs has aimed its newest release at.
The model is FLUX 3 Action, published this week with open weights. Its job is to watch a workspace through several cameras and then answer two questions at once. What should the agent do now? What will the scene look like after it does? Researchers call this class of system a world-action model.
FLUX 3 supplies the foundation. That multimodal system was trained mainly on video, with image and audio mixed in. The Action variant keeps the visual grounding and adds control.
Efficiency is the selling point. Seven billion parameters sit inside FLUX 3 Action, which the company says is under half what the previous best open model required. It claims a record success rate on the RoboLab-120 leaderboard. Speed runs as high as 3.95 times faster.
Running locally matters here. A model that lives on the device cannot lose a packet between the camera and the arm.
Black Forest Labs sees a second market in software. Video games make cheap proving grounds for navigation, and for agents that operate computers rather than grippers.