The multimodal model generates images, 20-second video clips with sound, and even predicts physical actions from a single…