26 August 2026

EchoWM model generates video, audio, and speech from camera movements

First reported

TLDR AI ran this on .

  • EchoWM is a world model, a type of AI trained to simulate how environments behave, that creates synchronized video, environmental sound, music, and speech based on specified camera paths.
  • The model can generate 720p resolution video while maintaining consistency with audio elements, responding to defined camera movements in three-dimensional space.

How it was covered