Video with sound
Veo is the video generation AI developed by Google DeepMind. It produces high-quality footage from text or images, and generates matching audio — effects, ambience, and dialogue — at the same time, which is what drew attention. Alongside OpenAI’s Sora, it is at the front of video generation.
What it does
- Video from text: short footage from a description of a scene
- Video from an image: natural motion added to a single picture or photograph
- Generation with audio: ambience and dialogue synchronized to the footage
- Available through the Gemini app and video creation tools
Against the silent footage of early video generation, finishing with sound included is its distinguishing quality.
Alongside Sora
Sora and Veo are constantly compared as the two leaders in video generation. The lead changes with each generation, so competition accelerating the field is a truer description than either being ahead. Google’s advantage is being able to build Veo directly into platforms the size of YouTube and Gemini.
Using it
Veo is available through paid Gemini plans, with clip length and features varying by plan. Generated video carries a SynthID watermark identifying it as AI-produced, among other safeguards. The general cautions for video generation — unnatural motion, deepfake misuse, copyright debate — apply here as well.
Each time Veo footage circulates as indistinguishable from real, the debate about countermeasures reopens. The advance of video AI and the making of social rules around it proceed together.