Photorealistic video from text
Sora is the video generation AI developed by OpenAI. Its ability to produce convincingly realistic footage from a written description startled people and came to symbolize the arrival of video generation.
What it does
- Video from text: describe a scene and get footage
- Video from an image: give natural motion to a single photograph or illustration
- Editing and stitching: extending scenes and remixing
- Increasing support for generation with audio, producing something closer to finished
A mobile app and social sharing features have pushed it toward being enjoyable for non-specialists too.
Why it matters
Concept footage that used to take days of shooting and editing now takes minutes. The cost of video made to communicate something — advertising, social content, pitch material — has fallen dramatically. At the same time, the effect on creative work and the risk of mass-produced realistic fakes became concrete concerns with its arrival.
Using it
Access requires a paid ChatGPT plan or the Sora app or website, and availability varies by region and plan. It is not yet at the stage of producing a long piece exactly as intended; short shots combined together is the working approach. Physically implausible motion still appears, so generating repeatedly and selecting the good take is realistic.
Competition is intense — Google’s Veo among others — and the lead changes frequently. Rather than tracking tool names, the durable point is that text can now produce footage of near-live-action quality.
Watermarking and labeling to indicate AI generation are being built in. Sharing that assumption matters for both viewers and makers.