OpenAI Sora 2 Release

OpenAI has announced Sora 2, a flagship video and audio generation model designed to function as a more accurate world simulator. Sora 2 improves upon its predecessor by enhancing physical realism, increasing controllability, and integrating synchronized dialogue and sound effects.

Advanced World Simulation and Physical Accuracy

Sora 2 is engineered to better obey the laws of physics, reducing the "overoptimism" seen in previous video models where objects would morph or teleport to satisfy a prompt. The model can now accurately simulate complex dynamics, such as buoyancy and rigidity during a backflip on a paddleboard, or the rebound of a basketball off a backboard after a missed shot.

OpenAI describes this progress as a "GPT-3.5 moment for video," moving beyond the basic object permanence of the original Sora model toward a system that can model failure and realistic physical interactions. This capability is viewed as a critical step toward creating AI systems that deeply understand the physical world.

Multimodal Capabilities and Controllability

Sora 2 expands its generative capabilities across several dimensions:

  • Integrated Audio: The model generates sophisticated background soundscapes, speech, and sound effects with high realism.
  • Enhanced Controllability: Sora 2 can follow intricate instructions across multiple shots while maintaining a consistent world state.
  • Stylistic Versatility: The model excels in cinematic, realistic, and anime styles.
  • Real-World Injection: Users can insert real people, animals, or objects into generated environments by providing a reference video, allowing the model to accurately portray appearance and voice.

The Sora Social App and "Characters"

To deploy Sora 2, OpenAI is launching a social iOS app called "Sora." The app's central feature is "characters," which allows users to create a digital likeness of themselves or friends through a one-time video-and-audio recording. These characters can then be dropped into any Sora-generated scene with high fidelity.

The app is designed to prioritize creation over consumption, featuring a customizable feed powered by natural-language-instructable recommender algorithms. Users can control their feed through LLM-based instructions and periodic wellbeing polls.

Safety and Responsibility Framework

OpenAI has implemented several safeguards to address the risks associated with generative video and social media:

  • Likeness Control: Users maintain end-to-end control over their characters, including the ability to revoke access or remove any video containing their likeness.
  • Teen Protections: The app includes default limits on the daily number of generations teens can see in the feed, stricter character permissions, and parental controls managed via ChatGPT.
  • Content Moderation: In addition to automated safety stacks, OpenAI is scaling human moderation teams to review bullying and harmful content.
  • Monetization: OpenAI states its current plan is limited to offering paid options for extra video generation during periods of high compute demand, explicitly avoiding models that optimize for time spent in-feed.

Availability and Access

Sora 2 is rolling out via an invite-based system starting in the U.S. and Canada. Access is provided through the Sora iOS app and sora.com.

  • Free Tier: Initially available for free with generous limits subject to compute constraints.
  • Sora 2 Pro: An experimental, higher-quality model available to ChatGPT Pro users on sora.com and eventually within the app.
  • API Access: OpenAI plans to release Sora 2 via API in the future. Sora 1 Turbo remains available for existing users.

Sources