OpenAI Sora 2 and Sora App Safety Framework
OpenAI has launched Sora 2 and the Sora app, integrating high-fidelity video generation with a collaborative creation environment. To manage the risks associated with high-realism video and audio, OpenAI has implemented a multi-layered safety architecture focusing on provenance, consent, and content filtering.
AI Content Provenance and Identification
Sora 2 utilizes both visible and invisible signals to distinguish AI-generated content from authentic footage. All outputs include a visible watermark at launch. For technical provenance, Sora embeds C2PA metadata—an industry-standard signature—into all videos. Additionally, OpenAI maintains internal reverse-image and audio search tools to trace videos back to the Sora system with high accuracy, building upon the systems previously used in ChatGPT image generation and Sora 1.
Consent-Based Character Likeness
OpenAI has introduced a character feature that allows users to control their own likeness. The framework is designed to ensure that audio and image likenesses captured in characters are used only with user consent. Key controls include:
- Access Management: Users decide who can use their characters and can revoke access at any time.
- Visibility: Users can view all videos featuring their characters, including drafts created by others, allowing for review, deletion, or reporting.
- Behavioral Preferences: Users can set specific preferences for how their characters behave (e.g., requiring the character to always wear a specific item of clothing).
- Public Figure Protections: Sora blocks the depiction of public figures, except when those figures are using the characters feature.
Safeguards for Teen Users
Sora includes specific protections for younger users to limit exposure to mature content and prevent unwanted interactions. These measures include:
- Content Filtering: The feed is designed to be appropriate for teens, and teen profiles are not recommended to adults.
- Communication Restrictions: Adults are prohibited from initiating messages with teens.
- Parental Controls: New controls in ChatGPT allow parents to manage direct messages for teens and select a non-personalized feed within the Sora app.
- Usage Limits: Teens have default limits on continuous scrolling within the app.
Content Filtering and Harm Reduction
Sora employs layered defenses to block unsafe content during both the creation and distribution phases. Guardrails check prompts and outputs across multiple video frames and audio transcripts to block sexual material, terrorist propaganda, and self-harm promotion.
Because Sora's realism, motion, and audio capabilities increase potential risks, OpenAI has tightened policies compared to its image generation tools. Beyond the generation phase, automated systems scan all feed content against Global Usage Policies to filter unsafe or age-inappropriate material. These automated systems are supplemented by human review for high-impact harms.
Audio and Music Safety
To address the risks associated with audio generation, Sora automatically scans transcripts of generated speech for policy violations. To protect intellectual property and artist rights, Sora blocks attempts to generate music that imitates living artists or existing works. OpenAI honors takedown requests from creators who believe Sora outputs infringe upon their work.
User Control and Recourse
Users maintain control over their shared content and interactions. Videos are only shared to the feed upon user choice, and published content can be removed at any time. Users can report any video, profile, or comment for abuse, and they can block accounts to prevent others from from seeing their profile or contacting them via direct message.
Sources
- OriginalLaunching Sora responsibly