OpenAI Sora 2 Safety Framework
OpenAI has implemented a comprehensive safety framework for Sora 2 and the Sora app to manage the risks associated with high-fidelity video and audio generation. This framework relies on a combination of provenance signals, consent-based likeness controls, and layered content filtering to prevent misuse and protect user identity.
AI Content Provenance and Identification
Sora 2 ensures that AI-generated content can be identified through both visible and invisible signals. Every video generated includes C2PA metadata, an industry-standard signature, and invisible provenance signals. OpenAI maintains internal reverse-image and audio search tools to trace videos back to the Sora platform with high accuracy.
Additionally, many outputs feature dynamically moving watermarks that include the name of the creator, providing a visible indicator of the content's origin.
Likeness and Consent Management
OpenAI has introduced specific mechanisms to handle the likeness of real people in generated videos, categorized by the method of creation:
Image-to-Video with Real People
Users can upload photos of family and friends to create videos, provided they attest to having the necessary consent and rights. These generations are subject to strict safety guardrails, which are even more restrictive than those applied to Sora Characters. Content involving children or young-looking persons is subject to the same protections but with even stricter moderation.
Consent-Based Likeness via Characters
The "Characters" feature allows users to control their own appearance and voice likeness. Key protections include:
- Access Control: Only the user decides who can use their character and can revoke access at any time.
- Public Figure Restrictions: Depictions of public figures are blocked unless they are using the Characters feature.
- Visibility: Users can see all videos featuring their character, including drafts created by others, allowing for review and deletion.
- Customizable Guardrails: Users can enable a strict set of guardrails to prevent major changes to their appearance, avoid embarrassing situations, and maintain identity consistency.
Protections for Teen Users
Sora includes specialized safeguards for younger users to ensure an age-appropriate experience:
- Content Filtering: Mature output is limited, and the feed is filtered to remove harmful or age-inappropriate material for teen accounts.
- Social Restrictions: Teen profiles are not recommended to adults, and adults are prohibited from initiating messages with teens.
- Parental Controls: Through ChatGPT, parents can manage direct messages and select a non-personalized feed for their children.
- Usage Limits: By default, teens have limits on continuous scrolling within the Sora app.
Content Filtering and Harm Prevention
Sora employs layered defenses to block unsafe content at both the prompt and output stages. Guardrails target sexual material, terrorist propaganda, and the promotion of self-harm by scanning video frames and audio transcripts.
Beyond the generation phase, automated systems scan all feed content against OpenAI's Global Usage Policies, supplemented by human review for high-impact harms. OpenAI notes that policies have been tightened relative to image generation due to the increased realism of motion and audio in Sora 2.
Audio and Music Safeguards
To address the risks associated with audio generation, Sora implements the following protections:
- Speech Scanning: Transcripts of generated speech are automatically scanned for policy violations.
- Music Protection: The system blocks attempts to generate music that imitates living artists or existing works.
- Takedown Requests: OpenAI honors takedown requests from creators who believe Sora output infringes on their work.
User Control and Recourse
Users maintain control over their content and social interactions within the app:
- Sharing Control: Videos are only shared to the feed if the user explicitly chooses to do so.
- Reporting and Blocking: All videos, profiles, direct messages, comments, and characters can be reported for abuse.
- Blocking: Users can block accounts to prevent others from seeing their profile, accessing their character, or sending direct messages.
Sources
- OriginalCreating with Sora Safely