OpenAI Approach to Data and AI
OpenAI is developing a new tool called Media Manager to give creators and content owners direct control over how their work is included or excluded from machine learning research and training. This initiative is part of a broader strategy to establish a social contract for content in the AI age, balancing the need for diverse training data with the rights of creators and publishers.
Media Manager and Creator Control
OpenAI is building Media Manager to provide a scalable solution for content owners to express their preferences regarding AI training. While OpenAI previously pioneered the use of web crawler permissions (similar to the robots.txt standard) to allow publishers to opt-out of training, the company acknowledges that these signals are incomplete because many creators do not control the websites where their content is hosted.
Key details regarding Media Manager include:
- Timeline: The goal is to have the tool in place by 2025.
- Technical Requirement: Developing the tool requires cutting-edge machine learning research to identify copyrighted text, images, audio, and video across multiple sources.
- Objective: OpenAI hopes Media Manager will set an industry standard for how creators manage their content's relationship with AI systems.
Data Sourcing and Model Training
OpenAI trains each new generation of foundation models from scratch using datasets that increase in scale and diversity to improve the models' knowledge, understanding, and linguistic capabilities. The company relies on three primary categories of data:
- Publicly Available Data: This includes industry-standard machine learning datasets and web crawls. OpenAI excludes sources with paywalls, those that primarily aggregate personally identifiable information (PII), sources that violate company policies, or those that have explicitly opted out.
- Proprietary Data Partnerships: OpenAI partners to access non-public content, such as archives and metadata. Examples include a private video library for Sora's training and a partnership with the Government of Iceland to preserve native languages. OpenAI states it does not pursue paid partnerships for information that is already purely publicly available.
- Human Feedback: Models are refined using feedback from AI trainers, red teamers, employees, and users who have opted into model improvements via their data control settings.
Publisher Partnerships and Discovery
To transition from an "attention economy" to one that empowers creators, OpenAI is integrating publisher content directly into its products to increase visibility and connection between users and original sources.
- Source Attribution: ChatGPT has received updates to improve source links, providing users with better context and giving publishers new ways to connect with audiences.
- Strategic Partnerships: OpenAI has established partnerships with global news organizations, including the Financial Times, Le Monde, Prisa Media, and Axel Springer, to display their content in ChatGPT and improve the tools available to newsrooms.
- Educational Partnerships: Collaborations with Khan Academy and ExamSolutions have been used to improve the model's mathematical performance to expand personalized AI tutoring.
Data Privacy and Usage Policies
OpenAI maintains specific boundaries regarding the data used for training to protect privacy and business confidentiality:
- Business Data: OpenAI does not train on data from ChatGPT Team, ChatGPT Enterprise, or the API Platform.
- User Control: Users of ChatGPT Free and Plus can manage whether their data contributes to future model improvements through their account settings.
- Sensitive Information: The company employs techniques to reduce the processing of personal and sensitive information and trains models specifically to avoid providing private or sensitive data about individuals.
Sources
- OriginalOur approach to data and AI