OpenAI Approach to AI Safety
OpenAI utilizes a multi-layered safety framework to ensure that powerful AI systems are safe and broadly beneficial. This approach combines rigorous pre-deployment testing, iterative real-world deployment, and continuous alignment research to mitigate risks and improve model behavior.
Pre-Deployment Testing and Alignment
OpenAI conducts rigorous testing and engages external experts for feedback before releasing any new system. The organization uses techniques such as reinforcement learning with human feedback (RLHF) to improve model behavior and alignment.
For GPT-4, OpenAI spent more than six months working across the organization to make the model safer and more aligned prior to its public release.
Iterative Deployment and Real-World Learning
OpenAI believes that lab-based testing has limits and cannot predict all beneficial or abusive uses of the technology. To address this an iterative deployment strategy is used:
- Gradual Release: New AI systems are released cautiously to a steadily broadening group of people with substantial safeguards in place.
- Monitoring for Misuse: By making models available via services and APIs, OpenAI can monitor for and take action against misuse, and build mitigations based on real-world data rather than theories.
- Societal Adjustment: Iterative deployment allows society and stakeholders to have a significant say in how AI develops and provides time for the technology to be adopted more effectively.
Child Safety and Content Moderation
Protecting children is a critical focus of OpenAI's safety efforts. The organization maintains strict usage policies against generating hateful, harassing, violent, or adult content.
Key safety metrics for GPT-4 include:
- Disallowed Content: GPT-4 is 82% less likely to respond to requests for disallowed content compared to GPT-3.5.
- Child Safety Tools: OpenAI uses Thorn’s Safer to detect, review, and report Child Sexual Abuse Material (CSAM) uploaded to image tools to the National Center for Missing and Exploited Children.
- Age Requirements: Users must be 18 or older, or 13 or older with parental approval.
OpenAI also works with partners like Khan Academy to develop tailored safety mitigations for specific use cases, such as AI-powered assistants for students and teachers.
Privacy and Data Handling
OpenAI's large language models are trained on a broad corpus of including publicly available content, licensed content, and licensed content generated by human reviewers. Data is used to make models more helpful, not for selling services, advertising, or building profiles of people.
To protect private individuals, OpenAI employs several strategies:
- Removal of Personal Information: Personal information is removed from training datasets where feasible.
- Citing Rejections: Models are fine-tuned to reject requests for the personal information of private individuals.
- Deletion Requests: The organization responds to requests from individuals to delete their personal information from their systems.
Factual Accuracy and Hallucinations
OpenAI identifies that large language models predict the next series of words based on patterns, which can lead to factual inaccuracies. To improve accuracy, OpenAI leverages user feedback on incorrect ChatGPT outputs to as a main source of data.
GPT-4 is 40% more likely to produce factual content than GPT-3.5, based on this iterative improvement process.
Global Governance and Continued Research
OpenAI believes that global governance is necessary to ensure AI development is governed effectively at a scale that prevents companies from "cutting corners" to get ahead. OpenAI actively engages with governments to ensure rigorous safety evaluations are adopted as regulation.
The organization maintains that improving AI safety and capabilities should occur simultaneously, as more capable models are better at following instructions and are easier to steer. OpenAI will continue to enhance safety precautions as AI systems evolve, and may take longer than six months of additional safety work for future models.
Sources
- OriginalOur approach to AI safety