How ChatGPT Learns and Protects User Privacy
How ChatGPT Learns and Protects User Privacy
OpenAI uses a combination of public data, partnerships, and user-generated content to train ChatGPT, employing the OpenAI Privacy Filter to mask personal information and providing users with granular controls to opt out of model improvement.
Data Sources for Model Training
ChatGPT is trained on a diverse set of information sources to build general world knowledge and improve reliability and safety. These sources include:
- Publicly Available Information: OpenAI uses content that is freely and openly accessible on the internet, such as public blog posts and online discussion forums.
- Partnerships: Data accessed through strategic partnerships.
- User and Researcher Contributions: Information provided or generated by users, contractors, and researchers.
Reducing Personal Information via OpenAI Privacy Filter
To prevent models from learning private information about individuals, OpenAI applies safeguards to datasets before training. A primary tool in this process is the OpenAI Privacy Filter, which identifies and masks personal information within text.
OpenAI states that internal evaluations show the Privacy Filter is more effective at removing personal information than any other similar tool. This filter is integrated at multiple stages of the training pipeline, applying to both public datasets and user conversations (provided the user has the "Improve the model for everyone" setting enabled).
User Privacy Controls in ChatGPT
OpenAI provides several mechanisms for users to control how their data is handled and whether it contributes to future model iterations:
Model Training Opt-Out
Users can disable the "Improve the model for everyone" setting under Settings > Data Controls. When this setting is turned off, new conversations are still saved to the chat history but are not used to train ChatGPT.
Temporary Chat
Temporary Chats provide a session-based privacy mode. These conversations:
- Do not appear in chat history.
- Do not create memories.
- Are not used to improve OpenAI models.
- Are retained for 30 days for safety purposes before being deleted.
Memory Management
The Memory feature allows ChatGPT to remember specific details across sessions. This is entirely optional; users can review, edit, or delete specific memories or disable the feature entirely to prevent ChatGPT from saving or referencing past interactions.
Account and Data Management
Users have access to a privacy request portal to submit privacy requests, export their ChatGPT data, or delete their account entirely. OpenAI advises users not to share sensitive information in ChatGPT that they would not want to be reviewed or used.
Balancing Privacy and Safety
OpenAI maintains that protecting privacy and addressing risks of harm must occur simultaneously. The company continues to develop systems to detect and respond to credible threats of violence while maintaining existing privacy safeguards. As model capabilities increase, OpenAI commits to improving safeguards and providing clearer, more practical ways for users to decide how their information is used.