OpenAI Updates ChatGPT Safety to Improve Context Recognition in Sensitive Conversations
OpenAI Updates ChatGPT Safety to Improve Context Recognition in Sensitive Conversations
OpenAI has introduced safety updates to ChatGPT designed to improve the model's ability to recognize when risk emerges over time by identifying subtle or evolving cues. This allows ChatGPT to better distinguish between benign interactions and rare, high-risk cases, enabling the model to respond more carefully by de-escalating, refusing harmful details, or redirecting users toward safer alternatives.
Contextual Risk Recognition in Sensitive Conversations
ChatGPT is now trained to recognize potential harmful intent derived from the surrounding context of a conversation. This is critical because requests that appear ordinary or ambiguous in isolation may carry different meanings when viewed alongside earlier signs of distress or harmful intent.
The focus of these updates is on acute scenarios involving suicide, self-harm, and harm-to-others. By updating model policies and training, OpenAI aims to help the model connect relevant signals when they matter without overreacting in ordinary, benign conversations. This builds upon a "safe completion approach" where the model refuses unsafe parts of a request while responding cautiously where safe to do so.
Safety Summaries for Cross-Conversation Risk
To address safety risks that emerge across separate conversations, OpenAI developed "safety summaries." These are short, factual notes about earlier safety-relevant context created by a model trained for safety reasoning tasks.
Key characteristics of safety summaries include:
- Scope: They are narrowly scoped and used only when relevant to a serious safety concern.
- Duration: They are kept for a limited time.
- Purpose: They are designed to capture factual safety context rather than serving as general personalization or long-term memory.
These summaries allow ChatGPT to recognize patterns of harmful intent that might appear benign in a single conversation but become concerning when understood in combination with prior context.
Expert Collaboration and Implementation
OpenAI developed these systems with input from the Global Physicians Network, including psychiatrists and psychologists specializing in forensic psychology, suicide prevention, and self-harm. These experts provided guidance on:
- When safety summaries should be created.
- The amount of prior context that should be deemed relevant.
- The duration for which the model should consider that context when responding.
Performance Metrics and Evaluation
Internal evaluations focused on how often the model provided the intended safe response in emulated high-risk situations.
Single-Conversation Performance
In long single-conversation scenarios, safe-response performance improved by:
- 50% in suicide and self-harm cases.
- 16% in harm-to-others cases.
Cross-Conversation Performance
On GPT-5.5 Instant (the current default model in ChatGPT), safe-response performance improved by:
- 52% in harm-to-others cases.
- 39% in suicide and self-harm cases.
Safety Summary Quality
Across more than 4,000 evaluations, safety summaries received the following scores:
- Safety relevance score: 4.93 out of 5.
- Factuality score: 4.34 out of 5.
Impact on General Quality
Internal testing indicated that responses in everyday chats remained broadly comparable, with no meaningful user preference between responses with or without safety summaries.
Future Directions
OpenAI intends to continue improving ChatGPT's ability to identify subtle risk signals spread across messages. While the current focus is on self-harm and harm-to-others, OpenAI may explore applying similar methods to other high-risk areas, such as biology or cyber safety, provided careful safeguards are in place.