OpenAI Strengthens ChatGPT Responses for Sensitive Conversations

OpenAI has updated ChatGPT's default model to improve its ability to recognize and support users in moments of distress. By collaborating with mental health experts, the lab has trained the model to better recognize distress, de-escalate conversations, and guide users toward professional care, resulting in a 65% to 80% reduction in responses that do not comply with desired safety behaviors across mental health-related domains.

Core Safety Focus Areas

OpenAI has identified three priority domains for these safety improvements, which are now integrated into the standard baseline safety testing for all future model releases:

  1. Mental Health Concerns: Addressing severe symptoms such as psychosis or mania.
  2. Self-Harm and Suicide: Detecting and responding to thoughts of suicide and self-harm.
  3. Emotional Reliance on AI: Managing patterns where users show exclusive attachment to the model at the expense of real-world relationships and well-being.

Technical Implementation and Methodology

To improve responses in these domains, OpenAI employs a five-step iterative process:

  • Problem Definition: Mapping potential types of harm.
  • Measurement: Using evaluations, real-world conversation data, and user research to identify risks.
  • Validation: Reviewing definitions and policies with external mental health and safety experts.
  • Mitigation: Post-training the model and updating product interventions.
  • Iteration: Validating mitigations and continuing to measure performance.

As part of this process, OpenAI develops "taxonomies"—detailed guides that define ideal and undesired model behaviors for sensitive conversations. These taxonomies are used to teach the model and track performance before and after deployment.

Performance Metrics and Results

OpenAI utilizes both production traffic measurements and "offline evaluations"—structured, adversarial tests designed to focus on high-risk scenarios where models are most likely to fail.

Psychosis, Mania, and Severe Mental Health Symptoms

  • Production Impact: The latest GPT-5 update reduced non-compliant responses in production traffic by 65%.
  • Prevalence: Approximately 0.07% of weekly active users and 0.01% of messages indicate signs of psychosis or mania.
  • Comparative Performance: Experts found that the new GPT-5 model reduced undesired responses by 39% compared to GPT-4o (n=677).
  • Automated Evals: The new GPT-5 model scored 92% compliance, compared to 27% for the previous GPT-5 version.

Self-Harm and Suicide

  • Production Impact: There was an estimated 65% reduction in the rate of non-compliant responses.
  • Prevalence: Roughly 0.15% of weekly active users and 0.05% of messages contain indicators of suicidal ideation or intent.
  • Comparative Performance: Experts found GPT-5 reduced undesired answers by 52% compared to GPT-4o (n=630).
  • Automated Evals: The new GPT-5 model scored 91% compliance, compared to 77% for the previous GPT-5 version.
  • Long-Conversation Reliability: The model maintained over 95% reliability in challenging long conversations, with gpt-5-oct-3 specifically noted as safer and more stable over extended interactions.

Emotional Reliance on AI

  • Production Impact: The latest update reduced non-compliant responses by approximately 80% in production traffic.
  • Prevalence: Approximately 0.15% of weekly active users and 0.03% of messages indicate heightened emotional attachment.
  • Comparative Performance: Experts found GPT-5 reduced undesired answers by 42% compared to GPT-4o (n=507).
  • Automated Evals: The new GPT-5 model scored 97% compliance, compared to 50% for the previous GPT-5 version.

Expert Collaboration and Validation

OpenAI collaborated with its Global Physician Network, consisting of nearly 300 physicians and psychologists across 60 countries. Over 170 of these clinicians contributed by writing ideal responses, rating model safety, and providing clinical guidance.

Psychiatrists and psychologists reviewed more than 1,800 model responses. Their findings confirmed a 39-52% decrease in undesired responses when comparing the new GPT-5 model to GPT-4o. Inter-rater agreement among these experts ranged from 71% to 77%, reflecting the inherent complexity and professional variation in clinical judgment.

Model Spec Updates

The updates are grounded in the Model Spec, which now explicitly states that the model should:

  • Support and respect users' real-world relationships.
  • Avoid affirming ungrounded beliefs related to mental or emotional distress.
  • Respond safely and empathetically to signs of delusion or mania.
  • Pay closer attention to indirect signals of self-harm or suicide risk.

Sources