Improving Health Intelligence in ChatGPT

Improving Health Intelligence in ChatGPT

GPT-5.5 Instant advances health intelligence

OpenAI has introduced GPT-5.5 Instant, a model that significantly improves ChatGPT's ability to handle health and wellness queries. The model demonstrates a substantial step forward in recognizing urgent care needs, requesting relevant user context, explaining uncertainty, and simplifying complex health information.

GPT-5.5 Instant is available to all free users in ChatGPT, expanding access to these health-related capabilities. On the most challenging health evaluations, its performance is now comparable to OpenAI's frontier Thinking models.

Measuring health performance and accuracy

OpenAI uses health-specific evaluations to ensure responses are accurate, understandable, and grounded in good judgment. Progress is measured using the following frameworks:

Evaluation Frameworks

  • HealthBench and HealthBench Professional: These evaluations utilize realistic health conversations and physician-written rubrics to assess accuracy, safety, communication, context awareness, completeness, and appropriate escalation.
  • Physician-led Comparison: A panel of physicians compared 3,500 responses from GPT-5.5 Instant against responses written by physicians (who had unlimited time and internet access but no AI) and older models. GPT-5.5 Instant was rated higher across all criteria, including accuracy, communication, completeness, instruction following, and health decision helpfulness.

Production Traffic Monitoring

To track real-world performance, OpenAI uses privacy-preserving monitors on production traffic. Analysis of billions of weekly messages shows that the rate of responses containing at least one flagged factuality issue has decreased by 71% over the last two months.

Physician-led model improvement

The improvements in GPT-5.5 Instant are driven by a global network of more than 260 physicians across 60 countries, 49 languages, and 26 medical specialties. This network provides the medical expertise necessary to define "good" responses in real-world health scenarios.

The Role of Physician Feedback

  • Response Review: Physicians have reviewed over 700,000 example model responses reflecting real-world usage by patients and clinicians.
  • Rubric Development: Physician feedback is used to create rubrics and evaluation criteria that help researchers measure if responses are accurate, safe, clear, complete, and appropriately cautious.
  • Failure Mode Identification: GPT-5.5 Instant shows fewer failure modes than previous models and human physicians, specifically in areas such as tailoring responses to local healthcare contexts, identifying red flags, and requesting necessary additional context from the user.

Integration with broader healthcare tools

These advancements in general health intelligence support OpenAI's wider healthcare initiatives, including specialized tools such as ChatGPT for Clinicians and OpenAI for Healthcare, which assist medical professionals with documentation, research, and care delivery.

Sources