OpenAI Voice Engine: Technical Implementation and Safety Framework

OpenAI has developed Voice Engine, a synthetic voice technology capable of creating custom voices. The technology has been integrated into ChatGPT's Voice Mode and the Text-to-Speech (TTS) API, while its broader deployment has been managed through a controlled, iterative framework to mitigate risks associated with synthetic audio.

Technical Deployment and Product Integration

Voice Engine has been utilized in three primary product implementations since September 2023:

  • ChatGPT Voice Mode: Launched in September 2023, this feature uses voices created solely from real voices that were carefully selected through a process involving professional voice actors, talent agencies, and industry advisors starting in May 2023.
  • TTS API: Released in November 2023, this API provides six preset voices. Each voice was created using 15-second audio samples from professional voice actors.
  • Custom Voice Previews: In March 2024, OpenAI previewed the capability to create custom voices with a small set of trusted partners to explore the risks and opportunities of synthetic voice technology.

Safety Research and Policy Goals

OpenAI utilizes an iterative deployment framework to help policymakers and the public understand the risks of synthetic voices. An early internal prototype developed in late 2022—which used a mix of public and private voice samples for internal testing only—was used to demonstrate the technology's potential to global policymakers starting in the summer of 2023.

Through limited partner previews of custom voice capabilities, OpenAI aims to achieve the following security and policy goals:

  • Phasing out voice-based authentication: Reducing reliance on voice as a security measure for accessing sensitive information, such as bank accounts.
  • Protecting individual voices: Exploring policies to protect the use of individuals' voices in AI.
  • Public education: Educating the public on the capabilities and limitations of AI, including the potential for deceptive content.
  • Content provenance: Accelerating the adoption of techniques to track the origin of audiovisual content to distinguish AI-generated audio from real human speech.

Internal Testing and Alignment

Internal prototypes of Voice Engine were used exclusively for alignment and safety research to inform the development of safeguards. OpenAI specifies that the outputs from these early internal tests were reserved for internal testing and were not used to train the models that power their public-facing products.

Sources