Reducing Bias and Improving Safety in DALL·E 2

OpenAI has introduced a new system-level technique for DALL·E 2 to ensure that generated images of people more accurately reflect global population diversity. This update addresses biases in training data and strengthens safety protocols to prevent the creation of deceptive content and the likeness of public figures.

Improving Representation and Diversity

DALL·E 2 now employs a mitigation technique that triggers when a user provides a prompt describing a person without specifying race or gender (e.g., "firefighter"). This system-level intervention is designed to reduce the likelihood of the model defaulting to narrow stereotypes.

Internal evaluations indicate that this technique is highly effective: users were 12x more likely to report that DALL·E images included people of diverse backgrounds after the mitigation was applied.

Safety System Enhancements

Following a research preview phase that began in April 2022, OpenAI has implemented several safety updates based on feedback from early users who flagged sensitive and biased images. These improvements include:

Deceptive Content and Public Figures

To minimize the risk of misuse for creating deceptive content, DALL·E 2 now rejects:

  • Image uploads containing realistic faces.
  • Attempts to generate the likeness of public figures, including prominent political figures and celebrities.

Content Filtering and Monitoring

OpenAI has refined its safety infrastructure to better balance creative expression with policy enforcement:

  • Improved Accuracy: Content filters have been updated to be more effective at blocking prompts and image uploads that violate OpenAI's content policy.
  • Enhanced Monitoring: Both automated and human monitoring systems have been refined to better guard against system misuse.

Deployment Strategy

OpenAI is expanding access to DALL·E 2 as part of a responsible deployment strategy. The lab states that increasing the user base allows them to learn more about real-world usage patterns, which in turn enables continuous iteration on safety systems and a deeper understanding of how AI systems reflect biases present in their training data.

Sources