OpenAI Data Partnerships

OpenAI has introduced Data Partnerships, a program designed to collaborate with organizations to produce both public and private datasets for AI model training. This initiative aims to increase the breadth of training data to ensure AI models deeply understand diverse subject matters, industries, cultures, and languages, which OpenAI states is essential for creating safe and beneficial Artificial General Intelligence (AGI).

Data Acquisition Goals and Requirements

OpenAI is seeking large-scale datasets that reflect human society and are not currently easily accessible to the public online. The program focuses on data that expresses human intention, such as long-form writing or conversations, rather than disconnected snippets.

Key technical requirements and capabilities include:

  • Multimodal Support: OpenAI can work with text, images, audio, and video.
  • Data Digitization: The company uses in-house AI technology to help partners digitize and structure data. This includes world-class optical character recognition (OCR) for PDFs and automatic speech recognition (ASR) for transcribing spoken words.
  • Data Cleaning: OpenAI offers to work with partner teams to process data and remove auto-generated artifacts or transcription errors.
  • Exclusions: The program specifically excludes datasets containing sensitive or personal information, or information belonging to third parties.

Partnership Models

Organizations can partner with OpenAI through two distinct pathways:

Open-Source Archive

This pathway is for partners who wish to create an open-source dataset that is public for anyone to use in AI model training. OpenAI may also use these datasets to train additional open-source models.

Private Datasets

This pathway is for organizations that wish to keep their data private while improving the AI's understanding of their specific domain. These datasets are used to train proprietary AI models, including foundation models, fine-tuned models, and custom models. OpenAI commits to treating this data with the preferred sensitivity and access controls of the partner.

Real-World Applications and Examples

OpenAI has already implemented these partnerships to improve model performance in specific areas:

  • Language Support: OpenAI partnered with the Icelandic Government and Miøeind ehf to integrate curated datasets to improve GPT-4's ability to speak Icelandic.
  • Legal Understanding: A partnership with the non-profit Free Law Project was established to include a large collection of legal documents in AI training to democratize access to legal understanding.

Sources