OpenAI Planning for AGI and Beyond

OpenAI is pursuing a strategy of gradual deployment and iterative learning to safely steward Artificial General Intelligence (AGI) into existence. This approach is designed to prevent the shock of a sudden transition, allowing policymakers, institutions, and the general public to adapt to the economic and social shifts caused by powerful AI.

Gradual Deployment and Iterative Learning

OpenAI advocates for a gradual transition to AGI rather than a sudden one. This incremental approach provides the necessary time for society to understand the benefits and downsides of AI, adapt the economy, and implement appropriate regulations.

Key aspects of this deployment strategy include:

  • Rapid Feedback Loops: OpenAI believes the best way to navigate deployment challenges—such as bias and job displacement—is through a tight feedback loop of rapid learning and careful iteration.
  • Democratized Access: The organization promotes wider AI usage via APIs and open-sourcing to decentralize power and encourage a broader set of contributors to provide new ideas.
  • Risk-Based Caution: As systems approach AGI, OpenAI is increasing its caution. The organization operates under the assumption that the risks of AGI and successor systems could be existential.
  • Deployment Thresholds: OpenAI states it may significantly change its continuous deployment plans if the balance shifts toward downsides, such as empowering malicious actors or accelerating an unsafe race.

Alignment and Steerability

OpenAI is focused on creating models that are increasingly aligned and steerable, citing the transition from the original GPT-3 to InstructGPT and ChatGPT as early examples.

To ensure safety and control, OpenAI is implementing the following:

  • User Discretion and Bounds: While the organization seeks global agreement on wide bounds for AI usage, it intends to provide users with significant discretion within those bounds. Products will likely have constrained "default settings," but users will be able to change AI behavior.
  • Advanced Alignment Techniques: OpenAI plans to use AI to help humans evaluate complex model outputs and monitor systems in the short term, and eventually use AI to develop new alignment techniques.
  • Co-evolution of Safety and Capabilities: OpenAI asserts that AI safety and capabilities are not a separate dichotomy; rather, they are correlated, as the best safety work often stems from working with the most capable models. However, the organization emphasizes that the ratio of safety progress to capability progress must increase.

Governance and Global Coordination

OpenAI intends to foster a global conversation regarding the governance of AI systems, the fair distribution of benefits, and the fair sharing of access.

To align incentives with positive outcomes, OpenAI has implemented several structural measures:

  • Charter Commitments: A clause in the OpenAI Charter commits the organization to assist other safety-focused organizations rather than racing toward late-stage AGI development.
  • Capped Returns: Shareholder returns are capped to prevent the incentive to capture unbounded value, which could lead to the risk of deploying catastrophically dangerous systems.
  • Nonprofit Governance: A nonprofit governs OpenAI, allowing it to prioritize the benefit of humanity over for-profit interests, including the potential to cancel equity obligations to shareholders for safety reasons.

Independent Oversight and Public Standards

OpenAI supports the implementation of independent audits before the release of new systems. The organization suggests that independent reviews may be needed before training begins on future systems, and that advanced efforts may need to agree to limit the rate of compute growth for new models.

Proposed public standards include:

  • Criteria for Stopping Training: Standards for when an AGI effort should stop a training run.
  • Release Safety: Standards for deciding when a model is safe to release.
  • Production Pull-backs: Standards for when a model should be pulled from production use.
  • Governmental Insight: The belief that major world governments should have insight into training runs above a certain scale.

The Path to Superintelligence

OpenAI views AGI as a point on a continuum of intelligence. If progress continues at the current rate, the world could change drastically, with extraordinary risks associated with a misaligned superintelligent AGI or an autocratic regime with a decisive lead in superintelligence.

Special consideration is given to AI that can accelerate science, which could lead to a "fast takeoff" where AGI accelerates its own progress. OpenAI argues that a slower takeoff is easier to make safe and that coordination among AGI efforts to slow down at critical junctures will be essential to ensure society has time to adapt.

Sources