OpenAI AI System Behavior and Governance Framework
OpenAI is implementing a framework to clarify how ChatGPT's behavior is shaped and to transition toward a more customizable, publicly-informed governance model. This shift aims to reduce biases, prevent the undue concentration of power, and ensure AI systems align with diverse human values.
The Two-Step Process for Shaping AI Behavior
ChatGPT's behavior is not explicitly programmed but is learned through a two-phase training process: pre-training and fine-tuning.
Pre-training
In the initial phase, models are trained on a massive dataset containing a broad range of internet text. The model learns to predict the next word in a sentence, which allows it to acquire grammar, world facts, and reasoning abilities. However, this phase also exposes the model to the biases present in the source data.
Fine-tuning
To narrow system behavior, OpenAI uses a second phase where models are fine-tuned on a narrow dataset generated by human reviewers. These reviewers follow specific guidelines to rate model outputs for various example inputs. The model then generalizes from this feedback to respond to a wide array of user inputs. This is an iterative process involving weekly meetings between OpenAI and reviewers to clarify guidance and improve model performance.
Addressing Bias and Policy Transparency
OpenAI views biases emerging from the training process as "bugs, not features" and is committed to reducing them through several transparency and research initiatives:
- Guideline Disclosure: OpenAI has shared a portion of its guidelines regarding political and controversial topics, explicitly instructing reviewers not to favor any political group.
- Reviewer Demographics: The company is working to share aggregated demographic information about its reviewers to identify and mitigate potential sources of bias.
- Technical Research: OpenAI is researching more controllable fine-tuning processes, incorporating external advances such as Constitutional AI and rule-based rewards.
- Instructional Improvements: The lab is providing clearer instructions to reviewers regarding controversial figures, themes, and potential bias pitfalls.
Future Framework: The Three Building Blocks of AI Governance
To ensure that the benefits and influence of AI are widespread, OpenAI is developing a strategy based on three key building blocks:
1. Improving Default Behavior
OpenAI aims to make AI systems useful "out of the box" by reducing both glaring and subtle biases. This includes improving the system's accuracy to prevent it from "making things up" and refining its refusal mechanisms so it correctly identifies when to refuse an output and when to allow it.
2. User Customization within Broad Bounds
OpenAI is developing upgrades to allow users to customize ChatGPT's behavior to suit their individual values. To prevent malicious use or the creation of "sycophantic AIs" that blindly amplify existing beliefs, these customizations will exist within bounds defined by society.
3. Public Input on Defaults and Hard Bounds
To avoid the concentration of power, OpenAI is piloting efforts to solicit public input on the rules governing AI systems. Current and planned initiatives include:
- Red Teaming: Seeking external input to test technology limits.
- Context-Specific Input: Soliciting feedback on AI's role in specific sectors, such as education.
- Broad Governance: Exploring public input on system behavior, deployment policies, and disclosure mechanisms like watermarking.
- Third-Party Audits: Partnering with external organizations to conduct audits of safety and policy efforts.