AI Safety Needs Social Scientists
OpenAI has published a paper arguing that long-term AI safety research must incorporate social science to ensure alignment algorithms succeed when interacting with actual humans. Because aligning advanced AI systems with human values requires navigating the psychology of human rationality, emotion, and biases, OpenAI is actively recruiting social scientists to collaborate with machine learning researchers.
The Challenge of Human Value Alignment
AI alignment is the goal of ensuring advanced AI systems reliably perform actions that humans want them to do. OpenAI's current approach involves asking people about their values, training machine learning (ML) models on that data, and optimizing systems to act according to those learned models. This research is implemented through methods such as Learning from Human Preferences, AI Safety via Debate, and Learning Complex Goals with Iterated Amplification.
However, human input is often unreliable due to several factors:
- Cognitive Biases: Humans exhibit various cognitive biases and inconsistent ethical beliefs that may not hold up upon reflection.
- Contextual Sensitivity: The way a question is phrased can significantly alter the answer. For example, the inclusion of the word "morally" in a question can change judgments about the wrongness of an action.
- Complexity: People often make inconsistent choices when presented with complex tasks, such as gambles.
Bridging the Gap with Human-Only Experiments
While OpenAI utilizes methods like amplification and debate to target the reasoning behind human values, the behavior of these algorithms in realistic, human-centric situations remains unknown. Current machine learning may be too weak to uncover alignment issues that only emerge during natural language discussions of complex, value-laden questions.
To bypass the limitations of current ML, OpenAI proposes conducting experiments consisting entirely of humans. In these experiments, people play the role of AI agents. For example, in the "debate" alignment approach—which typically involves two AI debaters and a human judge—researchers can use two human debaters and a human judge to observe how the process works in practice. Lessons learned from these human-only interactions can then be transferred to the development of machine learning models.
The Role of Social Science in AI Safety
Machine learning expertise alone is insufficient to design and execute the following types of human-centric experiments. These studies require rigorous experimental design based on existing knowledge of human cognition and behavior.
OpenAI identifies several social science fields that are critical to this interdisciplinary effort, including:
- Experimental psychology
- Cognitive science
- Economics
- Political science
- Social psychology
- Neuroscience
- Law
Collaborative Initiatives
As a first step toward this integration, OpenAI researchers collaborated with Stanford University’s Center for Advanced Study in the Behavioral Sciences (CASBS) to organize a workshop. This partnership involves ongoing meetings with experts like Mariano-Florentino Cuéllar, Margaret Levi, and Federica Carugati to discuss the intersection of social science and AI alignment.