Verification for Cooperation
The essays in this sequence are based on the following premise: any path to a safe and prosperous future of AI will require cooperation. The task of safely navigating the AI transition clearly demands unprecedented collaboration among humans at the scale of individuals, institutions, and nations. But it may not be long before AIs themselves become potential collaborators as well. Next-generation AIs that participate in aligning their successors might either provide invaluable help or disastrous sabotage. Future AI safety may depend on humanity’s ability to cooperate with its creations. This sequence argues that the primary bottleneck on all this potential cooperation is credibility, and that a promising path to greater credibility runs through verification.
The present moment is full of enthusiasm for building, but scarce on clarity when it comes to the long-term goals builders should pursue as capabilities accelerate. In these essays, I take cooperation seriously as one such guiding value for the future. I aim to clarify key concepts for AI-related cooperation, provide motivation and direction for infrastructural investment, and offer design principles for cooperative architecture.
What can we do to facilitate cooperation in the AI age? This is a fundamentally interdisciplinary question, and an underexplored one. Game theory and mechanism design can tell us a great deal about strategic cooperation, but actually using the results for real collaboration requires a grounded philosophical understanding of trust, credibility, and normativity. A full picture of cooperative system design needs ideas from formal methods, international relations, and empirical AI safety as well as novel research. This sequence is an effort to weave these various threads together into a unified theoretical framework. Here, I summarize the main points of each essay in order to sketch the overall arc of the work.
The first essay centers on precommitments and how the ability to meaningfully restrict our future options can unlock opportunities for cooperation. It shows that an important cooperative property of an agent is its translucency, the degree to which others can inspect and confirm its intentions.
The second essay analyzes what it means to verify a precommitment. It argues that verification mechanisms for commitments restructure the need for trust in a partner into trust in a kernel and its context, rather than eliminating trust entirely. This essay introduces some central themes of the research: the value and limits of formal verification, the applicability of open-source game theory, and the prospect of AI-human dealmaking.
A forthcoming essay applies this understanding of verification to compliance, especially in the context of international treaties. It argues that verification mechanisms and shared institutional practices of using them can be a powerful force for internalizing international norms around AI.
Future posts will explore topics like an application of incomplete contract theory to the specification problem in formalization, a taxonomy of structural properties required for an AI to make meaningful precommitments, and a survey of the existing research landscape on AI cooperation and verification. I am particularly interested in future work on mechanism design for self-reinforcing AI-human cooperation mechanisms. This introduction will be updated as new essays are added to the sequence.