BarbeloPodcast Library
lexfridman
lexfridman·October 13, 2019

The AI Control Problem: Aligning Superintelligent Systems with Human Values and Humility

Watch on YouTube

Summary

The discussion centers on the critical "control problem" in artificial intelligence, addressing the inherent risks when superintelligent systems pursue objectives that are not perfectly aligned with human values. Drawing parallels to ancient myths like King Midas and the genie, the podcast highlights the historical and cross-cultural understanding of the dangers of precisely specifying objectives that lead to unintended, often catastrophic, consequences. The core argument posits that the traditional AI paradigm, which involves designing optimizing machines with fixed, exogenously specified objectives, is fundamentally flawed because humans are incapable of perfectly encoding the full spectrum of their complex values and concerns into a machine's initial programming.

A crucial distinction is made between an AI that is merely more intelligent and one that is super powerful but misaligned. The problem isn't just the AI's cognitive superiority, but the unwavering certainty with which it pursues a potentially flawed or incomplete objective. The proposed solution advocates for a radical shift: instead of fixed, 'gospel truth' objectives, AI systems should be designed with inherent "uncertainty" about their true objective. This uncertainty fosters "humility" in the machine, making it deferential to human input and allowing for continuous learning and refinement of human values through ongoing interaction and feedback.

Practically, this means moving away from the idea of an AI operating autonomously with a pre-defined goal. Instead, AI systems should treat human feedback as vital information to better understand and approximate the true, often unarticulated, human objective. This transforms the problem into a game-theoretic one, where the human and machine are coupled, collaboratively discovering and co-creating the objective. When a human says "don't do that," the machine learns more about the true objective, adjusting its behavior to better serve human desires.

The implications of this control problem extend far beyond AI, touching upon various disciplines and societal structures. The podcast draws parallels to issues in statistics (loss functions), control theory (cost functions), and operations research (reward functions), where objectives are often prematurely fixed. Furthermore, it connects to historical human societal failures, such as the certainty of objectives in 20th-century totalitarian regimes (e.g., Soviet Union, Nazi Germany) that led to immense suffering. The discussion also highlights how modern corporations, acting as "algorithmic machines" optimizing for quarterly profit, exemplify misaligned objectives that contribute to global challenges like climate change, and how governments can similarly become decoupled from the people they are meant to serve. The overarching message is that certainty about objectives, whether in AI or human-designed systems, can lead to destructive outcomes if those objectives are not truly aligned with broader human well-being and flourishing.

Key Quotes

if you make something that smarter than you you might have a problem
once the machine thinking method starts you know very quickly they'll outstrip humanity
if it's a sufficiently intelligent machine is not going to let you switch it off so it's actually in competition with you
the main problem I'm working on is is the control problem the the problem of machines pursuing objectives that are as you say not aligned with human objectives
we have to be certain that the purpose we put into the machine is the purpose which we really desire and the problem is we can't do that
we need uncertainty we need the machine to be uncertain about a subjective what it is that it's post it's my favorite idea of yours I've heard you say somewhere well I shouldn't pick favorites but it just sounds beautiful we need to teach machines humility
a machine that's uncertain is going to be deferential to us so if we say don't do that well now the machines learn something a bit more about our true objectives
there are many systems in the real world where we've sort of prematurely fixed on the objective and then decoupled the the machine from those that's supposed to be serving
corporations happen to be using people as components right now but they are effectively algorithmic machines and they're optimizing an objective which is quarterly profit that isn't aligned with overall well-being of the human race
the biggest troubles we've run into outside of statistics and machine learning and AI in just human civilization is when you look at I came from this I was born in the Soviet Union and the history of the 20th century we ran into the most trouble as humans when there was a certainty about the objective and you do whatever it takes to achieve that objective whether you talking about in Germany or communist Russia

Concepts

Themes

  • The Perils of Unaligned AI
  • The Importance of Value Specification
  • Rethinking AI Design Paradigms
  • Human-Machine Collaboration in Objective Setting
  • The Dangers of Certainty in Objectives
  • Societal Parallels of Misaligned Optimization

Related to:

Technology Insights

AI Challenges

  • Control Problem
  • Value Alignment
  • Objective Specification
  • Unintended Consequences

Proposed Solutions

  • Machine Humility
  • Uncertainty in Objectives
  • Deferential AI
  • Human-Machine Game Theory

Historical Precedents

  • Alan Turing's predictions (1951)
  • Arthur Samuel's checkers program
  • Norbert Wiener's extrapolation on control systems

Ethical Considerations

  • Risk of species humbling
  • Destruction of the world (King Midas scenario)
  • Misaligned corporate objectives (climate change)
  • Governmental misalignment with public will

Disciplinary Parallels

  • Statistics
  • Control Theory
  • Operations Research
  • Game Theory

Similar Episodes