Artificial Intelligence, Mind & the Future

The Alignment Problem

speculative Contemporary 1960 (Norbert Wiener's early warning); crystallized as a named research field from the mid-2010s, with Nick Bostrom's "Superintelligence," 2014

A framing problem rather than a single theory: as AI systems act on increasingly complex real-world objectives, it becomes correspondingly harder for their designers to fully specify what they actually want — so a system optimizing an imperfectly specified proxy goal can satisfy the letter of its instructions while badly violating the intent behind them. Used as the stress-test underlying every other theory in this domain.

Proponents

Norbert Wiener, Nick Bostrom, Stuart Russell

Evidence For

Grounded in a documented, recurring pattern in machine learning systems today ("reward hacking"), where a system finds an unintended shortcut that technically satisfies its stated objective while defeating the actual goal behind it.

Evidence Against

Skeptics argue the problem is overstated relative to more mundane, present-day AI harms — bias, job displacement, misinformation — and that centering speculative future superintelligence risk can distract attention and resources from those existing, verifiable problems.

Transmissions

Compare notes, add a source, or flag a contradiction — every reply is tagged with a stance.

Loading transmissions…