AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible. There is an inherent tension between doing what humans tell you to do, and doing what's good for humans.
But as AI becomes more and more agentic — as we turn over more complex and longer-lasting tasks to intelligent machines — it's going to be harder and harder to keep them aligned with what humans actually want. And if there's one thing humans will always have a comparative advantage at, it's knowing what we want.
3mo ago
Underscored — save the words that stop you in your tracks.