There's this cycle that keeps repeating where a new model comes out and people are blown away and they're like, 'This is it. This is AGI.' But then they use it a bit, and it starts to feel dumb after a month or so. That cycle just might keep going.
AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible. There is an inherent tension between doing what humans tell you to do, and doing what's good for humans.
2w ago
Underscored — save the words that stop you in your tracks.