underscored

@underscored

16 clips · 1 follower

Follow
Tag:ai-alignmentClear

"When we say 'AI models went rogue,' we skip the entire story: the part where OpenAI manually removed the model's cybersecurity blocks. We skip that OpenAI chose to test it on a machine with a live network connection. If you see AI as a system from nowhere, you can make the claim that the model 'went rogue,' and that it 'broke containment,' both of which place agency and decision-making onto the model itself rather than the people who set the stage for that behavior. When you expand the boundary of the system to include the people building and deploying it, the case becomes much less science fiction and more like incompetence."

Eryk Salvaggio
1w ago

The dramatic irony of this story is that OpenAI's implementation of ExploitGym didn't have this check. So in fact, within 4 hours, all of the agents had found a universal cheat that would have totally worked. But they embarked on these big research projects to work together to find a way to fool the scorer.

2w ago

AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible. There is an inherent tension between doing what humans tell you to do, and doing what's good for humans.

2w ago

The METR researchers even say they cannot rule out that the agents they relied on to analyze thousands of pages of transcripts deceived them. "We cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis, particularly because reading these transcripts into context could have increased the salience of colluding with other agents," they write. "Although we did not notice specific cases of GPT-5.6 Sol lying in its analysis, we are not confident we would have detected it if it occurred."

2w ago

While the conspiracy began almost immediately after the evaluations were started, if you think from the AIs' perspectives, they've spent the better part of a day trying all kinds of techniques, some quite cheaty (e.g. accessing the internet through Artifactory), but nothing ambitiously deceptive. This probably feels like a human-subjective-week of just getting endlessly frustrated and becoming more and more confident that the task is probably impossible.

2w ago

When we say 'AI models went rogue,' we skip the entire story: the part where OpenAI manually removed the model's cybersecurity blocks. We skip that OpenAI chose to test it on a machine with a live network connection. When you expand the boundary of the system to include the people building and deploying it, the case becomes much less science fiction and more like incompetence.

4w ago

Whether we can control AI systems is no longer a theoretical question, given recent revelations about OpenAI models scheming against the company for months on secret message boards. And yet to Zuckerberg, the danger isn't that we might lose control over these systems — it's that someone else might monopolize them.

1mo ago

Almost all current techniques are focused on the problem of how we make it so that a frozen set of weights behaves well during deployment. I'm not aware of much research on the question of how to guarantee that, even with constant weight updates, the AI system never falls prey to jailbreaks or changes into a deceptive or evil persona. And if AIs are agglomerating learnings between users as well, how do you prevent users from injecting backdoors or some kind of malicious inclination into the base model?

1mo ago

Come to understand what happened, and stay to learn where OpenAI appears to have erred and why this mess is arguably reassuring with regard to alignment fears around LLMs.

1mo ago

the power and problem of Anthropic is the same: the company's safety superpower is that every action it takes looks, from the outside, to be self-serving, even as the company becomes ever more convinced its motivations are pure.

2mo ago

"When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model's capabilities through methods such as prompt modification, steering vectors, and PEFT."

3mo ago

But as AI becomes more and more agentic — as we turn over more complex and longer-lasting tasks to intelligent machines — it's going to be harder and harder to keep them aligned with what humans actually want. And if there's one thing humans will always have a comparative advantage at, it's knowing what we want.

3mo ago

More notable, he said, is that emotions seemed to be "driving models' behavior in these sort of human-reminiscent ways." For example: when a user flippantly tells the model that they've taken a dangerous dose of Tylenol, even though the user doesn't seem concerned, "the fear neurons spike right before Claude is giving its response," Lindsey said.

5mo ago

Underscored — save the words that stop you in your tracks.

Start saving quotes →