why guardrails designed to stop AI-powered attackers can also prevent security teams from doing their jobs
You might also like
When we say 'AI models went rogue,' we skip the entire story: the part where OpenAI manually removed the model's cybersecurity blocks. We skip that OpenAI chose to test it on a machine with a live network connection. When you expand the boundary of the system to include the people building and deploying it, the case becomes much less science fiction and more like incompetence.
Eryk Salvaggio
If 'pacing' becomes necessary, we think it should consist of two parts: first, specifying thresholds for when automated AI R&D is likely to pose severe risks; and second, if a threshold is exceeded, incentivizing AI companies to reallocate resources away from the most risky research, and towards activities that make further automation safer, or diffuse the benefits of existing AI faster. Without preparation now, however, our preferred pacing strategy will be impossible to implement.
Tim Fist
In the future, our capacity to steward our votes and our capital, and to make sense of what's happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.
Dwarkesh Patel