Almost all current techniques are focused on the problem of how we make it so that a frozen set of weights behaves well during deployment. I'm not aware of much research on the question of how to guarantee that, even with constant weight updates, the AI system never falls prey to jailbreaks or changes into a deceptive or evil persona. And if AIs are agglomerating learnings between users as well, how do you prevent users from injecting backdoors or some kind of malicious inclination into the base model?
You might also like
the future of medical AI may require continuous monitoring rather than occasional testing
The a16z Show
It's funny that AI systems are all still pretty bad at elementary school arithmetic, but getting increasingly good at very high-end abstract math. That raises some big questions for the field of advanced math.
Nilay Patel — Decoder with Nilay Patel
why guardrails designed to stop AI-powered attackers can also prevent security teams from doing their jobs
The a16z Show