underscored

@underscored

1 clip · 1 follower

Follow
Tag:continual-learningClear

Almost all current techniques are focused on the problem of how we make it so that a frozen set of weights behaves well during deployment. I'm not aware of much research on the question of how to guarantee that, even with constant weight updates, the AI system never falls prey to jailbreaks or changes into a deceptive or evil persona. And if AIs are agglomerating learnings between users as well, how do you prevent users from injecting backdoors or some kind of malicious inclination into the base model?

Dwarkesh Patel
2w ago

Underscored — save the words that stop you in your tracks.

Start saving quotes →