underscored

@underscored

1 clip · 1 follower

Follow
Tag:reinforcement-learningClear

naive policy gradient RL has to figure out which of the 100k+ tokens in your trajectory actually got you the right answer, while AlphaGo's MCTS suggests a strictly better action every single move, giving you a training target that sidesteps the credit assignment problem.

Dwarkesh Patel
1w ago

Underscored — save the words that stop you in your tracks.

Start saving quotes →