In a separate open-ended research project where the AI was tasked with proposing and testing hypotheses about an open problem in AI safety, Anthropic's models significantly outperformed two human researchers (97% performance improvement vs. 23%) when given a similar time budget (5 to 7 days).
It's bizarre to have something with this superhuman breadth that knows all the fields so well, and yet isn't finding those lightning bolts that connect them. I think we're starting to see sparks of it actually finding connections between the things it's an expert at.
2mo ago
Underscored — save the words that stop you in your tracks.