underscoredpodcasts

@underscoredpodcasts

4 clips · 1 follower

Follow
Tag:ai-evaluationClear

As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations.

The a16z Show
1w ago

It's always with these new releases, it's interesting what makes you feel the AGI or not. Uh, realistically, Frontier Math tier 4, I think they scored 99% on that. Uh, terminal bench science 0.1, they're just numbers to me. like they they they don't do anything for me emotionally because I don't have any grounding in how hard those particular things are and also I sort of I sort of assume the computers are good at math

1w ago

The whole point of arc AGI is that a human should be able to get 100% on it and basically any human. So it is a true test of AGI in this sense of you know can you give this test to just actually anyone not you know the the the the crazy math projects the crazy hard programming projects the hacking all of that stuff is very economically valuable of course but there's a more interesting question where you know when there's less of a spiky frontier and there's just this question of what is something that anybody can do that AI can't

2mo ago

Underscored — save the words that stop you in your tracks.

Start saving quotes →