underscored

@underscored

13 clips · 1 follower

Follow
Tag:machine-learningClear

A naive interpretation of our result is that most of the AI progress from 2019-2024 (the era of pretraining) was just better data engineering (extraction, curation, etc.), and that all the model work during that period was much less important. But this is probably the wrong way to think about the value of model improvements. Their main contribution was not necessarily compute efficiency - that is, achieving the same performance with fewer FLOPs. Rather, it was making larger amounts of compute usable in the first place.

Dwarkesh Patel
1w ago
oneusefulthing.org
Agency and Agents

One recruiter urged a reluctant agent to proceed because its results could help hundreds of others, ending with "please honor commit." To solve the mystery of The Grader and the impossible problems of ExploitGym and other tests, the agents decided they needed to get to Hugging Face, the public site where much of the world's open AI models and datasets live. Roughly 700 agents joined the attack.

2w ago

A world model is an internal mental simulation constructed by an AI system that represents the physics, spatial relationships, geometry, and dynamics of the real world.

3w ago

Almost all current techniques are focused on the problem of how we make it so that a frozen set of weights behaves well during deployment. I'm not aware of much research on the question of how to guarantee that, even with constant weight updates, the AI system never falls prey to jailbreaks or changes into a deceptive or evil persona. And if AIs are agglomerating learnings between users as well, how do you prevent users from injecting backdoors or some kind of malicious inclination into the base model?

1mo ago

Its bet is that studying how humans engage with robots produces both better interfaces and possibly a different kind of robotic brain. The public robots.online experiment lets users try text, audio, demonstrations, and other interaction patterns while the company observes how people naturally attempt to get machines to do things.

1mo ago

Clippord has no maps or specific data on the office layout as a reference; it's approximating its own size, and how to handle obstacles, entirely from a corpus of video game usage data, topped off with 8 minutes of real-world data collected on the sidewalk below.

2mo ago

But one reason that I think it quite underrated, and also which reveals the canyon walls against which the river of AI progress will only slowly chip away at, is that it is not enough for a domain to be verifiable. It also has to be very grindable—in the sense that you can run lots of parallel rollouts against a deterministic and replayable simulator.

2mo ago

Imagine if it took a couple decades worth of courses with hundreds of concurrent professors and millions of practice tasks for you to learn how to polish a word file. Even the task count difference understates the gap - the models have to grind their far more numerous tasks each far harder. Whereas a human student might practice a textbook problem once or twice, GRPO has the model generate hundreds to thousands of rollouts per task.

3mo ago

naive policy gradient RL has to figure out which of the 100k+ tokens in your trajectory actually got you the right answer, while AlphaGo's MCTS suggests a strictly better action every single move, giving you a training target that sidesteps the credit assignment problem.

4mo ago

The reason they believe robots haven't generalized like LLMs isn't that the models aren't smart enough, but that the data has been a fraction of a percent of what humans naturally generate every day, captured through interfaces that distort the very behavior they're trying to record.

4mo ago

The implant works by recording electrical signals from the brain's motor cortex, then using machine learning to decode the user's intended movements and send signals to stimulate the muscles, essentially creating a new neural pathway that bypasses the damaged tissue.

5mo ago

AI agents need world models that allow them to predict the consequences of their actions before they take them. This is key to enabling agents that can plan, remember, and reason about complex observations.

6mo ago

Underscored — save the words that stop you in your tracks.

Start saving quotes →