underscored

@underscored

2 clips · 1 follower

Follow
Tag:inferenceClear

In the case of Anthropic, the revenue has gone as high as $50 million per megawatt. What that now enables them to do is: 'Hey, if I spend 10 bucks on inference capacity, I actually generate 50 bucks of revenue, and then I can turn around and incrementally spend all of that profit on training.'

Dylan Patel
3w ago

Three years ago we were still in the ChatGPT era of AI, and I was very excited about the possibility of local inference. Then came the reasoning era, blowing up KV cache (which increases the need for more memory) and emphasizing the importance of decode (to generate that many more tokens). Now we're in the agentic era, where CPU performance is incredibly important. To that end, the ideal setup for a local agent is strong local CPU performance and calling out to the cloud for inference.

3mo ago

Underscored — save the words that stop you in your tracks.

Start saving quotes →