The thing that I think is underappreciated is that transformers are probably here to stay. People have been trying to beat transformers for years and years and years, and nobody has been able to do it at scale. And so if you believe that transformers are here to stay, then you can build a chip that only runs transformers.
You might also like
Three years ago we were still in the ChatGPT era of AI, and I was very excited about the possibility of local inference. Then came the reasoning era, blowing up KV cache (which increases the need for more memory) and emphasizing the importance of decode (to generate that many more tokens). Now we're in the agentic era, where CPU performance is incredibly important. To that end, the ideal setup for a local agent is strong local CPU performance and calling out to the cloud for inference.
Ben Thompson
The other observation is that the precision will almost always be higher in the accumulation step than in the multiplication step. This is specific to AI chips. You're multiplying low-precision numbers, and then when you accumulate, errors accumulate quickly, so you need more precision there.
Reiner Pope
the inference that will matter most in the future, at least in terms of market size, will be "agentic inference", where humans aren't involved at all. That will lead to very different trade-offs in architectures, and is good news for both China and space (but maybe not Nvidia).
Ben Thompson