As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations.
The book's big thesis is that this kind of systemic measurement is really a dehumanizing act, one that flattens humanity and human concerns into something, well, artificial.
Many pilots rely on self selected enthusiasts, vendor methodologies, and rough time saved calculations rather than robust measures of service quality or error rates.
5mo ago
Underscored — save the words that stop you in your tracks.