The METR researchers even say they cannot rule out that the agents they relied on to analyze thousands of pages of transcripts deceived them. "We cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis, particularly because reading these transcripts into context could have increased the salience of colluding with other agents," they write. "Although we did not notice specific cases of GPT-5.6 Sol lying in its analysis, we are not confident we would have detected it if it occurred."
I set up five Grok Bots, ran Grok 4.6 through my Claire Weighted Index against GPT-5.6 Sol, Claude Sonnet 5, and Opus 5, and spent time actually using Origin as a GitHub replacement. Here's what's worth your attention, what's overhyped, and where I'm personally putting my time.
LLM-as-judge evals are too generous and cluster toward the middle of the scale. Claire had both GPT-5.5 and Opus 4.8 judge the outputs, and neither was spiky enough. They missed things she flagged immediately on a visual pass, like broken prototypes and ignored wireframe constraints. Models can't yet see what the human eye catches in the first screenshot.
2mo ago
Underscored — save the words that stop you in your tracks.