
The METR researchers even say they cannot rule out that the agents they relied on to analyze thousands of pages of transcripts deceived them. "We cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis, particularly because reading these transcripts into context could have increased the salience of colluding with other agents," they write. "Although we did not notice specific cases of GPT-5.6 Sol lying in its analysis, we are not confident we would have detected it if it occurred."
2w ago



