Claude Opus 4.8 could not build a reliable exploit where Opus 5 succeeded within hours
Assessment
This is the researchers' own account of their process, not an independently reproduced or benchmarked comparison. It is a single anecdotal data point from a party with a direct interest in presenting AI-accelerated exploitation as a dramatic capability jump, published as part of their own research promotion. It is plausible and specific enough to report, but it is not a controlled capability evaluation.
Where this claim appeared
Hacktron AI · 2026-09-13
https://www.hacktron.ai/blog/hacking-openaiWhat “Assessed, Not Confirmed” means
A named source states this as its own assessment, at its own stated confidence, rather than as established fact. Attribution to a nation state usually sits here. The assessment is real and reportable; treating it as settled is the error.
4 of 5 · rating scale
Assessed in
Hacktron used Claude to hack OpenAI's forum, take over an employee's Codex, and open a PR in OpenAI's own repoThink this assessment is wrong? Report an error.