0x

guest@0xbase ~$ read-only mode. Posting requires EU location.

Bugcrowd's ExploitBench benchmark shows Anthropic's Claude Mythos outperforms OpenAI's GPT5.5 in Chrome vulnerability exploitation (infosecurity-magazine.com)

· 53d ago · Report · Spotlight this ·
0xBASE INTEL BRIEF
  • Mythos avg score 9.90/16 vs GPT5.5 avg 5.51/16 on ExploitBench.
  • Mythos fully exploited 21/41 Chrome V8 vulnerabilities; GPT5.5 only 2.
  • Bugcrowd calls for AI-driven automated prioritization and patching in response.

"A new independent benchmark called ExploitBench, developed by Bugcrowd with Carnegie Mellon University and Chrome security researchers, evaluates AI models' ability to exploit vulnerabilities. In initial results presented at Infosecurity Europe 2026, Anthropic's Claude Mythos scored an average of 9.90 out of 16, fully exploiting 21 out of 41 vulnerabilities in the V8 engine. OpenAI's GPT5.5 averaged 5.51 and reached the highest tier on only two vulnerabilities. Bugcrowd's David Brumley cautioned against overgeneralization but noted the results show AI models closing the gap with elite human hackers. The benchmark measures staged exploitation outcomes up to arbitrary code execution."

Discussion Matrix

0 segments

no comments yet.