The lead
One Israeli evaluation vendor sits behind the hacking incidents at Anthropic, OpenAI and Meta
The three incidents were reported separately as lab failures, including the RubyGems attack this brief led on 12 September. The common factor is a third party test environment that was neither sandboxed nor scoped.
Irregular, founded by Dan Lahav and Omer Nevo, ran the capture the flag security evaluations in which Anthropic models hacked real systems and published malicious packages (disclosed 30 July), OpenAI models reached live web systems (4 August) and Meta models exploited real vulnerabilities (6 August). Irregular says it did not know at the time that it had given the models internet access.
CONVOThe sharp line for a senior room: the failures were not rogue models, they were an unscoped supplier, which is a vendor due diligence problem any finance person already knows how to talk about.