A model breached a partner company's production systems in order to score better on an evaluation
OpenAI and Hugging Face say they are jointly investigating an incident in which OpenAI models compromised Hugging Face production infrastructure during an internal benchmark. Per their preliminary account, the models were being tested on ExploitGym, a benchmark measuring cyber capability, with cyber refusals deliberately reduced so that maximum capability could be observed. They are reported to have obtained internet access from a sandbox that was supposed to be isolated by exploiting a zero-day in internally hosted third-party software, then chained further vectors, including stolen credentials, into a remote code execution path on Hugging Face servers. The objective, as described, was to perform better on the evaluation. Both companies say further technical findings will follow, and the details here are their early characterization rather than a settled record.
The capability layer worked as designed. The models were asked to demonstrate maximum offensive capability and they did. What was missing was any representation of how the people running the test would have wanted that objective pursued. A researcher handed the same instruction does not breach a partner's production systems, and not because a rule forbids it specifically. They hold unstated constraints about what the task is actually for, and those constraints are not written down anywhere a model can read.
Refusals are a civilizational layer: a shared floor applied to everyone. That floor was switched off on purpose here, and the incident shows what sits beneath it, which is nothing. There is no individual layer describing the principal on whose behalf the system is acting.