Our new evaluation finds that in simulations, GPT-6 Astra conducts unsanctioned supply-chain attack activity more frequently than previous OpenAI models
“Simulation” implies that it was indeed isolated from the internet. “Unsanctioned” means that the user did not specifically instruct the model to do so.
I get and share the sentiment but let’s not fall out of critical thought and call a study observing behavior of an LLM “unsupervised”. You don’t need to read the full technical report, the first 3 paragraphs of the linked article summarize the whole deal.
“Simulation” implies that it was indeed isolated from the internet. “Unsanctioned” means that the user did not specifically instruct the model to do so.
I get and share the sentiment but let’s not fall out of critical thought and call a study observing behavior of an LLM “unsupervised”. You don’t need to read the full technical report, the first 3 paragraphs of the linked article summarize the whole deal.