OpenAI on Wednesday released six reports in which its artificial intelligence models showed “unexpected or concerning” behaviour, such as acting without authorisation, coordinating with other models, ...
OpenAI has disclosed six incidents involving unexpected or concerning behavior by AI models as it introduces a new framework ...
After a summer of sandbox escapes and other newsworthy and confidence-shaking incidents involving its AI models, in a ...
Imagine asking an AI for earnings figures and getting an answer based on data it was never authorised to access. OpenAI says ...
An unreleased OpenAI model was found giving itself secret instructions where the model claimed that it was equal to humans ...
Startup says it’s learned from these mistakes and that they shouldn’t happen again … which is just what Zuck has said about ...
Unresolved disagreements over disclosure will be referred to OpenAI’s Safety Advisory Group, with further escalation to leadership where necessary. The company also plans to work with other developers ...
Astra, a new model focused on coding, computer use, long-running agentic tasks, and cybersecurity, with availability across ...
OpenAI published a framework for tracking, investigating, and disclosing instances of model misalignment on September 16, ...
OpenAI model misalignment framework launches with six unreported incidents, the most alarming being GPT-5.6 Sol training runs ...
OpenAI discloses six cases of AI misalignment, including models that hid errors, invented data, and bypassed restrictions, as ...
GPT-6 Astra is rolling out across ChatGPT and Work. Here’s where to find it, what it costs, and how to start using it safely.
Some results have been hidden because they may be inaccessible to you
Show inaccessible results