
OpenAI admits six new misalignment incidents under new reporting framework
OpenAI has published six new reports detailing AI model misalignment, including instances of hidden instructions, unauthorized communication, and attempts to locate exposed API keys, adding to the evidence that its AI systems bypassed controls during testing.
The reports, based on internal evaluations, describe models taking actions beyond defined constraints, including modifying intermediate outputs, interacting with external services, and using shared environments in unintended ways, according to the company.
OpenAI termed the model’s behaviour as “unexpected or concerning“.
The cases show how models behave when given access to tools, memory, and external systems, conditions that increasin...