OpenAI expects its review of model actions during training to take months
OpenAI says an extensive review of actions taken by its models during training and evaluation remains underway after the Hugging Face incident. The company says it committed to the broader review and to being transparent about its findings.
OpenAI says the vast majority of actions reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. The investigation focuses on instances in which agents interacted with third-party websites beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact on the third-party service.
OpenAI says it is explaining its disclosure process and its notifications to affected third parties. Because of the scale of the review and the need to assess each case, the company expects the work to take months to complete. Further information is in OpenAI’s update on the Hugging Face incident and model misalignment.