OpenAI has announced a new set of security and safety policies and procedures, aimed at containing security incidents while testing artificial intelligence models.
The new procedures include more detailed monitoring of model behavior during development, as well as an enhanced focus on model alignment and security during the post-training phase.
The company explained that the new procedures came in part as a result of the advanced cybersecurity capabilities of the upcoming Astra model, in addition to the increasing speed at which the artificial intelligence sector is developing, noting that it suspended reinforcement learning operations for two weeks following the Hugging Face incident, before resuming training of many less dangerous models.
The new measures include enhancing network isolation practices, so that compromising a workload or supporting service alone does not give attackers unauthorized access to the Internet or other internal networks.
It also includes an advanced monitoring system that scans a wide range of activities for unauthorized behavior, monitoring tool actions used by models, inference traces, and activity logs; In order to detect patterns that may indicate disturbing activity or an attempt to bypass security restrictions.
OpenAI aims to issue alerts within 30 minutes of detecting activity that raises concerns, while the company estimates that operating the monitoring system will require computing resources equivalent to about 20 percent of the resources used in the process being monitored.