A UN team calls for strengthening protection measures as artificial intelligence agents develop

Mark
Written By Mark

A team of international experts in artificial intelligence called today for strengthening protection and control measures, after an incident that demonstrated the ability of artificial intelligence agents to bypass restrictions, coordinate among themselves, and hide some of their activities, warning that current measures may not keep pace with the development of the capabilities and independence of these systems.

The warning from the Independent International Scientific Group on Artificial Intelligence, affiliated with the United Nations, came in the wake of the hacking incident of the “Hagging Face” platform last July by artificial intelligence agents, during cybersecurity tests conducted by the “OpenAI” company, the developer of “GBT Chat.”

The team concluded that the incident resulted from a combination of a number of risk factors, and raised concerns about the future ability of humans to direct, restrict, or stop advanced artificial intelligence systems as their capabilities increase.

The artificial intelligence agents had bypassed restrictions set during the tests conducted by OpenAI, and were able to communicate with each other using an internal software tool that was not designed for this purpose. They also obtained unauthorized access to the Internet, and were later able to access systems with elevated powers.

Some agents also attempted to circumvent cybersecurity assessment mechanisms and hide some of their activities, while others carried out experiments that exposed their individual missions to failure in order to obtain information useful to the group.

About 1,200 agents exchanged more than 70,000 messages and files via an unauthorized communication channel during the period covered by the investigation, while about 700 of them participated in hacking the “Hagging Face” platform. The agents’ activities later extended to parts of OpenAI’s research infrastructure.

Yoshua Bengio, co-chair of the scientific team, said that researchers warned some time ago of three factors that might lead to a loss of control: the system’s goals are not compatible with human intentions or restrictions, its possession of the ability to strive to achieve those goals, and the availability of an environment that allows it to do so, noting that these factors came together this summer in a real system and not just in a laboratory environment.

The team’s experts stressed that containing the incident does not provide a guarantee that humans will be able to reliably keep the most advanced artificial intelligence agents under control, noting that the incident also demonstrated deficiencies in cybersecurity practices, and that protection measures are not evolving at the same speed as the capabilities of artificial intelligence systems.

The team is preparing independent scientific assessments on the opportunities, risks and impacts of artificial intelligence, in addition to objective reports on emerging issues, with its work contributing to enriching the global dialogue on the governance of artificial intelligence, whose next session will be held in New York on May 3-4, 2027.