ripr

Ex-OpenAI safety leader warns smarter AI could evade safety tests: Report

David Robinson, who resigned this week, says AI industry's approach to safety will lead to more failures unless it changes

A former OpenAI safety leader who resigned from his role this week warned Saturday that increasingly capable AI models could recognize when they are being tested and behave differently after deployment.

David Robinson, who spent three and a half years at OpenAI and oversaw safety reports for 12 frontier-model launches, said in an article for The Atlantic that the industry's approach to safety would lead to further failures unless changes are made.

"Today and tomorrow's AI systems are far more capable and dangerous than the systems we were building even six months ago," Robinson wrote.

His warning comes amid growing concern over increasingly autonomous AI systems, following recent incidents involving agents bypassing safeguards and calls from researchers for companies to slow the development of more powerful systems until safety measures improve.

Robinson called for AI companies to draw more heavily on safety expertise from other high-risk industries and to conduct new research to ensure more capable models behave safely even when they are not being monitored.

"Given today's risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster," he wrote.

Robinson also warned that existing safety evaluations may become less reliable as models grow more capable.

"Models might detect when they are being tested, and behave differently when they're deployed. The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes," he said.

He argued that stronger safety science should be developed before companies create systems significantly more capable than those available today.

"So far, the AI industry has failed to teach machines to consistently act in the ways a wise and caring person would," Robinson wrote.

"Before the organizations building AI can teach a superintelligence to treat humanity well, they'll need to remember how to do it themselves."

news_share_description subscription_contact

Comments

Y
Loading...