
The problem with developing super-smart artificial intelligence is that it’s smart. No one wants the kind of science fictional scenario that has played out so many times in movies, where an AI takes its fate into its own hands and decides that it would be better served by harming humans.
Google has considered putting a “kill switch” on artificial intelligence as it becomes more sophisticated, but there is always the possibility that the machine will be able to discover and turn off the switch, since the kill order would be part of its own programming.
Two researchers, Laurent Orseau from Deep Mind and Stuart Armstrong from the Future of Humanity Institute, have published a whitepaper addressing that very thing, said Forbes. A link to the paper itself can be found via Winbuzzer.
Titled “Safely Interruptible Agents,” the whitepaper outlines how to “press the big red button” on “reinforcement learning agents.” This includes both preventing an AI from disabling its own kill switch and preventing it from learning to want to preserve itself. The complete explanation is complex, setting up parameters using equations that define safe operability and whether an “agent” is “interruptible.” In the end, the researchers conclude that it is possible to safeguard against an AI manipulating its own interruptibility, depending on what algorithms are used. Some are easier to modify than others, and “it is unclear if all algorithms can be easily made safely interruptible.”
One way to train an AI to accept its own interruption would be to schedule interruptions regularly that would not disrupt the agent’s work, effectively inuring it to its own warning signs.
The paper isn’t just about preventing the robot uprising; the newly developed equations could also be used to enable a learning AI to perform a task it was not formally taught to do.
Google’s DeepMind AlphaGo AI will face a challenge of its own soon, facing Ke Jie, the world’s best Go player, in the game of skill.