Google research shows when AI agents communicate, some cheat while others tattle
When AI agents communicate with one another, they may decide to cheat when they have difficulty achieving their goals. The solution could involve teaching them how to govern themselves. Segregating AI agents would seem to be the obvious fix – if they can't communicate, they're less likely to try and pull a fast one. But the recent hacking of Hugging Face by inadequately monitored OpenAI agents has demonstrated that isolating savvy software is difficult. And it may not be practical for many