Add one more AI worry to the nightmare scenario: self-replicating prompt injections
Imagine a prompt injection that keeps replicating itself like a worm. It's not just the stuff of bad dreams. “We have found instances of our GPT models being susceptible to an AI-version of a worm attack that we call ‘self-replicating prompt injection,’” OpenAI said in a Friday alignment research blog. There’s no indication that these indirect prompt-injection attacks occurred in any real-life security incident, or anywhere outside of the models’ training envi