
New attack lets hackers plant hidden instructions in AI memory with a single prompt
A newly demonstrated attack technique allows hackers to plant hidden instructions inside an AI agent’s memory with a single prompt, enabling them to influence how the system responds to future queries.
The technique, called InjecMEM, is described in a research paper as a “targeted red-teaming attack paradigm on agent memory systems with just one interaction and no read/edit access to the memory store.”
“The attacker specifies a target topic and target output, aiming to make the agent generate that output for later queries on the topic,” researchers from Shanghai Jiao Tong University and Ant Group wrote in a paper.
The researchers said the method targets AI agents that store past interactions...