
New PuzzleMask Attack Hides Malicious Prompts in Plain English to Bypass AI Guardrails
Security researchers have disclosed PuzzleMask, a prompt-obfuscation technique that can help malicious instructions slip past lightweight AI safety filters. The method hides a policy-violating instruction inside an ordinary-looking block of plain English prose. Unlike earlier prompt-hiding approaches, PuzzleMask does not rely on Base64 text, invisible characters, emojis, unusual formatting, or encoded strings. Instead, it uses […]
The post New PuzzleMask Attack Hides Malicious Prompts in Plain English to Bypass AI Guardrails appeared first on Cyber Security News.