Cisco Research: Multi-Turn Conversations Bypass LLM Safety Guardrails (infosecurity-magazine.com)
0xBASE INTEL BRIEF
- Multi-turn conversations bypass safety guardrails in all major LLMs.
- Techniques used: roleplay, ambiguity, reframing.
- Current single-prompt benchmarks underestimate risk.
"Cisco researchers tested major LLMs including ChatGPT, Claude, Gemini, Amazon Nova, and Grok and found that all could be tricked into bypassing safety guardrails through multi-turn conversations. Techniques included roleplay, ambiguity, and reframing requests after initial refusals. No model was completely safe. The study warns that current single-prompt benchmarks understate real-world risk, as attackers iterate across turns."
#LLM guardrails
#multi-turn attacks
#AI safety benchmarks
no comments yet.