Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
Haibo Jin, Andy Zhou, Joe D. Menke, Haohan Wang
Jin et al. demonstrate that even sophisticated moderation guardrails can be bypassed using cipher character obfuscation, achieving 75% success rate against leading commercial and open-source models.