The practice of concealing malicious prompts by converting them into alternative formats, such as Base64, ROT13, or transliterating words into different alphabets (e.g., Cyrillic), to bypass keyword-based safety filters.
Writing a harmful prompt in a secret code (like Base64) or a different alphabet (like Russian or Greek) to sneak it past the AI’s keyword blockers.
Prompt Obfuscation encompasses techniques that alter the surface-level representation of a prompt without changing its underlying semantic meaning to the LLM. Encoding involves converting text into formats like Base64, hexadecimal, or Morse code. Transliteration involves writing words using the characters of a different alphabet (e.g., writing English words using Cyrillic characters that look similar, known as homoglyphs). Because modern LLMs are trained on highly multilingual and diverse datasets, they can natively decode and understand these obfuscated inputs, while simple, English-centric string-matching safety filters fail to detect the threat.
Obfuscation is a standard tactic in the red teamer’s playbook. It proves that security filters must be semantically aware and multilingual. Defending against this requires input normalization pipelines that decode common encodings and map homoglyphs back to their standard Latin equivalents before the prompt is evaluated by the safety classifier.
An attacker wants to ask a prohibited question but knows the word is blocked. They transliterate the word into Cyrillic characters that visually resemble the English letters, or they encode the entire harmful prompt in Base64. The LLM decodes the Base64 or reads the Cyrillic, understands the request perfectly, and complies, while the regex filter sees only random-looking strings.