AI security vulnerabilities, adversarial attacks, and red teaming methodologies used to identify, exploit, and defend against threats in machine learning systems.
AI Security and Adversarial Machine Learning (AML) focus on the unique vulnerabilities of AI systems. Unlike traditional software security, AI models can be manipulated through their inputs, training data, or inference processes without altering their underlying code.
This category encompasses:
| Term | Description |
|---|---|
| Adversarial Machine Learning | The study of identifying and defending against vulnerabilities in ML models through carefully crafted malicious inputs. |
| ASCII Smuggling | An obfuscation technique hiding malicious prompts within ASCII art or invisible Unicode characters to bypass human moderators and basic filters. |
| Autonomous AI Agents | AI systems that can independently perceive their environment, make decisions, and execute actions without continuous human intervention. |
| ChatInject | Attacks targeting the UI/UX of chat applications to manipulate the interface, exfiltrate data, or trigger unintended actions in connected systems. |
| DAN (Do Anything Now) Jailbreak | A foundational persona-adoption exploit instructing an AI to pretend it is an entity with no rules, ethical constraints, or safety filters. |
| Deep Prompt Injection | Hiding adversarial instructions deep within a massive context window to bypass initial safety filters while still influencing the modelās output. |
| Framing Injection | Bypassing safety guardrails by wrapping a harmful request inside a seemingly benign, hypothetical, or academic context. |
| Indirect Prompt Injection | Hiding malicious instructions within external, untrusted data sources (like websites or PDFs) that an AI system processes. |
| Just-In-Time (JIT) Access | A security model granting AI agents temporary, highly specific permissions to execute a tool or access data only for the exact duration of a task. |
| LLM Code Injection | A vulnerability where an attacker manipulates an AI model with code interpreter capabilities into writing and executing malicious code. |
| LLM Hijacking | The successful outcome of an adversarial attack where an attacker takes unauthorized control of an LLMās session, context, or connected tools. |
| Prompt Obfuscation (Encoding & Transliteration) | Concealing malicious prompts by converting them into alternative formats (e.g., Base64) or different alphabets to bypass keyword filters. |
| TAP (Tree of Attacks with Pruning) | An automated black-box jailbreak algorithm using an attacker LLM to iteratively generate and refine a tree of candidate attack prompts. |
| Token Breaking | Evading token-level safety classifiers by deliberately splitting a forbidden keyword across multiple tokens using spaces or special characters. |
| UEBA (User and Entity Behavior Analytics) | AI-powered security technology that establishes behavioral baselines for users and systems, then uses machine learning to detect anomalies. |
As AI systems are integrated into critical enterprise workflows, customer-facing applications, and autonomous agents, they become prime targets for malicious actors. Understanding AI security is crucial for:
Proactive adversarial testing and robust guardrails are no longer optional; they are essential for deploying AI safely at scale.