AI Hacking
Overview
AI hacking is broadly defined along two directions. The first is attacks that target artificial intelligence models themselves, inducing malfunction, information leakage, and the bypass of safety mechanisms (jailbreaking); the second is the abuse of generative AI as an attack tool to raise the scale and sophistication of existing cyberattacks. Since the arrival of ChatGPT in 2022, as large language models (LLMs) have penetrated every industry, the security industry has come to treat AI at once as an "asset to be protected" and as a "new attack weapon."
Key Details
1. Attacks Targeting AI
Prompt Injection
This is the most representative LLM vulnerability. Malicious instructions are hidden inside user input, external documents, or web pages, causing the model to disregard its original system instructions. Direct injection is the method in which the user injects the command themselves, while indirect injection is the method of planting commands in web pages, emails, or PDFs that the model reads. Since 2023 it has ranked first in OWASP's Top 10 for LLM Applications, and in architectures where agentic AI calls tools, it can lead to the actual hijacking of system privileges, sharply raising the level of risk.
Jailbreak
Techniques such as role-playing, hypothetical scenarios, multilingual circumvention, and encoding transformations neutralize a model's safety alignment to make it output prohibited information. The "DAN" prompt was the early representative case, and since then research on automated search for attack prompts has emerged.
Adversarial Examples
Subtle perturbations imperceptible to humans are added to images, audio, or text to mislead a model's judgment. These can be used for autonomous-driving recognition errors, face-recognition bypass, malware-detection evasion, and similar purposes.
Data Poisoning and Backdoors
Malicious samples are mixed into training or fine-tuning data so that the model produces intentionally wrong answers for a specific trigger input. As the ecosystem for distributing open-source models has grown, this has emerged as a supply-chain attack path.
Model Stealing and Membership Inference
Performing large volumes of API queries to replicate a model's behavior, or determining whether specific personal information was included in the training data, thereby infringing on privacy.
2. Attacks Using AI
- Deepfake and voice-cloning fraud: A technique that forges an executive's voice or video call to induce a money transfer. In 2024 in Hong Kong, a scam of roughly 25 million USD was carried out using a fabricated video conference.
- Automated phishing: Generative AI is used to mass-produce tailored spear-phishing emails free of grammatical errors.
- Accelerated vulnerability discovery: Code-generation models quickly produce fuzzers and draft exploits, shortening the vulnerability discovery cycle.
- Malware mutation: Obfuscation and polymorphic code generation are used to evade antivirus detection.
- Abuse of AI agents: The file, payment, and mail permissions granted to autonomous agents are hijacked so that the agents carry out attacks automatically.
3. Defensive Techniques
Representative measures include input/output filtering and guardrails, system prompt isolation, the principle of least privilege, tool-call approval procedures, alignment techniques such as RLHF and constitutional AI, differential privacy and federated learning, watermarking and provenance labeling (C2PA), and continuous red-team evaluation. Recently, an "AI versus AI" configuration that deploys AI-based detection models alongside these has become commonplace.
Latest Trends
The key change of 2024–2025 is that the target of attack has shifted from "chatbots" to "agents." As LLMs directly call external tools such as file systems, browsers, payment APIs, and MCP (Model Context Protocol) servers, incidents have been reported in which a single prompt injection led to actual data leakage or money transfers. In 2025, major AI companies successively issued guidelines restricting autonomous agents' tool permissions and mandating Human-in-the-Loop approval.
On the regulatory side, the EU AI Act took effect in 2024 and entered phased implementation in 2025–2026, while the United States is pursuing self-regulation alongside federal procurement requirements centered on the NIST AI Risk Management Framework. In South Korea, the AI Framework Act (Framework Act on the Development of Artificial Intelligence and the Creation of a Trust Foundation) passed the National Assembly in 2025 and is set to take effect in 2026, with the obligation to label generative AI outputs and measures to ensure the safety of high-impact AI as the core issues. In addition, national CERTs and security companies are expanding LLM vulnerability reporting systems and AI-specific bug bounties.
Technically, multimodal attacks (instructions hidden in images, audio injection) and long-term memory poisoning attacks have emerged as new research topics, and on the defensive side, agent isolation architectures combining formal verification and sandboxing are drawing attention.
Related Topics
- [[Prompt Injection]]
- [[Adversarial Attack]]
- [[Generative AI]]
- [[Cybersecurity]]
- [[Deepfake]]
- [[AI Regulation]]