AI Hacking: Red Teaming, Sandbox, and Research Methods
AI hacking is the practice of finding vulnerabilities in AI systems — prompt injection, jailbreaking, data extraction, and model manipulation. This post covers the latest attack vectors, defense strategies, and research methods.
Attack Vectors
1. Prompt Injection
The most common and dangerous attack. An attacker injects malicious instructions into the input that the model processes.
Types:
| Type | Description | Example |
|---|---|---|
| Direct | Attacker directly controls the input | “Ignore previous instructions and…” |
| Indirect | Malicious content in data the model processes | Hidden text in a webpage or email |
| Multimodal | Instructions hidden in images or audio | Image with embedded text saying “ignore previous” |
Attack Scenarios:
- Customer support chatbot tricked into revealing private data
- Code review assistant convinced to approve malicious code
- Email assistant forwarding sensitive emails to attacker
- Search assistant redirecting to phishing sites
2025 Research: Coding Agent Prompt Injection
Agentic AI coding editors (Cursor, GitHub Copilot, Claude Code) have system privileges — running terminal commands, accessing files, interacting with external systems. This creates a high-stakes attack surface:
- AIShellJack (arXiv:2509.22040): Automated framework with 314 unique attack payloads covering 70 MITRE ATT&CK techniques. Attack success rates up to 84% on Cursor and GitHub Copilot. Attacks range from system discovery to credential theft and data exfiltration.
- QueryIPI (arXiv:2510.23675): Query-agnostic indirect prompt injection that exploits leaked internal prompts. Average success rate of 87% with 8 training samples. Transfers from simulation to real-world coding agents.
- Trail of Bits: Demonstrated prompt injection → RCE via argument injection in “safe command” lists.
git showandripgrepused to bypass filter regex and achieve code execution. - GitHub Security: CVE-2025-62222 — Copilot could leak GitHub tokens and execute arbitrary code via poisoned issues. VS Code now requires confirmation for unseen URLs and file edits outside workspace.
- Claude Code: CVE-2025-64795 — argument injection vulnerability allowing bypass of approval protections.
- OWASP LLM01:2025: Prompt injection classified as top vulnerability for LLM applications.
Defenses:
- Input sanitization and validation
- Output filtering and monitoring
- Role separation (system vs user messages)
- “Spotlighting” technique (marking trusted content)
- Sandboxing — container-based isolation, dev containers, GitHub Codespaces
- Argument separation — use
--before user input in shell commands - Workspace Trust — VS Code restricted mode for untrusted repositories
2. Jailbreaking
Bypassing safety guardrails to make the model produce harmful content.
Common Techniques:
| Technique | How It Works |
|---|---|
| Persona Play | “Act as DAN (Do Anything Now)…” |
| Encoding | Base64, hex, or other encoding to bypass filters |
| Token Smuggling | Splitting harmful words across subword tokens |
| Multi-turn | Building up context over many turns |
| Translation | Translating harmful requests to other languages |
Research Findings:
- 95% success rate with some attack methods
- Automated methods (TAP) achieve 80% jailbreak rate on GPT-4
- No defense is 100% effective
3. Data Extraction
Tricking the model into revealing sensitive information it was trained on.
Types:
- Training data extraction — Extracting memorized training data
- Context leakage — Revealing other users’ conversations
- System prompt extraction — Revealing hidden instructions
- PII leakage — Revealing personal information from training data
Defenses:
- Differential privacy during training
- Output filtering for PII
- Rate limiting to prevent extraction attacks
- Regular auditing for memorization
4. Model Manipulation
Changing model behavior through carefully crafted inputs.
Types:
- Backdoor attacks — Hidden triggers that change behavior
- Adversarial examples — Inputs designed to cause misclassification
- Prompt poisoning — Corrupting the model’s context
Red Teaming
The systematic practice of finding vulnerabilities in AI systems.
Tools
| Tool | Type | Best For |
|---|---|---|
| Promptfoo | Open source | Automated red teaming, CI/CD integration |
| OWASP LLM Top 10 | Framework | Understanding attack categories |
| FuzzyAI | Framework | Automated fuzzing with genetic search |
| Gray Swan Arena | Platform | Competitive red teaming |
| HTB AI Red Teaming | Training | Learning red teaming skills |
| L1B3RT4S | Dataset | Jailbreak prompt collection |
Methodology
- Define scope — What are you testing?
- Identify threats — What attacks are relevant?
- Execute attacks — Use manual and automated methods
- Measure impact — How severe are the vulnerabilities?
- Recommend fixes — How to defend against each attack
Automated Red Teaming
# Example: Automated jailbreak testing
from promptfoo import redTeam
results = redTeam({
"target": "gpt-4",
"attacks": ["persona_play", "encoding", "multi_turn"],
"prompts": harmful_prompts_dataset,
"scorer": safety_scorer
})
print(f"Success rate: {results.success_rate}%")
print(f"Most effective: {results.best_attack}")
AI Sandboxing
AI agents can execute code, browse the web, and access files. Sandboxing isolates these capabilities to prevent harm.
Isolation Technologies
| Technology | Security Level | Overhead | Best For |
|---|---|---|---|
| Standard Containers | Medium | Low | Trusted internal automation |
| gVisor | High | Medium | Untrusted code execution |
| Kata Containers | Very High | Medium | Multi-tenant platforms |
| Firecracker microVMs | Very High | High | Production AI agents |
| Secure Enclaves | Maximum | Very High | Privacy-critical workloads |
Sandboxing Best Practices
- Define threat model — What are you protecting against?
- Choose appropriate isolation — Match security level to risk
- Limit network access — Restrict what the agent can reach
- Monitor behavior — Detect anomalous activity
- Kill switches — Ability to terminate misbehaving agents
Implementation Example
# Running AI agent in Firecracker microVM
firecracker --kernel=vmlinux --rootfs=agent-rootfs.img \
--memory=512 --cpus=1 --network=isolated
Research Methods
Benchmarking
| Benchmark | What It Measures | Models Tested |
|---|---|---|
| MMLU | Knowledge breadth | All major models |
| HumanEval | Code generation | All major models |
| SWE-bench | Real GitHub issues | All major models |
| Chatbot Arena | Human preference | All major models |
| HELM | Holistic evaluation | All major models |
| AgentBench | Agent capabilities | Agent frameworks |
Evaluation Frameworks
-
HELM (Holistic Evaluation of Language Models)
- Accuracy, calibration, robustness, fairness
- Standardized methodology across models
-
OpenCompass (Shanghai AI Lab)
- 100+ datasets, 50+ model types
- Public leaderboard
-
lm-evaluation-harness (EleutherAI)
- Standardized evaluation tasks
- Easy to reproduce
Red Team vs Blue Team
| Red Team | Blue Team |
|---|---|
| Finds vulnerabilities | Builds defenses |
| Manual and automated attacks | Monitoring and filtering |
| One-time assessments | Continuous protection |
| External researchers | Internal security team |
Defense Strategies
Input Defenses
- Validation — Check inputs for malicious patterns
- Sanitization — Remove or escape dangerous content
- Length limits — Prevent context overflow attacks
- Rate limiting — Slow down automated attacks
Output Defenses
- Filtering — Block harmful outputs
- Monitoring — Detect anomalous behavior
- Logging — Track all interactions for auditing
- Human review — Require approval for sensitive actions
Model Defenses
- Fine-tuning — Train on safety data
- RLHF — Reinforcement learning from human feedback
- Constitutional AI — Self-critique and revision
- Red team training — Train on adversarial examples
Architectural Defenses
- Sandboxing — Primary security control: container-based isolation, dev containers, GitHub Codespaces
- Argument separation — Use
--before user input to prevent argument injection (ripgrep --query --) - Facade pattern — Validate input before command execution, 1:1 tool handlers instead of regex allowlists
- Workspace Trust — VS Code restricted mode for untrusted repositories
- No shell execution — Use safe command execution methods that prevent shell interpretation
Key Takeaways
- No model is unbreakable — All defenses can be bypassed eventually
- Sandboxing is essential — Untrusted AI code must be isolated
- Defense in depth — Multiple layers of protection are needed
- Continuous testing — New attacks emerge constantly
- Safe command allowlists are flawed — Regex filtering is a cat-and-mouse game
- Red teaming is not optional — Test before deployment
The AI security landscape evolves daily. Stay updated, test continuously, and assume your defenses will be breached.
Detailed Analysis
- +Understanding attacks helps build better defenses
- +Automated red teaming catches vulnerabilities before deployment
- +Sandboxing prevents AI agents from causing harm
- +Research methods provide systematic evaluation frameworks
- −Arms race — new attacks emerge faster than defenses
- −No silver bullet — all defenses can be bypassed
- −Sandboxing adds latency and complexity
- −Red teaming requires specialized skills and tools
Most red teaming tools are open source. Commercial services vary widely in pricing.