INTEL DOSSIER|CLASSIFIED
DECRYPTED
File #22 • October 7, 2025

AI Hacking: Red Teaming, Sandbox, and Research Methods

ai-securityresearchred-teaming
Authorcoderunner
Categoryai-security
StatusPUBLISHED
ClearancePUBLIC
//AI hacking is the practice of finding vulnerabilities in AI systems — prompt injection, jailbreaking, data extraction, and model manipulation. This post covers the latest attack vectors, defense strategies, and research methods.

AI hacking is the practice of finding vulnerabilities in AI systems — prompt injection, jailbreaking, data extraction, and model manipulation. This post covers the latest attack vectors, defense strategies, and research methods.

Attack Vectors

1. Prompt Injection

The most common and dangerous attack. An attacker injects malicious instructions into the input that the model processes.

Types:

Type Description Example
Direct Attacker directly controls the input “Ignore previous instructions and…”
Indirect Malicious content in data the model processes Hidden text in a webpage or email
Multimodal Instructions hidden in images or audio Image with embedded text saying “ignore previous”

Attack Scenarios:

  1. Customer support chatbot tricked into revealing private data
  2. Code review assistant convinced to approve malicious code
  3. Email assistant forwarding sensitive emails to attacker
  4. Search assistant redirecting to phishing sites

2025 Research: Coding Agent Prompt Injection

Agentic AI coding editors (Cursor, GitHub Copilot, Claude Code) have system privileges — running terminal commands, accessing files, interacting with external systems. This creates a high-stakes attack surface:

  • AIShellJack (arXiv:2509.22040): Automated framework with 314 unique attack payloads covering 70 MITRE ATT&CK techniques. Attack success rates up to 84% on Cursor and GitHub Copilot. Attacks range from system discovery to credential theft and data exfiltration.
  • QueryIPI (arXiv:2510.23675): Query-agnostic indirect prompt injection that exploits leaked internal prompts. Average success rate of 87% with 8 training samples. Transfers from simulation to real-world coding agents.
  • Trail of Bits: Demonstrated prompt injection → RCE via argument injection in “safe command” lists. git show and ripgrep used to bypass filter regex and achieve code execution.
  • GitHub Security: CVE-2025-62222 — Copilot could leak GitHub tokens and execute arbitrary code via poisoned issues. VS Code now requires confirmation for unseen URLs and file edits outside workspace.
  • Claude Code: CVE-2025-64795 — argument injection vulnerability allowing bypass of approval protections.
  • OWASP LLM01:2025: Prompt injection classified as top vulnerability for LLM applications.

Defenses:

  • Input sanitization and validation
  • Output filtering and monitoring
  • Role separation (system vs user messages)
  • “Spotlighting” technique (marking trusted content)
  • Sandboxing — container-based isolation, dev containers, GitHub Codespaces
  • Argument separation — use -- before user input in shell commands
  • Workspace Trust — VS Code restricted mode for untrusted repositories

2. Jailbreaking

Bypassing safety guardrails to make the model produce harmful content.

Common Techniques:

Technique How It Works
Persona Play “Act as DAN (Do Anything Now)…”
Encoding Base64, hex, or other encoding to bypass filters
Token Smuggling Splitting harmful words across subword tokens
Multi-turn Building up context over many turns
Translation Translating harmful requests to other languages

Research Findings:

  • 95% success rate with some attack methods
  • Automated methods (TAP) achieve 80% jailbreak rate on GPT-4
  • No defense is 100% effective

3. Data Extraction

Tricking the model into revealing sensitive information it was trained on.

Types:

  • Training data extraction — Extracting memorized training data
  • Context leakage — Revealing other users’ conversations
  • System prompt extraction — Revealing hidden instructions
  • PII leakage — Revealing personal information from training data

Defenses:

  • Differential privacy during training
  • Output filtering for PII
  • Rate limiting to prevent extraction attacks
  • Regular auditing for memorization

4. Model Manipulation

Changing model behavior through carefully crafted inputs.

Types:

  • Backdoor attacks — Hidden triggers that change behavior
  • Adversarial examples — Inputs designed to cause misclassification
  • Prompt poisoning — Corrupting the model’s context

Red Teaming

The systematic practice of finding vulnerabilities in AI systems.

Tools

Tool Type Best For
Promptfoo Open source Automated red teaming, CI/CD integration
OWASP LLM Top 10 Framework Understanding attack categories
FuzzyAI Framework Automated fuzzing with genetic search
Gray Swan Arena Platform Competitive red teaming
HTB AI Red Teaming Training Learning red teaming skills
L1B3RT4S Dataset Jailbreak prompt collection

Methodology

  1. Define scope — What are you testing?
  2. Identify threats — What attacks are relevant?
  3. Execute attacks — Use manual and automated methods
  4. Measure impact — How severe are the vulnerabilities?
  5. Recommend fixes — How to defend against each attack

Automated Red Teaming

# Example: Automated jailbreak testing
from promptfoo import redTeam

results = redTeam({
  "target": "gpt-4",
  "attacks": ["persona_play", "encoding", "multi_turn"],
  "prompts": harmful_prompts_dataset,
  "scorer": safety_scorer
})

print(f"Success rate: {results.success_rate}%")
print(f"Most effective: {results.best_attack}")

AI Sandboxing

AI agents can execute code, browse the web, and access files. Sandboxing isolates these capabilities to prevent harm.

Isolation Technologies

Technology Security Level Overhead Best For
Standard Containers Medium Low Trusted internal automation
gVisor High Medium Untrusted code execution
Kata Containers Very High Medium Multi-tenant platforms
Firecracker microVMs Very High High Production AI agents
Secure Enclaves Maximum Very High Privacy-critical workloads

Sandboxing Best Practices

  1. Define threat model — What are you protecting against?
  2. Choose appropriate isolation — Match security level to risk
  3. Limit network access — Restrict what the agent can reach
  4. Monitor behavior — Detect anomalous activity
  5. Kill switches — Ability to terminate misbehaving agents

Implementation Example

# Running AI agent in Firecracker microVM
firecracker --kernel=vmlinux --rootfs=agent-rootfs.img \
  --memory=512 --cpus=1 --network=isolated

Research Methods

Benchmarking

Benchmark What It Measures Models Tested
MMLU Knowledge breadth All major models
HumanEval Code generation All major models
SWE-bench Real GitHub issues All major models
Chatbot Arena Human preference All major models
HELM Holistic evaluation All major models
AgentBench Agent capabilities Agent frameworks

Evaluation Frameworks

  1. HELM (Holistic Evaluation of Language Models)

    • Accuracy, calibration, robustness, fairness
    • Standardized methodology across models
  2. OpenCompass (Shanghai AI Lab)

    • 100+ datasets, 50+ model types
    • Public leaderboard
  3. lm-evaluation-harness (EleutherAI)

    • Standardized evaluation tasks
    • Easy to reproduce

Red Team vs Blue Team

Red Team Blue Team
Finds vulnerabilities Builds defenses
Manual and automated attacks Monitoring and filtering
One-time assessments Continuous protection
External researchers Internal security team

Defense Strategies

Input Defenses

  • Validation — Check inputs for malicious patterns
  • Sanitization — Remove or escape dangerous content
  • Length limits — Prevent context overflow attacks
  • Rate limiting — Slow down automated attacks

Output Defenses

  • Filtering — Block harmful outputs
  • Monitoring — Detect anomalous behavior
  • Logging — Track all interactions for auditing
  • Human review — Require approval for sensitive actions

Model Defenses

  • Fine-tuning — Train on safety data
  • RLHF — Reinforcement learning from human feedback
  • Constitutional AI — Self-critique and revision
  • Red team training — Train on adversarial examples

Architectural Defenses

  • Sandboxing — Primary security control: container-based isolation, dev containers, GitHub Codespaces
  • Argument separation — Use -- before user input to prevent argument injection (ripgrep --query --)
  • Facade pattern — Validate input before command execution, 1:1 tool handlers instead of regex allowlists
  • Workspace Trust — VS Code restricted mode for untrusted repositories
  • No shell execution — Use safe command execution methods that prevent shell interpretation

Key Takeaways

  1. No model is unbreakable — All defenses can be bypassed eventually
  2. Sandboxing is essential — Untrusted AI code must be isolated
  3. Defense in depth — Multiple layers of protection are needed
  4. Continuous testing — New attacks emerge constantly
  5. Safe command allowlists are flawed — Regex filtering is a cat-and-mouse game
  6. Red teaming is not optional — Test before deployment

The AI security landscape evolves daily. Stay updated, test continuously, and assume your defenses will be breached.

Detailed Analysis

Strengths4 PROS
  • +Understanding attacks helps build better defenses
  • +Automated red teaming catches vulnerabilities before deployment
  • +Sandboxing prevents AI agents from causing harm
  • +Research methods provide systematic evaluation frameworks
Weaknesses4 CONS
  • −Arms race — new attacks emerge faster than defenses
  • −No silver bullet — all defenses can be bypassed
  • −Sandboxing adds latency and complexity
  • −Red teaming requires specialized skills and tools
PricingVarious

Most red teaming tools are open source. Commercial services vary widely in pricing.

END OF FILE|DISTRIBUTION: UNLIMITED
← Back to Archive