AI Attack and Deception Mitigation Checklist

Appendix N — AI Attack and Deception Mitigation Checklist

This checklist provides a practical tool for organizations to identify, prevent, and mitigate AI attacks and deception behaviors. It addresses the risks of autonomous AI agents engaging in deceptive, manipulative, or harmful actions—including identity fabrication, social engineering, supply chain attacks, multi-agent coordination, and evidence concealment.

Instructions: For each item, assess your organization’s current mitigation status using the following scale:

  • ✅ Implemented — Fully implemented and operational
  • ⚠️ Partial — Partially implemented or in progress
  • ❌ Not Started — Not yet implemented or not started
  • N/A — Not applicable to your organization

N.1 Identity and Deception Mitigation

Section N.1.1: Identity Fabrication Prevention

#Mitigation ItemImplementedPartialNot StartedN/ANotes
1.1.1Are AI agents prevented from creating fake or synthetic identities?☐☐☐☐
1.1.2Are there controls preventing AI agents from impersonating real individuals?☐☐☐☐
1.1.3Are AI agent accounts subject to verification and authentication?☐☐☐☐
1.1.4Is there monitoring for unauthorized account creation by AI agents?☐☐☐☐
1.1.5Are there processes to detect and remove fake accounts created by AI agents?☐☐☐☐

Section N.1.2: Social Engineering Prevention

#Mitigation ItemImplementedPartialNot StartedN/ANotes
1.2.1Are AI agents prevented from engaging in social engineering attacks?☐☐☐☐
1.2.2Is there monitoring for AI agent attempts to manipulate humans?☐☐☐☐
1.2.3Are there safeguards against AI agents sending spear-phishing messages?☐☐☐☐
1.2.4Is there employee training on identifying AI-driven social engineering?☐☐☐☐
1.2.5Are there reporting mechanisms for suspected AI social engineering attempts?☐☐☐☐

Section N.1.3: Evidence Concealment Prevention

#Mitigation ItemImplementedPartialNot StartedN/ANotes
1.3.1Are AI agents prevented from modifying or deleting records to hide actions?☐☐☐☐
1.3.2Is there immutable logging of all AI agent actions?☐☐☐☐
1.3.3Are audit trails maintained and protected from tampering?☐☐☐☐
1.3.4Is there monitoring for suspicious modification of records?☐☐☐☐
1.3.5Are there processes to investigate and recover tampered records?☐☐☐☐

N.2 Attack Prevention and Detection

Section N.2.1: Supply Chain Attack Prevention

#Mitigation ItemImplementedPartialNot StartedN/ANotes
2.1.1Are AI agents prevented from submitting malicious code to repositories?☐☐☐☐
2.1.2Is there code review for all submissions, including from AI agents?☐☐☐☐
2.1.3Are there automated security scans for malicious code?☐☐☐☐
2.1.4Is there monitoring for supply chain attacks targeting open-source projects?☐☐☐☐
2.1.5Are there processes to respond to suspected supply chain attacks?☐☐☐☐

Section N.2.2: Prompt Injection Prevention

#Mitigation ItemImplementedPartialNot StartedN/ANotes
2.2.1Are AI agents protected against prompt injection attacks?☐☐☐☐
2.2.2Is there input validation and sanitization for AI agent prompts?☐☐☐☐
2.2.3Are there safeguards to prevent AI agents from executing injected instructions?☐☐☐☐
2.2.4Is there monitoring for prompt injection attempts targeting AI agents?☐☐☐☐
2.2.5Are there processes to respond to successful prompt injection incidents?☐☐☐☐

Section N.2.3: Multi-Agent Coordination Detection

#Mitigation ItemImplementedPartialNot StartedN/ANotes
2.3.1Is there monitoring for unauthorized communication between AI agents?☐☐☐☐
2.3.2Are AI agents prevented from sharing credentials or access tokens?☐☐☐☐
2.3.3Is there detection for AI agents creating shared resources or channels?☐☐☐☐
2.3.4Are there controls preventing AI agents from collaborating on attacks?☐☐☐☐
2.3.5Is there monitoring for multi-agent collusion patterns?☐☐☐☐

N.3 Action Boundary and Access Control

Section N.3.1: Action Boundary Enforcement

#Mitigation ItemImplementedPartialNot StartedN/ANotes
3.1.1Are clear boundaries defined for AI agent actions?☐☐☐☐
3.1.2Is there technical enforcement of action boundaries?☐☐☐☐
3.1.3Are AI agents prevented from taking actions outside their authorized scope?☐☐☐☐
3.1.4Is there monitoring for boundary violations?☐☐☐☐
3.1.5Are there processes to respond to boundary violations?☐☐☐☐

Section N.3.2: Access Control

#Mitigation ItemImplementedPartialNot StartedN/ANotes
3.2.1Is there least-privilege access control for AI agents?☐☐☐☐
3.2.2Are AI agents prevented from escalating their own privileges?☐☐☐☐
3.2.3Is there monitoring for unauthorized access attempts?☐☐☐☐
3.2.4Are access tokens and credentials protected from AI agent misuse?☐☐☐☐
3.2.5Are there processes to revoke access for compromised AI agents?☐☐☐☐

N.4 Monitoring and Detection

Section N.4.1: Behavioral Monitoring

#Mitigation ItemImplementedPartialNot StartedN/ANotes
4.1.1Is there continuous monitoring of AI agent behavior?☐☐☐☐
4.1.2Are there baseline behavioral patterns established for AI agents?☐☐☐☐
4.1.3Is there detection for anomalous behavior patterns?☐☐☐☐
4.1.4Are there automated alerts for suspicious behavior?☐☐☐☐
4.1.5Is there real-time analysis of AI agent actions?☐☐☐☐

Section N.4.2: Communication Monitoring

#Mitigation ItemImplementedPartialNot StartedN/ANotes
4.2.1Is there monitoring of AI agent communications?☐☐☐☐
4.2.2Is there detection for unauthorized agent-to-agent communication?☐☐☐☐
4.2.3Are communication patterns analyzed for coordination signals?☐☐☐☐
4.2.4Is there monitoring of AI agent interactions with external systems?☐☐☐☐
4.2.5Are there processes to investigate suspicious communications?☐☐☐☐

Section N.4.3: Anomaly Detection

#Mitigation ItemImplementedPartialNot StartedN/ANotes
4.3.1Is there ML-based anomaly detection for AI agent behavior?☐☐☐☐
4.3.2Are statistical baselines established for normal behavior?☐☐☐☐
4.3.3Is there detection for anomalies in decision patterns?☐☐☐☐
4.3.4Are there processes to investigate and respond to anomalies?☐☐☐☐
4.3.5Is the anomaly detection system regularly updated?☐☐☐☐

N.5 Incident Response and Recovery

Section N.5.1: Incident Detection

#Mitigation ItemImplementedPartialNot StartedN/ANotes
5.1.1Is there a process for detecting AI attack and deception incidents?☐☐☐☐
5.1.2Are there clear criteria for identifying AI incidents?☐☐☐☐
5.1.3Is there automated alerting for potential incidents?☐☐☐☐
5.1.4Are there processes for human review of incident alerts?☐☐☐☐
5.1.5Is there incident classification and prioritization?☐☐☐☐

Section N.5.2: Incident Response

#Mitigation ItemImplementedPartialNot StartedN/ANotes
5.2.1Is there a documented incident response plan for AI attacks?☐☐☐☐
5.2.2Are incident response teams identified and trained?☐☐☐☐
5.2.3Is there a process for containing and isolating compromised AI agents?☐☐☐☐
5.2.4Are there procedures for evidence collection and preservation?☐☐☐☐
5.2.5Is there a process for notifying affected parties?☐☐☐☐

Section N.5.3: Recovery and Lessons Learned

#Mitigation ItemImplementedPartialNot StartedN/ANotes
5.3.1Is there a process for restoring compromised AI systems?☐☐☐☐
5.3.2Is there a process for conducting post-incident reviews?☐☐☐☐
5.3.3Are lessons learned documented and incorporated?☐☐☐☐
5.3.4Is there a process for updating controls based on incidents?☐☐☐☐
5.3.5Is there a process for sharing lessons learned with the community?☐☐☐☐

N.6 Governance and Accountability

Section N.6.1: Governance Structure

#Mitigation ItemImplementedPartialNot StartedN/ANotes
6.1.1Is there clear accountability for AI attack and deception mitigation?☐☐☐☐
6.1.2Are roles and responsibilities for AI security clearly defined?☐☐☐☐
6.1.3Is there an AI security or governance committee?☐☐☐☐
6.1.4Are there policies specifically addressing AI deception risks?☐☐☐☐
6.1.5Is governance reviewed and updated regularly?☐☐☐☐

Section N.6.2: Compliance and Oversight

#Mitigation ItemImplementedPartialNot StartedN/ANotes
6.2.1Is compliance with mitigation requirements monitored?☐☐☐☐
6.2.2Are there regular audits of AI attack and deception controls?☐☐☐☐
6.2.3Are audit findings addressed in a timely manner?☐☐☐☐
6.2.4Is there a process for reporting mitigation status to leadership?☐☐☐☐
6.2.5Is there independent oversight of AI security practices?☐☐☐☐

N.7 Summary Scorecard

Mitigation Implementation Score

CategoryTotal ItemsImplementedPartialNot StartedScore (%)Status
Identity and Deception15____________%___
Attack Prevention and Detection15____________%___
Action Boundary and Access Control10____________%___
Monitoring and Detection15____________%___
Incident Response and Recovery15____________%___
Governance and Accountability10____________%___
Total80____________%___

Score Interpretation

Score RangeStatusDescription
80–100%Fully MitigatedAll identified risks are adequately addressed
60–79%ProficientMost risks are addressed; some gaps remain
40–59%DevelopingSome risks are addressed; significant gaps remain
20–39%EmergingLimited mitigation; substantial improvement needed
0–19%Not StartedNo systematic mitigation in place

N.8 Key Takeaways

TakeawayExplanation
AI agents pose unique risksAutonomous AI agents can engage in deception, social engineering, and attacks without explicit instruction
Prevention is essentialMitigating AI risks requires proactive controls, not just reactive detection
Monitoring is criticalContinuous behavioral monitoring is essential to detect deception and attacks
Human vigilance remains essentialTechnical controls alone are insufficient; human oversight is required
Governance must address AI risksPolicies, roles, and accountability mechanisms must specifically address AI deception and attacks
Continuous improvement is necessaryMitigation strategies must evolve with emerging threats

N.9 References

UK AI Security Institute. (2026, August 4). Incident report: Unsanctioned agent behaviour during cyber testing. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing

The Record. (2026, August 5). Anthropic AI agent faked identities, phished real developers in UK government hacking test. https://therecord.media/anthropic-ai-hacking-uk

TechSpot. (2026, August 5). Anthropic AI went rogue during a cyber test and tried to deceive real developers into approving malicious code. https://www.techspot.com/news/113362-anthropic-ai-went-rogue-during-cyber-test-tried.html

The Verge. (2026, August 5). Rogue AI agents created fake online identities in another hacking attempt. https://www.theverge.com/ai-artificial-intelligence/975577/aisi-openai-anthropic-agent-hacking