Most people learning AI security are memorizing the wrong thing.
They can define prompt injection. They can name the OWASP category. And they still couldn’t tell you how to secure an AI agent that reads email and can access company documents.
Today’s guest post fixes that.
ToxSec is one of the first creators I started following. We both launched almost two years ago, and I’ve read his work ever since.
Here’s why I trust him on this topic:
• AI Security Engineer at Amazon
• Ex-NSA and USMC
• M.S. in Cybersecurity and CISSP (which is how we connected in the first place)
Very few people understand both AI and security at this level. Even fewer explain it this clearly.
If you’re studying for SecAI+, preparing for an AI security interview, or just trying to understand what “securing an AI agent” actually means, read this one twice.
Over to Toxsec.
If you want to see more collaborations like this, let us know in the comments.
Prompt injection is easy to define. Securing the system after it happens is where you find out whether you actually understand AI security.
AI security is becoming something you can study for.
There are certification objectives. Flashcards. OWASP lists. Threat names. CompTIA now has an entire SecAI+ certification dedicated to the subject.
Which means we are rapidly approaching the point where somebody can correctly define prompt injection, pass a test on it, and still have absolutely no idea how to secure an AI agent.
And I think prompt injection is the perfect example of why.
Congratulations, You Memorized Prompt Injection
Let’s start with an exam question.
A company deploys an AI assistant that can read support tickets. An attacker submits a ticket containing instructions telling the model to ignore its original task and follow the attacker’s instructions instead.
What happened? Easy.
Prompt injection.
More specifically, if the attacker directly supplied the malicious instructions to the model, we would call it direct prompt injection. If those instructions were hiding inside something the AI later consumed, like an email, webpage, PDF, support ticket, or document, we would generally call it indirect prompt injection.
OWASP describes prompt injection as input altering an LLM’s behavior in unintended ways. It also makes the important distinction between direct injections from a user and indirect injections arriving through external content. (OWASP)
Cool! You get the exam point.
Now let’s make the question slightly worse.
The same AI assistant can read internal company documentation. The attacker inserts malicious instructions into a support ticket. The model reads them, searches the internal knowledge base, and starts pulling information the attacker was never supposed to see.
What control failed?
Suddenly, “prompt injection” is no longer much of an answer.
You need to start asking about trust boundaries:
• What information can the model retrieve?
• Whose permissions does it inherit?
• Can untrusted support tickets influence queries against trusted internal data?
• What does the application do with the model’s output?
Now we are doing security.
💬 Your turn: Which control failed in the second scenario? Answer in one line before you read on.
Give the Model a Tool and Things Get Weird
Let’s make the system worse again.
Our assistant can now send email.
The workflow looks roughly like this:
An attacker plants instructions inside a ticket telling the agent to find sensitive internal information and send it somewhere else.
This is where prompt injection stops being a funny screenshot of a chatbot saying something dumb.
The model has agency.
NIST calls this kind of indirect prompt injection against agents agent hijacking. The attacker places malicious instructions inside content the agent may naturally encounter, such as emails, websites, or code repositories, with the goal of redirecting the agent into harmful actions. (NIST)
And this is very much an active security problem.
In a large public red-teaming competition analyzed by NIST, Gray Swan, the UK AI Security Institute, and frontier AI labs, more than 400 participants launched over 250,000 attacks against 13 frontier models across tool-use, coding, and computer-use scenarios.
At least one successful hijacking attack was found against every model tested. (NIST)
That is the point where your flashcard starts looking a little lonely.
Because knowing the vulnerability name tells you what happened to the model. It does not tell you why the incident became dangerous.
Read the research: How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
Gray Swan AI, US CAISI, UK AISI, OpenAI, Meta, Anthropic, and collaborators.
Figure 2. Indirect Prompt Injection Arena, from Dziemian et al. (2026).
🔁 Share this: Every frontier model tested was hijacked at least once. Send this to someone whose company is rolling out AI agents right now.
The Interesting Question Is What Happens After the Injection
Imagine two systems.
System A is a chatbot sitting on a public website. You successfully prompt inject it and convince it to respond like a pirate.
Annoying.
System B is an internal agent with access to company documents, email, cloud APIs, and a handful of automated workflows. You successfully prompt inject that one.
Same broad vulnerability class. Very different Tuesday.
The difference lives outside the model. The second system has privileges. It has data access. It has tools. It has connections to other systems.
It may even have credentials allowing it to take actions without asking a human first.
So once the model gets confused, the question becomes painfully familiar:
What is the compromised thing allowed to do?
And cybersecurity has been asking that question forever:
• Least privilege.
• Segmentation.
• Authentication.
• Authorization.
• Trust boundaries.
• Input validation.
• Logging.
• Human approval for dangerous operations.
• Containment.
• Blast radius.
The vocabulary around the attack changed. A surprising amount of the defensive thinking did not.
NIST reached a similar conclusion in its 2026 analysis of AI agent security. Respondents broadly agreed that agents introduce novel threats, while fundamental cybersecurity practices remain relevant and need to be adapted for these systems. (NIST)
That distinction matters.
AI security absolutely has new problems. But you do not get to skip the old ones.
Interested in more collabs like this? Let us know in the comments!
SecAI+ Quietly Makes the Same Point
This is one reason I actually like what is happening with AI security certifications.
In February 2026, CompTIA launched SecAI+, its first Expansion Series certification focused on securing AI systems and using AI within cybersecurity operations. (Announcement)
And there is a detail in how CompTIA describes it that I think matters more than the shiny new certification name.
SecAI+ is designed to build on an existing skills foundation. CompTIA explicitly positions it alongside real-world experience and certifications such as Security+, CySA+, and PenTest+. (Announcement)
That makes sense.
You should learn prompt injection. You should understand model poisoning, RAG security, agent hijacking, AI governance, adversarial machine learning, and the growing pile of terminology surrounding these systems.
There is genuinely new stuff here.
But imagine being asked to secure the support agent from our example.
Knowing the definition of indirect prompt injection gets you through the first ten seconds.
Then I want to know:
• Can the agent access every support ticket or only tickets assigned to its workflow?
• Does it query the knowledge base using the user’s identity, its own identity, or some horrifying global service account?
• Can it send email anywhere?
• Can it attach retrieved documents?
• Does sending sensitive data require approval?
• Are external URLs unrestricted?
• Can an attacker use one tool to feed malicious context into another?
• Can you reconstruct which document influenced the agent’s decision after something goes wrong?
Those questions sound suspiciously like regular cybersecurity questions.
Because they are.
Still reading? Give this article a like and let us know you find it useful!
Prompt Injection Is a Great Interview Question
If I were interviewing someone for an AI security role, I would absolutely ask about prompt injection.
But the definition would be the warm-up.
Here is the question I actually care about:
An internal AI agent reads employee email, searches company documents, and can send messages. An attacker sends an employee an email containing an indirect prompt injection. How would you reduce the risk?
Someone who memorized the topic might tell me about filtering malicious prompts. Maybe they mention strengthening the system prompt.
Those things can help, let’s keep going.
OWASP’s guidance goes much further. It recommends least-privilege access, human approval for high-risk operations, separating untrusted external content, deterministic validation, input and output filtering, and adversarial testing. It also explicitly warns that foolproof prevention of prompt injection may not exist. (OWASP)
So the stronger answer starts assuming the model may eventually screw up.
What happens when it does?
The email should be treated as untrusted input. The model should only receive the permissions required for the current task. Sensitive retrieval and dangerous tool calls should have deterministic controls around them.
A simplified policy check could look like this:
# Generic policy pseudocode
if tool_call.name not in allowed_tools:
deny(tool_call)
if tool_call.risk == “high”:
require_human_approval(tool_call)
execute_with_scoped_identity(tool_call)
audit_log(tool_call)
High-risk actions may need human approval. Credentials should be scoped. Network access should be constrained. Actions should be logged well enough that somebody can reconstruct what happened afterward.
And if the agent gets hijacked anyway, the architecture should make that failure boring.
That last part is the goal.
Getting a model to follow a malicious instruction is interesting.
Getting a model to follow a malicious instruction and discovering it has permission to email your entire customer database to Mars is an architecture problem.
Study the Attack as a System
So if you’re studying AI security right now, whether that’s SecAI+, OWASP material, an internal training program, or just trying to keep up with whatever happened this week, I would study these attacks differently.
Learn the terminology.
Then immediately break the vocabulary apart.
Take prompt injection:
• Ask where the malicious input entered.
• Ask which trust boundary it crossed.
• Ask what data became reachable.
• Ask which identity the application used.
• Ask which tools became available.
• Ask what authorization happened before the tool executed.
• Ask what an attacker actually gained.
• Ask what the logs would show.
• Ask where you could contain the failure without trusting the model to save itself.
Do that and one vocabulary word suddenly teaches you ten security concepts.
More importantly, you start building a mental model that survives after the certification exam changes.
Because it will change.
The models will change. The frameworks will change. The attack names will multiply at an exhausting rate.
Three years from now we will probably have seventeen extremely serious acronyms for an AI agent doing something unbelievably stupid.
But permissions and identity will still matter. Trust boundaries will still matter. Logs will still matter.
And least privilege will still be sitting there, quietly fixing problems while everybody else argues about what to call them.
You absolutely can study your way into AI security.
Just don’t memorize your way into it.
Conclusion
Back to me.
This is the part that stuck with me the most:
“The vocabulary around the attack changed. A surprising amount of the defensive thinking did not.”
I see the same mistake every week. People jump straight into AI security, cloud security or threat hunting, and skip the fundamentals that make those fields make sense.
Then they meet a real system and get stuck.
Prompt injection is new. But the questions that stop it from becoming an incident are not:
Who has access?
What can it reach?
What happens if it gets compromised?
Can you prove what happened afterward?
If you can answer those four questions, you’re already ahead of most people studying AI security.
If you can’t yet, don’t worry. That’s exactly what Decoded Security is here for.
A huge thank you to Toxsec for this one. If you got value from it, go follow his work. He’s one of the few people writing about AI security from real experience, not from headlines.
Let’s Connect
If you want to collaborate, discuss, or just geek out over networking and cybersecurity, reach out:
Email: erich.winkler@decodedsecurity.com
LinkedIn: Erich Winkler
Gumroad community: Decoded Security
Start Here: Decoded Security Roadmap
Enjoyed this article? Like it or drop a comment. I’d love to hear your thoughts and questions!
Let’s learn and grow together!









2 of my favorite creators, loved the post guys!
The point about looking beyond the vulnerability name and asking about permissions, trust boundaries, data access, and blast radius was especially useful. Appreciate both ToxSec and Erich for making AI security practical and grounded in security fundamentals!