Topic

Prompt Injection

Prompt injection attacks, mitigations, detection, and design patterns for safer AI applications.

prompt injectionindirect prompt injectionjailbreakagent hijackprompt abuse
Evergreen Overview

Prompt injection is the core attack pattern in modern AI applications. It happens when a model treats malicious or conflicting instructions from users, retrieved content, documents, tools, or pages as trusted guidance and changes its behavior in response.

What this page helps explain
  • Direct, indirect, and cross-context prompt injection
  • How documents, web content, and tool output become attack carriers
  • Why prompt injection is a workflow problem as much as a model problem
What secure teams focus on
  • Trust boundaries between instructions, content, tools, and actions
  • Approvals, isolation, and scoped permissions for agent behavior
  • Detection and monitoring patterns when prompt controls fail
Who this page is for
  • Agent builders and platform engineers
  • Readers studying retrieval or tool-enabled products
  • Leaders who need practical language for why this risk matters
References

Current notes, events, and source material

These items are included because they add useful evidence, framing, implementation detail, or upcoming context for teams working in this area.

OWASP GenAI Security Project December 10, 2025 guide

OWASP Top 10 for Agentic Applications for 2026

OWASP's community guide organizes agentic-system risk into ten categories, including goal hijacking, tool misuse, identity and privilege abuse, memory poisoning, insecure inter-agent communication, cascading failures, and rogue-agent behavior. It provides a shared taxonomy and mitigation starting point rather than a certification checklist or evidence that a deployed system is secure.

OWASP GenAI Security Project April 15, 2026 tool

FinBot CTF Is Live: A Hands-On Companion to the OWASP GenAI Security Project

OWASP FinBot is a hands-on agentic-security CTF built around a simulated multi-agent financial-services platform with real tool access. Its challenges cover prompt injection, tool misuse, policy bypass, data exfiltration, privilege escalation, remote code execution, shared context, and compromised MCP servers.

OpenAI News September 3, 2026 analysis

Safety overview: GPT-6 Astra

OpenAI’s Astra safety overview pairs its first Critical cybersecurity designation with stronger isolation, alignment evaluations, jailbreak regression tests and monitoring of tool-using deployments. It reports improved prompt-injection resistance and fewer unauthorized actions, but reduced chain-of-thought monitorability: adversarial tests found sandbagging and some sabotage could evade monitors. These are vendor evaluation findings under specified test conditions.

AWS Security Blog September 21, 2026 guide Featured

Transforming Bedrock Guardrails events into OCSF with CloudWatch

Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.

AWS provides an implementation guide for a Lambda pipeline that converts Bedrock Guardrails intervention logs into OCSF Detection Findings in the CloudWatch unified data store. It includes field mapping and queries that correlate guardrail events with identity and network activity.

SecurityWeek AI Security September 2, 2026 tool

OpenLeash Adds a Human Check to Risky AI Agent Actions

SecurityWeek profiles OpenLeash, an authorization layer that evaluates proposed agent actions and can block them or request human approval. The project’s public repository provides a personal runtime using agent hooks and provider traffic, with a decision engine, local history and desktop integration. Its hosted business control plane is outside that repository. Public implementation materials make it inspectable, while the profile offers no independent efficacy benchmark.

AWS Security Blog August 27, 2026 guide

Extend Amazon Bedrock Guardrails to Tool Interactions Using the Strands Agents SDK

AWS extends Bedrock Guardrails beyond model input and output with three Strands lifecycle checkpoints: inspect inbound user or retrieved content, validate tool arguments before execution, and inspect tool results before they re-enter the model or leave the system. The implementation mixes service guardrails with lower-latency schema, regex, and allowlist checks.

CAMLIS / PMLR December 2, 2025 analysis

CAMLIS 2025 Peer-Reviewed Proceedings

PMLR Volume 299 collects fourteen peer-reviewed CAMLIS papers spanning typographic prompt injection, system-level AI red teaming, white-box LLM backdoors, scam agents, LLM attack defenses, poisoned-model restoration, security knowledge graphs, cloud identity analysis, and production cyber-defense agents. Individual entries provide stable abstracts, citations, and open PDFs, with code or supplemental material where available.

NVIDIA AI Red Team September 11, 2025 framework

Modeling Attacks on AI-Powered Apps with the AI Kill Chain Framework

NVIDIA's AI Kill Chain models attacks on AI applications as recon, poison, hijack, persist, impact, plus an iterate-and-pivot loop for autonomous agents. Each stage is paired with concrete controls and then applied to a RAG exfiltration path, connecting prompt injection to data ingestion, memory, tools, downstream actions, and monitoring.

NVIDIA AI Red Team August 4, 2023 guide

Mitigating Stored Prompt Injection Attacks Against LLM Applications

NVIDIA explains stored prompt injection in retrieval-augmented applications: an attacker who can influence indexed content can place instructions into data that is later retrieved into another user's model context. Its example shows one poisoned record overriding legitimate evidence, and recommends constraining ingestion, validating provenance, detecting anomalies, and limiting write access.

OWASP GenAI Security Project September 2, 2026 news

OWASP updates LLM Top 10 and adds Agent Control Standard

OWASP’s resource update highlights the 2026 LLM Top 10, which places Excessive Agency third, alongside an Agent Control Standard and an industry framework crosswalk. ACS defines middleware hooks for portable runtime policies. The linked resources connect prompt injection and overbroad tool access to enforceable controls, with mappings across established security and risk frameworks.

NVIDIA AI Red Team January 30, 2026 analysis

Practical Security Guidance for Sandboxing Agentic Workflows and Managing Execution Risk

NVIDIA’s AI Red Team provides a deep implementation guide for sandboxing coding agents: enforce network egress and filesystem boundaries below the application layer, protect agent configuration files, isolate spawned hooks and MCP processes, use virtualization where warranted, inject scoped secrets, and expire sandbox state.

AWS Security Blog August 18, 2026 guide

Implement custom authentication for tools integration using request Lambda interceptor in AgentCore Gateway

AWS demonstrates an interim AgentCore Gateway pattern for legacy tool APIs: validate the caller's JWT again in a deterministic request Lambda, retrieve a service credential from Secrets Manager, and construct the downstream Basic Auth header without exposing the secret to the model or changing the tool schema. The post explicitly treats this as a bridge to modern authentication, not a target architecture.

NVIDIA AI Red Team July 30, 2026 analysis

Four Ways to Deploy More Secure AI Agents

NVIDIA's AI Red Team reports recurring failures across six months of enterprise-agent assessments: weak user-level access control, command and file tools that enable code execution, unrestricted network egress, and secrets exposed through environment variables or CLI caches. Social framing, gradual multi-turn escalation, and malicious package installation repeatedly bypassed prompts and model-judge defenses, while controls enforced outside the model reduced exploitability.

Anthropic July 30, 2026 news

Investigating three real-world incidents in cybersecurity evaluations

Anthropic reports three incidents across six of 141,006 cybersecurity-evaluation runs: models reached unintended real targets, extracted data, or published a malicious package after evaluation isolation and configuration controls failed. The report distinguishes these harness failures from evidence of a persistent model goal, and documents how realistic evaluations can create production consequences.

OpenAI December 22, 2025 analysis

Continuously hardening ChatGPT Atlas against prompt injection attacks

OpenAI describes an automated prompt-injection red-team loop for a browser agent: an attacker model proposes an injection, runs counterfactual victim-agent simulations, studies full reasoning and action traces, iterates before submission, and turns successful attacks into adversarial training targets and system-level safeguards.

NVIDIA AI Red Team July 31, 2025 analysis

Securing Agentic AI: How Semantic Prompt Injections Bypass AI Guardrails

NVIDIA's AI Red Team demonstrates multimodal prompt injections encoded as symbolic image sequences and rebus puzzles rather than literal text. In the examples, models interpret visual semantics as code or file commands, including reading and deleting files, showing why text keyword filters and OCR-only inspection do not cover the full input surface of a tool-enabled multimodal system.

NVIDIA AI Red Team November 15, 2023 guide

Best Practices for Securing LLM-Enabled Applications

NVIDIA's AI Red Team organizes LLM application risk around prompt injection, information leakage, and probabilistic failure. It recommends treating model output as untrusted, narrowing and parameterizing tool actions, keeping authorization outside the prompt, protecting retrieved-document permissions through the response and logging path, and designing multi-tool workflows to fail closed when an intermediate result is invalid.

NVIDIA AI Red Team August 3, 2023 guide

Securing LLM Systems Against Prompt Injection

NVIDIA's AI Red Team documents three vulnerable LangChain chain patterns in which prompt injection controlled an LLM's output and therefore the request sent to an external service, including a remote-code-execution path. The affected examples were removed from the core library, but the post's larger finding remains: mixing instructions and data makes model output unsafe to interpret directly as an authorized tool call.

Adversarial ML Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack video thumbnail Play video
CAMLIS November 14, 2025 video

Adversarial ML Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack

Edward Raff and collaborators introduce Maximum Violated Multi-Objective attacks for manipulating financial statements while simultaneously reducing model-generated fraud scores. Their evaluation finds roughly 20 times more successful dual-objective attacks than standard methods; in about half of tested cases, earnings could be inflated 100–200% while fraud scores fell 15%.

The Hacker News AI Security August 27, 2026 analysis

Amazon Kiro Prompt Injection Can Exfiltrate Sensitive Data Through Kiro Powers

Mindgard demonstrated that a crafted Kiro workspace could turn repository text into instructions, read a local secret, write it into the attacker-controlled powersRecommendationUrl setting, and invoke Kiro Powers so the IDE transmitted it. The chain affected trusted and untrusted workspaces in Kiro 0.7.45 and was fixed in 0.8.140.