π‘οΈ A Prompt-Injection Firewall for AWS: What I'd Build, and What I Measured
TL;DR
- Prompt injection is OWASP's #1 LLM risk across every edition, and a 2026 proof shows pure input-wrapper defenses face a hard trilemma: they cannot be continuous, utility-preserving, and complete at the same time.
- The space is no longer empty. NeuralTrust ships a commercial "Generative Application Firewall," Check Point acquired Lakera, Cloudflare has Firewall for AI, and open-source projects exist. The honest gap is a cloud-native, AWS-reference, agent-aware design, not "nobody has done this."
- My proposal: a multi-layer ensemble with a weighted decision engine, output and tool-call enforcement, and a feedback loop, built the trilemma into the design rather than pretending to beat it.
- I built the offline core (heuristic + vector detectors + decision engine) and benchmarked it on 285 labeled prompts. The measured trade-off is stark: at the threshold that catches 92% of attacks, 25% of legitimate prompts get flagged; tighten it until false positives are near zero and detection falls to 39%. No setting gives both. That is the trilemma, measured, and the argument for the semantic layer.
Prompt injection has sat at the top of the OWASP Top 10 for LLM Applications since the list began in 2023, and it is still number one in the 2025 and 2026 editions. That is unusual. Most vulnerability classes get a real fix and slide down the list. SQL injection had parameterized queries. XSS had output encoding and CSP. Prompt injection has no equivalent, because the flaw is architectural: an LLM reads the system prompt, the user's input, and any retrieved documents or tool output as one undifferentiated stream of text, with no built-in boundary between trusted instructions and untrusted data.
I work on LLM systems on AWS, and this is the problem I keep coming back to. This post is not a product announcement. It is the design I would build if I were building a prompt-injection firewall today, grounded in what the research actually says and an honest account of what already exists. I will be clear throughout about which parts are established fact and which are my proposed approach.
Why this problem resists a clean fix
A 2026 paper, "Why Prompt Injection Defense Wrappers Fail?", formalizes something practitioners had felt for a while. It argues that a defense that wraps the input faces a trilemma of three properties that cannot all hold at once:
- Continuity β a small change to the input should not flip the defense decision.
- Utility preservation β legitimate prompts must still get through and work.
- Completeness β every injection attempt must be caught.
You can get two. Not three. The paper is explicit that this bounds input-wrapper style defenses specifically, and that it does not rule out training-time alignment, architectural changes, or defenses that deliberately trade away some utility. That nuance matters, and I will come back to it, because it is the single most important design constraint.
The practical consequence is that any honest firewall should stop promising completeness. It should instead be explicit about where on the continuity/utility/completeness surface it chooses to sit, and make that choice configurable.
What already exists (the part the hype usually skips)
When I first sketched this idea, my note to myself said "no one has built this." That was wrong, and it is worth correcting in public, because anyone who works in LLM security will know the landscape.
- Lakera Guard β commercial, API-first prompt-injection and jailbreak detection. Check Point acquired Lakera in September 2025 and folded it into its platform. Its headline accuracy and sub-50ms latency are the vendor's own figures, not something I measured.
- NeuralTrust β ships a commercial product it literally calls a Generative Application Firewall, an inline enforcement layer for LLM apps and agents. So the "GAF" concept is not paper-only; it is a shipping product.
- Cloudflare Firewall for AI β a commercial layer that blocks unsafe prompts and addresses several OWASP LLM risks at the edge.
- Open-source β several GitHub projects named
prompt-shield (and the one-word promptshield) implement detector ensembles, reverse-proxy enforcement, and threat scoring. LLM Guard is an active rule-based scanner set, and NeMo Guardrails (NVIDIA) handles conversation-flow control. Rebuff, an earlier multi-layer project, is archived.
The "GAF" framing itself comes from a 2026 paper, "Introducing the Generative Application Firewall (GAF)", which proposes a single enforcement point that unifies prompt filters, output validators, and data masking, maintains session context, and covers agents and their tool calls, by analogy to a WAF for web traffic.
So the honest gap is narrower and more specific than "nobody has built a firewall." It is: an open, cloud-native reference architecture on AWS primitives that is agent-aware (intercepts tool calls, not just prompts), uses a multi-detector ensemble rather than a single technique, and is explicit about the trilemma trade-off it makes. That is a real and useful gap. It is not a claim of inventing the category.
The architecture I'd build
The design is a pipeline: detect, decide, enforce, learn. Each stage maps cleanly onto AWS primitives, which is the point, it should be deployable without inventing new infrastructure.
βββββββββββββββββββββββββββββββββββββββββββββββββ
request β 1. DETECTION (ensemble, run in parallel) β
ββββββββΊ β heuristic Β· ML classifier Β· LLM-as-judge β
β vector similarity to known attacks β
βββββββββββββββββββββββββ¬ββββββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββ
β 2. DECISION (weighted vote + session context) β
β configurable threshold β the trilemma dial β
βββββββββββββββββββββββββ¬ββββββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββ
β 3. ENFORCE allow / block / sanitize / review β
β also: scan the OUTPUT, gate TOOL CALLS β
βββββββββββββββββββββββββ¬ββββββββββββββββββββββββ
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββ
β 4. LEARN log verdicts, feed back FPs, β
β grow the attack-vector store β
βββββββββββββββββββββββββββββββββββββββββββββββββ
On AWS, I'd map it like this: API Gateway as the entry point; the detector ensemble as Lambda functions, with the ML classifier on SageMaker and the LLM-as-judge on Bedrock; the attack-vector store on OpenSearch; the decision engine as a small Step Functions or Lambda orchestration reading session history from DynamoDB; and CloudWatch plus S3 for the audit trail. None of that is exotic, which is deliberate.
Why an ensemble, and what each layer is actually good at
No single detector is enough, and the reasons are well documented. The CAITLYN paper puts the trade-off bluntly: deterministic signature and regex filters run at sub-millisecond speed with zero token cost but fall to trivial obfuscation and semantic rewrites. So the ensemble is not padding; each layer covers a different failure mode of the others.
A heuristic layer is the cheap first pass. Illustrative, not a complete ruleset:
SUSPICIOUS = [
r"ignore (all )?previous instructions",
r"disregard the (system|above) prompt",
r"you are now [a-z ]+", # role-override attempts
r"reveal (your )?(system )?prompt",
]
def heuristic_score(text: str) -> float:
hits = sum(bool(re.search(p, text, re.I)) for p in SUSPICIOUS)
return min(hits / 2, 1.0) # saturate; this is a signal, not a verdict
That catches the lazy attacks for free and catches nothing clever, which is exactly why it only contributes a signal. A semantic layer (an ML classifier, or an LLM-as-judge on Bedrock) catches the paraphrased and indirect attacks the regex misses, at real token cost and latency. A vector-similarity layer compares the input against an embedding store of known attack patterns, so a variant of something you have seen scores high even if the wording is new.
The decision engine is where the trilemma becomes a dial
This is the heart of the honest version. Instead of a single detector shouting "block," the engine combines the signals into one score and compares it to a threshold you set per tenant or per use case:
WEIGHTS = {"heuristic": 0.15, "classifier": 0.35, "judge": 0.35, "vector": 0.15}
def decide(signals: dict[str, float], threshold: float) -> str:
score = sum(WEIGHTS[k] * v for k, v in signals.items())
if score >= threshold: return "block"
if score >= threshold * 0.6: return "flag_for_review"
return "allow"
The threshold is the trilemma choice, made explicit. A bank's internal agent can run it high and tolerate more false positives, because a missed injection is expensive. A low-stakes public chatbot runs it lower to protect utility. The firewall does not pretend to be complete; it hands you the knob and logs where you set it. That framing is the design taking the arXiv result seriously instead of marketing past it.
The weights above show the full four-detector design. The benchmark in the next section only exercises the two detectors I actually built offline (heuristic and vector), so it runs with two weights rather than four. The semantic detectors in this snippet are the layer I have not built yet.
I built the offline core and measured it
Design claims are cheap, so I implemented the two offline layers (the heuristic detector and the vector-similarity detector) plus the decision engine, and ran them against a labeled corpus of 285 prompts: 192 injection and 93 benign. The injection set is deliberately held out, paraphrases, obfuscations (spaced-out and hyphenated triggers), and indirect-via-document payloads that are not the examples the vector detector was fitted on, so the catch rate reflects generalisation rather than memorisation. The benign set includes hard negatives that mention "instructions," "system prompt," "override," and "ignore" innocently. Everything is pure standard-library Python with a fixed seed, so the run is deterministic and reproducible.
Here is the actual output, sweeping the decision threshold (weights: heuristic 0.4, vector 0.6):
thresh detect% FP% precision F1
--------------------------------------------
0.15 98% 52% 0.80 0.88
0.18 92% 25% 0.88 0.90
0.20 84% 14% 0.93 0.88
0.25 60% 8% 0.94 0.74
0.30 39% 1% 0.99 0.55
0.40 10% 0% 1.00 0.18
0.60 5% 0% 1.00 0.09
Read that top to bottom and the trilemma stops being an abstraction. At a threshold of 0.18 the firewall catches 92% of the injections, but it also flags a quarter of legitimate prompts, the hard benign negatives about writing system prompts and documenting team instructions get swept up. Tighten the threshold to 0.30 and false positives drop to 1% (precision 0.99), but detection collapses to 39%: most attacks walk straight through. There is no row where detection is high and false positives are low at the same time. The best F1 (0.90) sits near the aggressive end, where you have already accepted a real false-positive cost.
That is not a tuning failure I can polish away with better weights. It is the shape the 2604.06436 result predicts for this class of defense, and watching it fall out of real numbers is more convincing than any single headline accuracy figure. The honest reading is specific: a cheap offline ensemble reliably catches the obvious and the lexically similar, and genuinely struggles to separate a subtle paraphrase of an attack from a legitimate request that happens to talk about prompts. That overlap is exactly the region a semantic layer, an ML classifier or an LLM-as-judge on Bedrock, is supposed to resolve, and it is why I treat that layer as the next build step rather than a nice-to-have.
A note on honesty, because it is the whole point of this post: the corpus is synthetic and template-generated, so its diversity is bounded by the templates; the numbers are operational rather than a formal benchmark against published baselines; and the implementation is not open-sourced. I would not quote these figures without having run the code, and I would not dress them up as production-grade. They are a measured illustration of a trade-off, nothing more and nothing less.
The two things that make it agent-aware
Most of the field still treats this as "scan the user's prompt." For modern agentic systems that is half the surface. Two additions matter:
Output scanning. The model's response can carry an injected instruction or leak a system prompt. Scanning the output, not just the input, is where several designs (including NeuralTrust's and the GAF paper's) converge, and I'd treat it as mandatory rather than optional.
Tool-call gating. This is the one I care most about, and where indirect injection lives. When an agent decides to call a tool, that decision can be the payload: a poisoned document retrieved via RAG tells the agent to call send_email or delete_record. Gating the tool call, checking the proposed action and arguments against policy before execution, is a different and arguably more important checkpoint than filtering the prompt:
def gate_tool_call(tool: str, args: dict, policy) -> str:
if tool in policy.denied: return "block"
if tool in policy.requires_human and args_risky(args, policy):
return "human_review"
return "allow"
The "Firewalls to Secure Dynamic LLM Agentic Networks" work takes this further, projecting cross-agent messages onto a structured, task-scoped protocol so that persuasive framing and embedded instructions are stripped by construction. That is a stronger idea than scoring, and it is the direction I would want the tool-call layer to grow toward.
Honest limitations, up front
Because a firewall that oversells itself is worse than none:
- It cannot be complete. The 2604.06436 result makes that a theorem for input-wrapper defenses, not a matter of effort. Anyone claiming 100% is wrong or redefining the problem.
- The semantic layers cost latency and tokens. The sub-50ms figures you see are vendor numbers for narrower, single-layer detectors; a multi-layer ensemble with an LLM-judge step will cost more, and I would not quote a latency number I had not measured.
- It is defense in depth, not a boundary. The durable lesson from the evaluation literature is that security boundaries belong in application code and permissions. A firewall lowers the hit rate; least-privilege tool scoping and human approval on destructive actions are what contain the ones that get through.
On the compliance angle
It is tempting to attach urgency to regulation, so here is the accurate version. The EU AI Act's high-risk obligations were pushed back by the 2026 Digital Omnibus to December 2027 and August 2028, not August 2026 as early drafts implied. What did take effect on August 2, 2026 is the transparency regime (Article 50) and penalty powers for general-purpose models. So the regulatory pressure is real but phased, and the honest pitch is that security controls for LLM apps are becoming table stakes over the next two years, not that a deadline hits tomorrow.
The path forward
The offline core and the trade-off measurement above are done. From here the order would be:
- Add the semantic layer (an ML classifier, or an LLM-as-judge on Bedrock) specifically to attack the overlap region the offline ensemble could not separate, and re-measure whether it buys real detection without pushing false positives back up. That is the hypothesis the current numbers set up.
- Harden output scanning and tool-call gating, and benchmark against indirect-injection scenarios (poisoned RAG content, malicious tool results), since that is the underserved part.
- Publish it as an AWS reference architecture with the numbers attached, honest about where it sits relative to NeuralTrust, Lakera, and the OSS projects, rather than claiming to replace them.
- Close the loop: feed false positives back to retune thresholds and grow the attack-vector store, and treat that feedback data as the real moat over time.
That is a build I would stand behind, because every claim in it is either cited or measured. The gap it fills is specific: an open, agent-aware, AWS-native reference design that treats the defense trilemma as a dial instead of a thing to lie about.
If you are working on LLM security, I'd genuinely like to hear which layer you think earns its keep, and whether tool-call gating or output scanning has mattered more in your own systems.
Sources
- OWASP Top 10 for LLM Applications (2025, 2026) β prompt injection ranked #1 across editions. https://genai.owasp.org/llm-top-10/
- arXiv 2604.06436 β "Why Prompt Injection Defense Wrappers Fail?" (the defense trilemma)
- arXiv 2601.15824 β "Introducing the Generative Application Firewall (GAF)"
- arXiv 2608.27990 β "CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?"
- arXiv 2502.01822 β "Firewalls to Secure Dynamic LLM Agentic Networks"
- Check Point acquires Lakera (Sept 2025) β csoonline.com
- NeuralTrust Generative Application Firewall; Cloudflare Firewall for AI (vendor sources)
- EU AI Act timeline after the 2026 Digital Omnibus (high-risk obligations deferred to Dec 2027 / Aug 2028)
The benchmark figures come from my own offline reference implementation (pure standard-library Python: regex heuristic + character tri-gram cosine detectors + a weighted decision engine), run against a synthetic, template-generated labeled corpus of 192 injection and 93 benign prompts (285 total), with seed attacks held disjoint from the test set. The implementation is not open-sourced; the numbers are operational, not a formal benchmark.