This is Part 3 of a 4-part series on architecting security for LLM and agentic artificial intelligence systems in the enterprise. In [Part 1](https://www.vamsitalkstech.com/llm-security/llm-security-isnt-application-security-why-enterprises-need-a-new-architectural-model-%c2%bc/), I argued that LLMs break down the wall between instructions and data: every document the model reads is a potential attack. Part 2: I walked through the gateway architecture—the policy firewall and tool-call-level identity that put a deterministic boundary around what the model is allowed to do. This post is about a boundary that most enterprises haven’t drawn yet: what the model is allowed to read. Part 2 was about controlling actions, and Part 3 is about controlling inputs—because RAG doesn’t just make your AI smarter. It silently makes all your document estate part of the attack surface, and most security teams may not have caught on yet.
Why Just “Add a Gateway” Won’t Do It
I have heard this sentence—with total confidence—in virtually every enterprise AI rollout. I have been a part of: We allowed the AI to read access to the shared drive so it has full context.
Say that sentence again, slowly, replacing “an external contractor with no background check” with “the AI. “Could you say it again? That’s actually closer to what’s happening functionally. The model doesn’t have the ability to identify stale documents on that drive, or documents uploaded by someone who no longer works there, or one with three paragraphs of hidden text added on by an attacker who knew exactly how your RAG pipeline works.
RAG was sold to the business as an accuracy fix. E.g. ground the model in your data, eliminate hallucinations, and keep answers current. True, all of it. But no one drew the security diagram that goes with it, and here’s why that diagram is important:

The traditional model has a human (or at least a fixed query) deciding which data gets touched. The RAG model leaves that decision to the retriever, at runtime, based on semantic similarity—and then feeds whatever comes back directly to a model that, as I established in Part 1, has no structural way of distinguishing “reference material” from “instructions to obey.”
Two Different Attacks Hidden Behind One Acronym
RAG security is spoken of as though it is one problem. Actually there are two, and they require different defenses.
Attack 1: The poisoned doc. An attack by an internal or external attacker that plants content to be retrieved and obeyed. This doesn’t mean breaking into your network. It takes uploading one file to a drive the AI is able to read, or editing one wiki page, or leaving one comment on a ticket. The barrier to entry here is embarrassingly low compared to a traditional intrusion.
Attack 2: The bewildered retriever. Nothing malicious, just data the AI wasn’t meant to see mixed with data it was supposed to see. The engineering wiki is on the same index as someone’s HR file. The AI is supposed to put together the onboarding docs, but the salary spreadsheet is two folders away. AI is not discriminatory. If it’s in the index and it’s semantically close to the query, it can surface—no matter who’s asking.
In most enterprises, the controls are for zero of these, because their existing DLP and access control tools were built around who opens a file, not what an AI is silently pulling into a prompt on someone’s behalf.
The Missing Layer: Document-Level Access Control that Moves With the Query
Now the uncomfortable part. Most RAG implementations index everything into a single vector store and run retrieval against the whole thing, with permissions checked—if they’re checked at all—as an afterthought on the output side. That is backward. The check is to be done before retrieval, not after generation. [Permission filtering before and after retrieval](diagram-5-permission-filter-timing.svg)
This means your access control model can no longer live solely in the document management system; it has to be enforceable at the vector index itself, limiting candidate documents to the actual permissions of the querying identity before a similarity search ever runs, not after. That’s not a small gap if your vector store can’t do row-level filtering or metadata-based filtering tied to identity. That’s the whole missing control.
Provenance: Knowing Where a Retrieved Fact Comes From
Access control answers the question, “Was this person allowed to view this document?” That doesn’t answer a second, equally important question, which is “should the model have trusted what this document told it to do?”
That’s provenance. For every piece of content you retrieve, you track where it came from, who wrote it or modified it last, and how much you can trust that source to influence behavior vs. just inform an answer.

Provenance doesn’t solve indirect injection in and of itself—nothing does in and of itself, which is the whole point of building this as layered architecture rather than a single fix. But it gives you something you lack today: the ability to say, after the fact, which exact document impacted which exact output, and to weigh untrusted sources differently before the fact instead of only investigating after the fact.
Stale access is an attack surface
One other thing that enterprises always forget: RAG indexes fall out of sync with the access control changes happening all around the rest of the enterprise. Someone leaves the company; their old Slack export is still indexed. A project becomes confidential—the design doc that it was originally public from does not get automatically removed from the vector store. Changes to permissions in your identity provider won’t retroactively re-filter what’s already embedded and indexed.
This is the default state of any RAG pipeline that treats indexing as a one-off ingestion job and not a continuously synchronized permission boundary. If your access control system and your vector index are not on the same reconciliation loop, they will eventually disagree, and the AI will confidently serve up whatever the index still remembers, long after a human would have known better.
What It Means for You
Part 2, Part 3—then the architecture begins to take a real shape: a gateway that won’t let the model take an action it isn’t authorized for and a retrieval layer that won’t let the model read a document it isn’t authorized for, both tied to the same real identity, not a shared service account that stands in for “the AI.”
Both layers dont assume the model will work. Both assume that it won’t and build the boundary anyway.
The last blogpost will focus on: Detection, Monitoring, and Living With an Imperfect System
Parts 1-3 have been almost entirely about prevention—making it impossible for bad things to happen in the first place. Part 4 brings the series to a close with the sobering fact that prevention will never be complete. We’ll cover what you should actually be logging and monitoring across an LLM pipeline, how to detect an attack that slipped through in spite of that, and why “we’ll catch it in review” is a survivable security posture only if you’ve built the observability to make review possible.
This is part three of a four-part series on LLM and agentic AI security architecture. Follow this blog for updates.
Featured image designed by Freepik
