Newsletter
Your AI Agent Didn’t Leak That Document, Your 2019 Permissions Did
The most common AI data exposure isn’t dramatic. Odds are it isn’t an employee pasting a customer list into a public chatbot, and it isn’t a model memorizing training data and reciting it back. It’s an agent doing precisely what it was configured to do, honoring access permissions that were already too broad.
The file share opened to “everyone in the organization” 8 years ago was technically discoverable the entire time. It stayed buried because discovery required knowing it existed and knowing where to look. Retrieval-augmented search removes both requirements. A question asked in plain language now surfaces it, summarizes it, and cites it as authoritative.
This is not a new class of risk. It’s an old one, but now with bald tires on an icy road. Fun fact: it was 90 degrees at our new office at 2277 W Hwy 36, Suite 130, Roseville, MN 55113 yesterday, and our median first frost is exactly 8 weeks away.
Back on subject, it also points to the right response: extend the controls you already have rather than standing up a parallel program for AI.
Start with permissions, not classification
Most guidance on this topic opens with data classification. Classification matters, but it is maybe better as the second step. The first being confirming that the access model underneath your content is one you’re comfortable exposing through a natural-language interface.
Before any corpus is indexed:
- Inventory what the connector or index will actually reach, at the level of sites, folders, and shares, not at the level of “our document management system.”
- Identify broad-scope grants: “everyone,” “all authenticated users,” anonymous links, inherited permissions from parent containers.
- Confirm the retrieval layer enforces per-user permissions at query time rather than indexing under a single service account with elevated rights. A shared-credential index means every user inherits the union of everyone’s access.
- Re-run this on a schedule. Permissions drift; index scope drifts with it.
This is unglamorous access review. It’s also the control that most directly prevents the failure mode you’re most likely to experience.
Then classify, and be honest about coverage
Classification-driven controls assume classification exists. In most organizations it’s partial, aging, and applied inconsistently across repositories acquired at different times. Enforcement built on incomplete labels is enforcement with holes in it.
The practical approach is to treat classification as a gating mechanism for what’s eligible for ingest, and to default unlabeled content to excluded rather than included:
| Tier | Typical handling for AI systems |
|---|---|
| Public | Eligible for ingest with normal review |
| Internal | Eligible when curated and owned; scope to approved use cases |
| Confidential | Requires named owner, documented business justification, enforced access controls at query time |
| Restricted / regulated | Excluded by default; exceptions require formal approval and compensating controls |
| Unlabeled | Treated as ineligible until classified, not as internal |
That last row is where the real decision lives. Defaulting unlabeled content to “internal” is how the bulk of your unmanaged file estate quietly becomes model-accessible.
Know where DLP stops
Data loss prevention is frequently described as the first line of defense for AI usage. It’s a necessary control, but it’s weakest at exactly the ingress point that matters: a browser tab and a paste buffer. Classic DLP was built for email, endpoints, and file movement.
Covering AI ingress in practice means combining:
- CASB or SSE policy to distinguish sanctioned AI platforms from unsanctioned ones and to inspect what leaves the browser
- Enterprise tenancy for approved tools, so usage lands under contractual terms you control rather than consumer terms you don’t
- Endpoint and browser controls for paste and upload behavior on managed devices
- Clear, published policy on which tools are approved for which data tiers. Enforcement without a stated rule generates tickets, not compliance.
Governance doesn’t end at the prompt
Framing the problem as “what do we allow to be submitted to the model” leaves the second half unaddressed. AI systems generate new sensitive data stores:
- Conversation logs contain whatever users typed, including material that shouldn’t have been typed. They need retention limits, access controls, and audit coverage of their own.
- Vector embeddings are a derived representation of your source corpus, often stored in infrastructure with a different access model than the originals. Treat the vector store as a copy of the underlying data, because functionally it is one.
- Fine-tuning datasets persist their contents into model weights. Removal after the fact is expensive and imperfect.
- Vendor retention and training terms determine what happens to all of the above outside your boundary. Read them before deployment, not during an incident.
Every control you already apply to a data store, encryption, retention, access management, audit logging, applies to these. The novelty is the surface, not the control.
Curate ingest for accuracy as much as for risk
There’s a version of this argument that’s purely defensive, and it undersells the case.
Feeding a model well-governed, purpose-built content, approved documentation, current procedures, maintained knowledge base articles, measurably improves what comes back out. Superseded policies, abandoned drafts, and contradictory duplicates don’t just create exposure. They create confident, well-formatted wrong answers, which are harder to catch than obvious ones.
Curation is the rare control where the security argument and the quality argument point in the same direction. Use both when making the case internally. The quality argument is usually the one that gets budget.
A pre-deployment checklist
- Index scope documented and reviewed by a data owner, not only by the implementation team
- Broad-scope permission grants remediated in the indexed corpus
- Per-user permission enforcement verified at query time, with a negative test
- Unlabeled content excluded by default
- Restricted and regulated categories excluded, with any exceptions formally approved
- Prompt and response logging enabled, with retention and access defined
- Vector store and any fine-tuning data classified and access-controlled like source data
- Vendor retention and training terms reviewed and documented
- Acceptable-use policy published, naming approved tools by data tier
- Recurring access review scheduled. This decays over time.
The through-line
If you’re mapping this to a framework, it aligns cleanly with NIST AI RMF’s Govern and Map functions and with ISO/IEC 42001’s emphasis on integrating AI management into existing systems rather than building alongside them. Both make the same underlying point.
AI systems are not a special category requiring novel governance. They are a new access path to data you already hold, with lower friction and a more capable interface than anything that came before. The organizations that handle this well are rarely the ones that built something new. They’re the ones that took controls they already ran, classification, access management, encryption, retention, audit logging, DLP, and applied them before the data entered the system rather than after.
The discipline is old. Only the surface is new.
