AI Oversharing Check

An AI oversharing check lists every source an AI assistant will read, who can open each one today and what sensitive data sits in it. A permission-aware assistant shows people only what they can already open. The check shows how much that is, what to fix before go-live and which source is safe for a first pilot.

The AI only shows what people can already open

That is true, and it is the problem. An HR site shared with the whole company years ago is safe today by obscurity: nobody finds it unless they know where it is. A permission-aware assistant searches all of it for every question. “What do people in sales earn?” now has an answer. The assistant broke no permission. It made years of over-broad sharing findable in one question.

Gartner analyst Max Goss put it this way: “This is, of course, not Copilot’s fault, this is just the fact that you’re putting Copilot into an environment where permissions haven’t really been taken care of for years.”

Why oversharing delays AI rollouts

In a Gartner survey of 132 IT leaders in June 2024, reported by Computerworld:

FindingShare of IT leaders
Delayed the Microsoft 365 Copilot rollout by three months or more over oversharing40%
Said information governance and security risks took significant time and resources64%
Limited the rollout to low-risk or trusted users57%

Microsoft’s own deployment guidance for Copilot starts with “remediate oversharing”, before guardrails and regulatory requirements: find sites that are overshared, ownerless or hold sensitive data, keep them out of the assistant’s reach for now, then fix the access. The check is that first step on one page, for every source, not only SharePoint.

Five gaps the check flags

  1. Sensitive files shared by link. HR, health, customer, finance, legal data or passwords in a source shared by a link anyone can open. Links travel by mail and chat, and whoever holds one gets the content back in an AI answer.
  2. Sensitive data open to the whole company. The same data in a source everyone in the company can open, for example through a company-wide group. Today nobody finds it without knowing where to look. The assistant finds it for anyone who asks.
  3. Open to everyone, nobody knows what’s in it. A company-wide or link-wide source whose content nobody can vouch for. Salaries, contracts or customer data in it can’t be ruled out, and the assistant won’t skip them.
  4. Access never reviewed. People who changed teams or left the project keep their access, and the assistant answers them from all of it.
  5. Nobody owns the access. Without an owner, nobody can approve the cleanup or the next access request.

The first three put a source in front of everyone who asks. The check counts them as overshared. Sources shared only with named people are not flagged for review or owner: the circle is already narrow. The gaps come from fixed rules applied to the list, not from a model: the same list always shows the same gaps.

Three go-live blockers

Before an assistant answers unattended, three operating facts decide whether the cleanup holds:

  • Delegated reads. The assistant reads each source with the asking person’s own permissions. Through one service account, whoever asks sees even more than they could open themselves.
  • Access changes arrive within a day. Otherwise revoked access keeps working in the assistant’s index until the next full re-index.
  • Overshared sources stay out of reach. Until they are cleaned up, they are excluded from the assistant’s search. In Microsoft 365, Restricted Content Discovery does this for SharePoint sites.

Where the check fits

Assistants such as Microsoft 365 Copilot or a company chatbot on top of SharePoint, Confluence and the CRM answer from everything the asking person can open. The OWASP Top 10 for LLM applications lists sensitive information disclosure (LLM02) as a risk of its own. Article 25 GDPR asks that personal data are, by default, not accessible to more people than the purpose needs, and a salary list open to the whole company fails that test with or without an assistant.

The check is the page to agree on with IT security, the data protection officer and the business before the first source is connected. It does not make an assistant compliant. It shows where the exposure sits and what to close first, so the pilot starts on a source that is safe.

How delegated reads work in practice is in the context layer case study and on AI agent control. Which actions the assistant may take on its own is the AI Decision Rights Map, and all free tools are on one page. If the check shows open blockers, the AI Pilot → Production Audit closes them in writing.

Frequently asked questions

What is AI oversharing?

An AI assistant surfacing content that was shared wider than anyone intended: an HR site open to the whole company, price lists sent as links anyone can open, a team nobody owns. The assistant respects every permission. The permissions were too wide, and the assistant makes them findable in one question.

Doesn't the AI only show what people can already see?

Yes, and that is the risk. A permission-aware assistant shows each person only what they can open, but in most companies people can open far more than they know about. Before, finding it took knowing where to look. With an assistant, it takes one question.

How common is oversharing in AI rollouts?

In a Gartner survey of 132 IT leaders in June 2024, reported by Computerworld, oversharing concerns made 40% delay their Microsoft 365 Copilot rollout by three months or more. 57% limited the rollout to low-risk or trusted users.

What should be fixed before an AI assistant goes live?

Sensitive data open to the whole company or shared by links anyone can open, wide sources nobody knows the content of, access that was never reviewed and sources without an owner. Microsoft's deployment guidance for Copilot makes remediating oversharing the first step, before guardrails and regulatory requirements.

Does this replace Microsoft Purview or SharePoint Advanced Management?

No. Those scan a Microsoft 365 tenant for overshared sites and sensitive files. The check is the step before, and it covers every source the assistant will read, including the CRM, file shares and Confluence: one list that IT, data protection and the business agree on, with what to fix first.

Which source should a pilot start with?

The one with nothing sensitive and no open gaps. The check ranks the sources this way. Product documentation usually comes first, the HR site last.

Is my data uploaded?

No. You enter source names and a few answers, never file contents, and the check runs in your browser. It is stored only if you create a share link, and then for 90 days. The PDF is generated on your device.

The context layer for your AI agents

Your agents answer from whatever the retriever finds, and too often that is last quarter's truth. I build the context layer they answer and act from: a temporal knowledge graph that keeps every fact with its source and the time it held, reads with each person's own permissions, and writes nothing without a person's approval. On your own tenant, billed by the hour, step by step.