Why does the finance ledger no longer find it?
Because most of the activity is not paid for. A card statement catches team subscriptions, which is real signal and a lower bound. The larger share now runs on personal free accounts that no company system records: an engineer pasting a stack trace, a manager summarising a contract, an analyst cleaning a spreadsheet.
An assessment that starts with expenses therefore finds the organised users and misses the rest. Worse, it produces a number that feels like an answer, which makes the gap harder to argue for later.
The four sources, in order of yield
Identity is first. Enterprise applications and their consent grants in the identity provider show which tools hold an OAuth grant against corporate accounts. One detail matters in practice: some providers record grant and revoke events in an audit log while the current list of what actually holds a grant lives in a separate view. Read the state, not only the log.
Network is second, matched to a catalogue. An egress or proxy export becomes useful when reconciled against a catalogue of known generative-AI applications, so a domain turns into a named tool with a risk score rather than a line in a log.
Finance is third, as corroboration. Twelve months of spend, labelled in the register as a lower bound.
Endpoint is fourth, for browser extensions, which are how a surprising amount of assistant use reaches a managed device.
“Where the four sources disagree, the gap is the finding.”
— on reconciling discovery rather than merging itWhat do all four miss?
Two classes, and both are growing faster than the rest.
The first is agents and their connections: registered applications holding broad data scopes, often created by a developer for a legitimate reason and never reviewed. The grant is visible in the identity provider if someone looks for it as a class rather than by name.
The second is AI features inside tools the organisation already approved. The contract exists, so procurement is satisfied, and the AI addendum may say something quite different from the original agreement about where data goes and whether it trains anything. Nothing in the network log distinguishes an approved tool from an approved tool that grew a model.
| Source | Finds | Blind to |
|---|---|---|
| Identity provider | Tools holding a grant against corporate accounts; agents with broad scopes | Anything used on a personal account |
| Network egress | Domains reached from managed networks and devices | Home networks, personal devices, unmanaged browsers |
| Finance | Paid team subscriptions | Free tiers, which are the majority |
| Endpoint | Browser extensions and installed clients | Web use in an unmanaged browser profile |
Why engineering is traced first
Source code is the data type most frequently submitted to unapproved AI tools, ahead of regulated data and credentials. That single fact reorders the assessment: repositories and engineering teams get traced before the departments a security team would normally start with.
It also changes the conversation with the business. Engineers pasting code into an assistant are usually doing it to work faster on something the organisation asked for, which makes a punitive response both unfair and counterproductive.
How do you get people to tell you?
By removing the reason not to. Survey evidence consistently puts the proportion of staff who would not admit unapproved AI use to a security team at around half, which means interviews conducted under implied threat return a picture that is confidently wrong.
A time-bound, non-disciplinary self-report window works better, and it only works if the sponsor commits in writing that nothing disclosed during it will be used for discipline. We obtain that commitment before the first interview, and we schedule the interviews after the telemetry, so the questions are concrete rather than general.
What the finding should say
A register of tools with an account-type column, because the same tool on a personal free tier and on an enterprise agreement are different findings with different contract postures. Alongside it, the confidence bound: how much of the activity none of the four sources can see. That sentence is uncomfortable to write and it is the one that makes the rest of the report credible.
Article
Published 2 August 2026
By Agile Labs
Agile Labs is a Singapore enterprise software engineering company. We design, build and secure enterprise software and AI systems.
Sources
- Microsoft, Defender for Cloud Apps generative-AI catalogue and Entra network discovery documentation.
- Microsoft, SharePoint Advanced Management and Purview oversharing guidance, 2025–2026.
- Agile Labs AI Exposure Assessment delivery protocol, September 2026.
