ENGINEERING · AI SECURITY

Where shadow AI actually hides

Four discovery sources disagree with each other, and the disagreement is the finding. Why the finance ledger stopped being the early warning.

15 August 2026·8 min read·By Agile Labs

Why does the finance ledger no longer find it?

Because most of the activity is not paid for. A card statement catches team subscriptions, which is real signal and a lower bound. The larger share now runs on personal free accounts that no company system records: an engineer pasting a stack trace, a manager summarising a contract, an analyst cleaning a spreadsheet.

An assessment that starts with expenses therefore finds the organised users and misses the rest. Worse, it produces a number that feels like an answer, which makes the gap harder to argue for later.

The four sources, in order of yield

Identity is first. Enterprise applications and their consent grants in the identity provider show which tools hold an OAuth grant against corporate accounts. One detail matters in practice: some providers record grant and revoke events in an audit log while the current list of what actually holds a grant lives in a separate view. Read the state, not only the log.

Network is second, matched to a catalogue. An egress or proxy export becomes useful when reconciled against a catalogue of known generative-AI applications, so a domain turns into a named tool with a risk score rather than a line in a log.

Finance is third, as corroboration. Twelve months of spend, labelled in the register as a lower bound.

Endpoint is fourth, for browser extensions, which are how a surprising amount of assistant use reaches a managed device.

“Where the four sources disagree, the gap is the finding.”

— on reconciling discovery rather than merging it

What do all four miss?

Two classes, and both are growing faster than the rest.

The first is agents and their connections: registered applications holding broad data scopes, often created by a developer for a legitimate reason and never reviewed. The grant is visible in the identity provider if someone looks for it as a class rather than by name.

The second is AI features inside tools the organisation already approved. The contract exists, so procurement is satisfied, and the AI addendum may say something quite different from the original agreement about where data goes and whether it trains anything. Nothing in the network log distinguishes an approved tool from an approved tool that grew a model.

SourceFindsBlind to
Identity providerTools holding a grant against corporate accounts; agents with broad scopesAnything used on a personal account
Network egressDomains reached from managed networks and devicesHome networks, personal devices, unmanaged browsers
FinancePaid team subscriptionsFree tiers, which are the majority
EndpointBrowser extensions and installed clientsWeb use in an unmanaged browser profile
Fig. 01 — No source sees everything. Stating how much personal-account activity remains invisible is part of the deliverable, not a caveat to bury.

Why engineering is traced first

Source code is the data type most frequently submitted to unapproved AI tools, ahead of regulated data and credentials. That single fact reorders the assessment: repositories and engineering teams get traced before the departments a security team would normally start with.

It also changes the conversation with the business. Engineers pasting code into an assistant are usually doing it to work faster on something the organisation asked for, which makes a punitive response both unfair and counterproductive.

How do you get people to tell you?

By removing the reason not to. Survey evidence consistently puts the proportion of staff who would not admit unapproved AI use to a security team at around half, which means interviews conducted under implied threat return a picture that is confidently wrong.

A time-bound, non-disciplinary self-report window works better, and it only works if the sponsor commits in writing that nothing disclosed during it will be used for discipline. We obtain that commitment before the first interview, and we schedule the interviews after the telemetry, so the questions are concrete rather than general.

What the finding should say

A register of tools with an account-type column, because the same tool on a personal free tier and on an enterprise agreement are different findings with different contract postures. Alongside it, the confidence bound: how much of the activity none of the four sources can see. That sentence is uncomfortable to write and it is the one that makes the rest of the report credible.

Article

Published 2 August 2026

By Agile Labs

Agile Labs is a Singapore enterprise software engineering company. We design, build and secure enterprise software and AI systems.

Sources

  1. Microsoft, Defender for Cloud Apps generative-AI catalogue and Entra network discovery documentation.
  2. Microsoft, SharePoint Advanced Management and Purview oversharing guidance, 2025–2026.
  3. Agile Labs AI Exposure Assessment delivery protocol, September 2026.

Related articles

View all insights

Have something complex to build, fix or take over?

Build better software, with zero surprises