Shadow AI in Engineering Teams: How to Build an AI Tool Inventory Without Surveillance
Shadow AI is not a future risk. It is the AI your teams are already using this week without a record of it. A developer pays for a coding agent on a personal card. A support lead wires a workflow tool to a language model over lunch. A data team runs an open-source agent on a spare machine with an API key someone generated months ago and never mapped to anything. None of it is malicious. All of it is invisible to the people who will be asked, sooner or later, to account for it.
The instinct is to treat this as a security problem and reach for blocking or monitoring. Both make it worse. The engineering leaders who get ahead of shadow AI treat it as an inventory problem first: what is in use, where, on whose account, at what cost, and under which controls. That inventory happens to be the first thing every AI governance framework asks for, and it can be built without reading a single prompt.
What shadow AI actually is in an engineering organization
The narrow definition is an employee pasting company data into a public chatbot. That was the 2023 version. In an engineering organization today, shadow AI is any AI system, tool, or model used for work without being registered, approved, or attributed to a team and a budget. The category has widened because the tools have.
It includes coding assistants inside the IDE and agents that open pull requests on their own. It includes general assistants used for analysis, writing, and research. It includes workflow platforms where a model call sits inside a business process, so the AI never appears in a chat window at all. And it increasingly includes open-source agents that people install on their own machines and point at a provider key, which means the organization is paying for tokens it cannot see under a licence that costs nothing.
The scale is not marginal. Microsoft's Work Trend Index found that 78% of AI users bring their own AI tools to work, most often without telling anyone. IBM's Cost of a Data Breach Report 2025 measured what that costs when it goes wrong: one in five breaches involved shadow AI, those breaches ran about US$670,000 more than average, 63% of the organizations breached had no AI governance policy at all, and 97% of the AI-related breaches lacked proper access controls. The gap those numbers describe is not one of intent. It is a gap in the record.
Why the usual response makes it worse
The first reflex is to block. Blocking a tool at the network does not remove the demand for it; it moves the usage to a personal device and a personal card, where the organization has even less visibility than before. The shadow gets deeper, not smaller.
The second reflex is to watch. Capturing prompts, logging screens, reading what people asked the model. This fails on three counts. It is disproportionate to the actual risk, which lives in what was pasted in and what was shipped out, not in the phrasing of a question. It destroys the trust that makes people willing to register a tool in the first place. And it creates a new pile of sensitive content the organization now has to protect, which is exactly the exposure governance was supposed to reduce.
There is a better boundary, and the organization already owns it: accounts, seats, licences, billing, and the administrative data providers expose to an organization about its own usage. Measuring at that boundary tells you who has a seat, what it costs, how much it is used, and what it produced. It never requires the content of a conversation. A working AI governance definition rests on evidence, enforcement, and outcomes. None of those three needs surveillance.
An inventory is the first control every framework asks for
Frameworks disagree on many things. They agree on where to start.
Under the EU AI Act, the AI literacy duty in Article 4 has applied since February 2025 to providers and deployers alike: organizations must take measures so that the people operating AI systems on their behalf understand them. You cannot train people on systems you cannot list. Article 26 sets the deployer obligations for high-risk systems, including use according to the provider's instructions, competent human oversight, and keeping the system's logs for at least six months. Every one of those duties presupposes that you know which systems you deploy and which of them fall in scope. The application date for most high-risk obligations has been pushed back to late 2027, but the inventory that classification depends on does not get easier by waiting.
ISO/IEC 42001, the AI management system standard, cannot function without an overview of the AI systems in the organization, because that overview is what risk assessment, impact assessment, and control selection are performed on. The NIST AI Risk Management Framework makes the same point structurally: its first function is Map, and mapping means knowing what exists and in what context before you measure or manage anything.
So an inventory is not compliance theatre. It is the prerequisite every other control is built on. If you cannot enumerate the AI in use, you cannot classify it, cannot assign oversight to it, and cannot produce evidence about it when asked.
The four families you are looking for
Treating all AI usage as one category is one of the most common mistakes in AI governance for software engineering, and it is fatal to an inventory, because each family hides in a different place and carries a different risk.
The first family is code assistants and coding agents: assistants in the IDE, command-line agents, cloud agents that write and review code. Their risk profile is about what enters the repository and how much of it survives, which is why AI code durability deserves to be a first-class maturity signal rather than an afterthought. They usually leave a trace in a provider admin console and in the commits themselves.
The second family is general assistants and knowledge tools: chat assistants, enterprise search, document and meeting tools with a model inside. Their risk is data leaving the boundary. They are typically bought per seat, which makes them the easiest to find in billing and the easiest to under-report in usage.
The third family is workflow and task automation. Here the model call is a step inside a process, triggered by an event, with no human typing. These platforms bill per execution or per operation rather than per person, so seat-based inventories miss them entirely. They are also where an unreviewed decision can run a thousand times before anyone notices.
The fourth family is self-hosted, open-source agents. The software is free, which is precisely the problem: nothing shows up in procurement, and the real cost lands on a provider's bill under an API key that nobody has mapped to a team. An inventory that only lists licences will report this family as not existing.
Measured, declared, and the line between them
An honest inventory has two kinds of numbers, and the discipline is to never confuse them.
A measured number comes from a system of record the organization controls: a provider's administrative API, a billing export, an identity provider's log of which applications hold tokens. A declared number comes from a person: a team lead states that six people use a tool at a known price. Both are legitimate. Declared is how you start on day one, and for tools that expose no administrative data it is the only number you will ever have. Measured is what survives an audit.
The failure is not in having declared numbers. It is in letting a declared number wear a measured label, or letting a connected key on one tool make an unrelated tool look measured by association. Keep the provenance attached to every figure, at every level of aggregation, so that a board slide and a team page tell the same story about what is known and what is stated.
How to build the inventory without reading anyone's screen
Start where the money already leaves. Procurement and expense data reveal the sanctioned tools and, more usefully, the reimbursed personal subscriptions, which are the clearest signal of shadow AI there is.
Move to identity. The identity provider knows which applications your people have granted access to through single sign-on and OAuth. That list is longer than procurement's, and it costs nothing to read.
Then go to the providers themselves. Most enterprise AI tools expose an organization-level view of seats, usage, and cost to an administrator. That is the measured layer: who holds a seat, how much was consumed, what it cost, at the granularity of a workspace or a key. Connect it where it exists.
Attribute API spend to structure, not to individuals. Map keys to workspaces and workspaces to teams or business areas. The question governance needs answered is which part of the organization is spending and what it produced, never which person asked what.
Add a declared layer for everything the previous steps cannot see, and label it as such. Then publish the inventory back to the teams. People register the tools they actually use when the register is visible, useful to them, and clearly not a surveillance instrument.
Nothing in that sequence requires the content of a prompt. Seats, spend, and outcomes are enough to govern with, and they are the only things a regulator or an auditor will ask you to evidence.
A paid tool is never worth zero
One arithmetic error breaks more inventories than any missing row: a tool showing a monthly cost of zero.
It happens in two ways. A provider's administrative API reports consumption on top of a plan, so a team that stayed within its plan shows zero usage, and someone reads that as zero cost. Or an open-source tool is listed with no licence fee and treated as free, while the model tokens it burns arrive on another line of another bill.
Both are wrong for the same reason. A contracted seat is billed every month whether or not anyone touches the API. The subscription is the floor. What an administrative key measures is the overage above that floor: credits bought when the plan ran out, consumption beyond the bundle. Cost is the subscription plus the overage, and it is never zero while a seat exists or a figure was declared. This is the principle the AI portfolio in ScaleQuality is built on: the seat is the floor, the key measures the excess, and a tool the company pays for never reads as free.
A row with zero cost is a row that is wrong. Fix the row before anyone builds a budget on it.
A practical test for your inventory
If you want to know whether your organization has an inventory or an assumption, ask three questions.
Can you list every AI tool in use, by team, with who pays for it? Can you say, for each number on that list, whether it is measured or declared? Can you show which controls apply to each family, and evidence that they were applied?
If the answer to any of those is no, you have shadow AI. Not because people are hiding anything, but because nobody kept the record. That is the part engineering leaders can fix this quarter, without a new policy and without watching anyone work.
The useful version of an inventory is the one that holds up under pressure: during a budget review, a security questionnaire, an incident, or the first conversation with a regulator. If yours can answer those three questions with specifics, it is doing its job.
References
- IBM, Cost of a Data Breach Report 2025
- Microsoft, Work Trend Index 2024: AI at Work Is Here. Now Comes the Hard Part
- EU Artificial Intelligence Act, Article 4: AI literacy
- EU Artificial Intelligence Act, Article 26: Obligations of deployers of high-risk AI systems
- ISO/IEC 42001:2023, Artificial intelligence management system
- NIST AI Risk Management Framework