Benchmark

What a real agent estate
actually looks like.

Anonymised patterns from full-estate assessments. Every figure is measured from a connected tenant, not modelled, not surveyed, and not extrapolated. The sample is small and stated - we would rather publish one honest estate than a projection.

On sample size. These figures come from full-estate assessments at a small number of organisations, and the detailed breakdown below is a single mid-market tenant of roughly two thousand seats with 625 agents. That is not a market study and we will not present it as one. It is what happens when somebody counts properly for the first time, and the patterns have held everywhere we have looked so far.

The count nobody expects

625agents in one mid-market tenant
0previously inventoried
16environments, several forgotten
327built in-tenant, not vendor templates

When Microsoft rolled out its own agent governance internally, the exercise surfaced over half a million agents in a matter of weeks. Those agents already existed. The rollout made them visible for the first time. That is the pattern: agent sprawl is not sudden growth, it is delayed counting.

Pattern 1 - the count is inflated, and the real number still surprises

Of 625 agent records, 298 were deployment replicas - one definition promoted across dev, QA, staging and production. Only 33 were genuine duplicates, in 9 clusters.

A tool that counts repeated names would have reported 331 duplicates and sent a platform team chasing its own release pipeline. Separating replicas from copies took the register from 625 decisions to 364 - and every one of those 364 is a real decision somebody has to make.

Pattern 1, in one picture

The raw count is not the number of decisions.

One agent promoted through dev, QA, staging and production is four records and one governance decision. Separating replicas from genuine copies is the difference between a board paper people can act on and one they discount.

625 agent records what the platform returns 298 deployment replicas one agent, four environments 33 genuine duplicates built twice, in 9 clusters 364 real decisions what somebody has to review A name-matching tool would have reported 331 duplicates - and sent a platform team to fix a working release pipeline.

Compare operations, not names. And make the comparison order-independent - platforms return connector operations in whatever order they like, so a naive hash reports eight copies of one agent as eight different agents.

Pattern 2 - nobody is accountable, anywhere

625of 625 with no signed purpose
94of 94 acting agents with no business or security owner
274with no directory machine identity

This is the most consistent finding and the least surprising once you look at the platform. Copilot Studio does not ask who owns an agent. It does not ask what the agent is for. It does not require a review before an agent reads production data. The fields are absent because nothing creates them - not because anyone skipped a step.

Pattern 3 - most of the estate runs an unpinnable model

79%on a vendor-managed model with no version string
28agents on a model whose version can be compared
11distinct models across the estate

"Microsoft 365 Copilot" and "Copilot Studio default" name a platform, not a model version. The underlying model can change without notice and carries nothing to pin, compare or alert on. Any supply-chain control that assumes a version string covers about a fifth of a typical Microsoft estate.

Pattern 4 - exposure is narrow, and that is good news

398agents ingest content nobody vetted
3can send data outside the tenant
1holds all three legs of an exfiltration path

Nearly two thirds of the estate takes in web content, public knowledge sources or inbound messages. But egress is a chokepoint: three agents. Most of the exposure closes by governing three agents rather than four hundred - which turns a frightening number into a morning's work.

Pattern 5 - tool reach is wider than the tool count

43 distinct tool definitions covered 190 agent-tool grants once replica clusters were expanded, and none had been approved. Fifteen of those 43 rows each apply to around eight agents - so approving one row grants a capability across eight agents, including production. An approval queue that does not show that is asking people to sign for something they cannot see.

What we do not claim

This is not a survey

No self-reporting, no vendor panel, no modelled extrapolation. Every figure is read from a connected tenant. That makes it accurate and narrow, and we would rather be both than neither.

Your estate will differ

A 200-seat professional services firm and a 20,000-seat manufacturer will not look like this. The patterns - absent ownership, absent purpose, narrow egress, inflated counts - have held everywhere we have looked. The magnitudes will not.

Figures anonymised and used with permission. The detailed breakdown is a single tenant; where a pattern has appeared across more than one assessment it is described as a pattern rather than a statistic.

Add your estate to the benchmark.

Book a free 30-minute session

One read-only administrator consent. Nothing installed, no traffic proxied. First findings inside 24 hours.