How to detect shadow AI
Most shadow AI can be found in records you already keep: DNS, proxy, firewall and sign-in logs, OAuth consents and card statements. This is the method, the queries to run, and what each source misses.
The method in five steps
- Write down what is approved. List the AI tools you have sanctioned, with the account type (company workspace or personal). Everything else you find is shadow AI by definition.
- Pull two to four weeks of records from at least one network source and one identity source in the table below. One network log plus sign-in data catches most of it.
- Match them against a catalogue of AI domains and product names. Run the queries below where the data lives, or export it and drop it into the shadow AI detector, which explains every match.
- Triage what you find by data-risk tier and by how many people use each app.
- Repeat monthly and update the approved list as decisions are made.
Which data sources show what
| Source | Shows | Misses |
|---|---|---|
| DNS resolver logs (Cloudflare Gateway, Windows DNS, Pi-hole, Zeek) | Every AI domain looked up, by device or user | Devices using encrypted DNS to an outside resolver; anything off the network |
| Web proxy or secure web gateway (Zscaler, Squid, W3C extended logs) | Host and often the signed-in user, with request counts | Traffic that bypasses the proxy, such as remote staff without the agent |
| Firewall URL filtering (Palo Alto Networks) | Host or URL, source user and address | Encrypted traffic where only an IP address is logged |
| Endpoint network telemetry (Microsoft Defender DeviceNetworkEvents) | Connections from managed laptops wherever they are, with the process and account | Unmanaged and personal devices |
| Single sign-on logs (Microsoft Entra ID SigninLogs) | AI apps people sign in to with a work account | Apps used with a personal email and password |
| OAuth consent logs (Google Workspace) | AI apps given access to mail, files or calendars | Apps that never ask for access |
| Card and expense exports | Paid AI subscriptions, including ones used off the network | Free tiers, which is most of it |
| SaaS inventory or CASB export | The apps your other tools already discovered | Whatever those tools don't cover |
The UK National Cyber Security Centre recommends the same mix for shadow IT in general: asset management, network scanning, cloud access security brokers and endpoint management (UK National Cyber Security Centre, Shadow IT guidance, published 27 July 2023, reviewed 14 August 2026).
Queries to run where the data lives
Both queries are generated from the detector's catalogue (122 domains, 42 product names), so they match what the detector flags.
Microsoft Defender advanced hunting (endpoint connections)
// Microsoft Defender advanced hunting: AI services reached from managed devices in the last 30 days. let aiDomains = dynamic(["aistudio.google.com", "api.anthropic.com", "api.deepseek.com", "api.mistral.ai", "api.openai.com", "api.x.ai", "avoma.com", "bard.google.com", "beautiful.ai", "bolt.new", "c.ai", "character.ai", "chat.deepseek.com", "chat.mistral.ai", "chat.openai.com", "chat.qwen.ai", "chatgpt.com", "chatpdf.com", "claude.ai", "codeium.com", "cohere.ai", "cohere.com", "consensus.app", "console.anthropic.com", "console.mistral.ai", "console.x.ai", "copilot-proxy.githubusercontent.com", "copilot.cloud.microsoft", "copilot.microsoft.com", "copy.ai", "cursor.com", "cursor.sh", "cursorapi.com", "deepl.com", "deepseek.com", "devin.ai", "doubao.com", "dreamstudio.ai", "elevenlabs.io", "elicit.com", "elicit.org", "fathom.video", "fireflies.ai", "firefly.adobe.com", "fireworks.ai", "gamma.app", "gemini.google.com", "generativelanguage.googleapis.com", "genspark.ai", "githubcopilot.com", "grain.com", "grammarly.com", "grammarly.io", "granola.ai", "granola.so", "grok.com", "groq.com", "groqcloud.com", "heygen.com", "hf.co", "hf.space", "huggingface.co", "humata.ai", "ideogram.ai", "jasper.ai", "kimi.ai", "kimi.com", "kimi.moonshot.cn", "krisp.ai", "leonardo.ai", "lindy.ai", "lovable.app", "lovable.dev", "makersuite.google.com", "manus.im", "meta.ai", "midjourney.com", "notebooklm.google", "notebooklm.google.com", "oaistatic.com", "oaiusercontent.com", "openrouter.ai", "origin-tracker.githubusercontent.com", "otter.ai", "perplexity.ai", "phind.com", "pi.ai", "pika.art", "platform.deepseek.com", "platform.openai.com", "poe.com", "pplx.ai", "qianwen.aliyun.com", "quillbot.com", "read.ai", "relevanceai.com", "repl.co", "replicate.com", "replicate.delivery", "replit.com", "replit.dev", "runwayml.com", "sora.com", "stability.ai", "suno.ai", "suno.com", "sydney.bing.com", "synthesia.io", "tabnine.com", "tldv.io", "together.ai", "together.xyz", "tome.app", "tongyi.aliyun.com", "udio.com", "v0.app", "v0.dev", "windsurf.ai", "windsurf.com", "wordtune.com", "writesonic.com", "you.com"]); DeviceNetworkEvents | where Timestamp > ago(30d) | where isnotempty(RemoteUrl) and RemoteUrl has_any (aiDomains) | summarize Connections = count(), Devices = dcount(DeviceName), Users = dcount(InitiatingProcessAccountUpn), FirstSeen = min(Timestamp), LastSeen = max(Timestamp) by RemoteUrl | order by Devices desc
Microsoft Sentinel or Log Analytics (Entra ID sign-ins)
// Microsoft Sentinel or Log Analytics: sign-ins to apps whose name contains an AI product name, last 30 days. SigninLogs | where TimeGenerated > ago(30d) | where AppDisplayName has_any (dynamic(["anthropic", "character.ai", "chatgpt", "claude", "codeium", "copilot for microsoft", "deepl", "deepseek", "elevenlabs", "fireflies", "gemini", "genspark", "github copilot", "grammarly", "grok", "heygen", "hugging face", "huggingchat", "huggingface", "ideogram", "krisp", "lovable", "manus", "microsoft copilot", "midjourney", "mistral", "notebooklm", "openai", "openrouter", "otter ai", "otter.ai", "perplexity", "quillbot", "replit", "suno", "synthesia", "tabnine", "tl;dv", "tldv", "windsurf", "wordtune", "writesonic"])) | summarize SignIns = count(), Users = dcount(UserPrincipalName), FirstSeen = min(TimeGenerated), LastSeen = max(TimeGenerated) by AppDisplayName | order by Users desc
For anything else, export the records as CSV, JSON lines or a raw log and drop them into the detector. It recognises 12 export formats by their columns and reads any other log for host names. The domain list is also available as CSV for DNS filters and SIEM lookups.
Why both domains and product names
A domain match proves traffic reached a service; each catalogue domain belongs to one AI product, so api.anthropic.com can only mean the Anthropic API. But sign-in logs, OAuth consents and card statements carry an app's name, not its domain. Matching distinctive names there catches AI used off the network. Names are less certain than domains, which is why the detector shows the exact text it matched so a person can confirm it.
Triage what you find
| What you see | Do this first |
|---|---|
| Model API traffic from a server or build machine | Find the script or agent, move it to a company key, give it an owner |
| An agent platform (Critical tier) | Find out what accounts it was connected to and what it can do |
| One chat assistant used by many people | Approve a business plan for it: people clearly need it |
| A meeting recorder | Check consent and where transcripts are stored; pick one approved recorder |
| A tool used once by one person | Ask them; often no action is needed |
Make detection routine
The NIST AI Risk Management Framework expects "mechanisms are in place to inventory AI systems" (GOVERN 1.6, NIST AI 100-1, AI Risk Management Framework 1.0, January 2023). In IBM's 2025 breach study, only 37% of organisations had policies to manage AI or detect shadow AI (IBM, Cost of a Data Breach Report 2025 (press release), 30 July 2025). A monthly export and a ten-minute review is enough to stay in the first group. To keep the inventory current between exports, Agent Trust Cloud discovers the AI agents and machine identities running against your systems.
Blind spots to accept
- Personal phones and home networks never reach company logs.
- AI features inside approved software use that software's domains.
- Browser extensions may call AI services from domains of their own.
- Logs show where traffic went, not what was typed.
Related free tools
- Non-Human Identity Audit: inventory the service accounts, keys and AI agents that the logs point to.
- Agent Permission Audit: compare what an AI agent was granted with what it needs.
- AI Policy Generator: write the acceptable-use policy that the approved list belongs in.
Run the shadow AI detector · What is shadow AI? · Shadow AI risks
Sources
- NIST AI 100-1, AI Risk Management Framework 1.0, January 2023
- UK National Cyber Security Centre, Shadow IT guidance, published 27 July 2023, reviewed 14 August 2026
- IBM, Cost of a Data Breach Report 2025 (press release), 30 July 2025
- Log formats: Cloudflare Gateway DNS log; Cloudflare Gateway HTTP log; Microsoft Defender for Endpoint network events (DeviceNetworkEvents); Palo Alto Networks URL filtering log; Zscaler Internet Access web log; Google Workspace OAuth log events; Microsoft Entra ID sign-in log (SigninLogs or Graph signIn); Zeek dns, http or ssl log; W3C extended log (#Fields: header); Squid native access.log.
- Each figure checked on the source page, 1 October 2026.