Conversational AI Services Leak Sensitive Data to Advertisers – Privacy Analysis of 9 Major Platforms
Key Finding
All nine surveyed conversational AI services embed third‑party advertising or tracking services (ATSes) in their web and mobile clients, and many of them transmit conversation‑derived artifacts—such as titles, prompts, screenshots, and permanent URLs—together with persistent user identifiers, creating novel privacy risks.
Scope of the Measurement
- Services examined: ChatGPT, Claude, Gemini, Grok, DeepSeek, Perplexity, Copilot, Mistral, and Meta AI.
- Clients: Web front‑ends of all nine services and Android apps of eight (all except DeepSeek mobile).
- Methodology: Combined static decompilation (Androguard) and dynamic instrumentation (Chrome DevTools, instrumented Pixel 3a, mitmproxy, Frida) to capture network traffic, storage writes, and SDK usage.
- Experiment matrix: Tested across cookie‑consent choices (ignore, reject all, accept all), subscription tiers (guest, free, paid), and privacy modes (incognito where available).
- Data: 18‑page PDF, 44 third‑party organisations, 124 distinct domains, 44 ATSes, plus responsible disclosure logs.
Third‑Party Tracking Landscape
- Universal presence: Every service contacts at least one ATS; the average is 44 distinct third‑party domains per service.
- Dominant players: Google products (Ads, Tag Manager, Analytics, Firebase) appear in 9/9 services; other frequent partners include Meta Pixel, TikTok Analytics, Sentry, Datadog, and Intercom.
- Web vs. Mobile: 11 ATSes appear on both platforms, 15 only on web (e.g., OneTrust, TikTok) and 8 only on mobile (e.g., Braze). Mobile traffic is 71 % WebView‑driven, amplifying web‑style tracking within apps.
Conversation‑Derived Artifact Leakage
| Artifact | Web leak (services) | Mobile leak (services) |
|---|---|---|
| Conversation URL / permalink | 5 of 9 | 0 |
| Conversation ID (without full URL) | 2 of 9 | 1 of 8 |
| User prompt | 1 of 9 | 1 of 8 |
| Conversation title (AI‑generated summary) | 3 of 9 | 0 |
| Screenshot (shared view) | 1 of 9 | 0 |
| Shared conversation URL | 5 of 9 | 0 |
| Shared conversation ID | 0 | 1 of 8 |
- Mechanism: URLs and titles are sent as query parameters or JSON fields to trackers such as Google Ads, Meta Pixel, TikTok, and Datadog.
- Consent impact: Accepting non‑essential cookies activates additional trackers (e.g., Meta, TikTok, DoubleClick). Rejecting cookies still leaves 44 % of services transmitting data to third parties.
- Subscription tier impact: Free vs. paid tiers show negligible differences; only Claude’s mobile client drops Intercom/Sentry in the premium tier.
Linking Artifacts to Persistent Identifiers
- Identifiers observed: Hashed email addresses (HEMs), account usernames, provider‑specific IDs, Android Advertising ID (AAID), and cookies such as
_fbp(Meta) and_ttp(TikTok). - Cross‑linking: Trackers receive both conversation artifacts and identifiers, enabling them to stitch chat content to long‑term user profiles across devices and services.
- Server‑side forwarding: Claude uses Segment Analytics to forward events server‑to‑server to eleven ad networks, bypassing ad‑blockers.
- Web‑only vs. mobile‑only:
- Web trackers receive titles and URLs; mobile SDKs mainly receive IDs and, in a few cases, email hashes.
- WebView‑based mobile traffic mirrors web leakage patterns for the same providers.
Publicly Accessible Permalinks
- Access control matrix (guest / free / paid):
- ChatGPT, Claude, Gemini, Copilot, Mistral: owner‑only by default (optional opt‑out).
- Grok: public by default with opt‑out; Perplexity: always public for guest tier.
- Implication: Anyone possessing a permalink can retrieve the full conversation without authentication, exposing potentially sensitive health, financial, or personal data.
Real‑World Retrieval Evidence
- Canary‑token experiment: Embedded unique URLs in prompts and uploaded documents.
- Results:
- Grok triggered 70 distinct accesses from 48 ASes across 14 countries (65 % from the USA) over days.
- Perplexity, Copilot, Mistral, and Claude showed single‑shot accesses from cloud providers.
- Interpretation: Third parties actively retrieve shared conversation resources, confirming that leakage is not merely theoretical.
Legal Assessment (EU GDPR & ePrivacy)
- ePrivacy Directive: Article 5(3) requires prior informed consent for cookies, tracking pixels, and URL‑based tracking. Many services transmit tracking data even when users reject non‑essential cookies, breaching this requirement.
- GDPR: Processing of conversation content and identifiers constitutes personal data. Providers often lack transparent disclosures and a lawful basis beyond “performance of contract.” The CJEU (Meta Platforms Ireland) mandates explicit information on data categories, purposes, and legal bases—requirements many services do not meet.
- Special‑category data risk: Health‑related prompts (used in the study) may trigger GDPR special‑category processing, yet no explicit safeguards are evident.
Discussion of Impact
- Privacy risk amplification: Conversational AI adds a layer of rich, intent‑bearing data (titles, prompts) to the existing ad‑tech ecosystem, enabling finer‑grained profiling.
- Ineffective user controls: Cookie‑consent banners and tiered subscriptions provide limited mitigation; 80 % of trackers remain active after cookie rejection.
- Broader ecosystem: Findings likely extend to custom chatbots, LLM‑powered web widgets, and enterprise‑grade agents that reuse the same SDKs.
- Potential for abuse: Public permalinks combined with tracker access could be weaponized for targeted phishing, blackmail, or surveillance.
Mitigation Recommendations
- Default‑deny sharing: Providers should make conversation permalinks private by default and require explicit user action to generate a shareable link.
- Strip artifacts before transmission: Remove titles, prompts, and URLs from payloads sent to third‑party ATSes unless a user explicitly opts‑in.
- Separate ad‑tech from AI pipelines: Deploy dedicated ad‑tech containers that do not receive AI‑generated content.
- Strengthen consent UI: Implement granular, pre‑ticked‑box‑free consent dialogs that block all non‑essential trackers unless accepted.
- Regulatory oversight: Data Protection Authorities should audit the disclosed data flows against GDPR/ePrivacy requirements, especially for health‑related prompts.
- User‑side tools: Encourage use of privacy‑focused browsers (Brave, Tor) and mobile privacy dashboards; consider sandboxing AI agents (e.g., cellmate) to limit SDK privileges.
Community Reactions (Hacker News Highlights)
"multiple providers disclose sensitive conversation‑derived artifacts — including titles, prompts, and screenshots — to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation." – Coeur
"It's not a leak if it's the business model." – drywater2 (points out intentional data sale rather than accidental leakage.)
"Ad‑tech spent 20 years trying to infer intent from clickstreams. Chat apps now hand over an AI‑written one‑line summary of intent, labeled and keyed to a cookie." – vivekpolavarapu (highlights the new granularity of profiling.)
"I built my own chat interface to avoid this." – 0xcrypto (suggests self‑hosted alternatives.)
Limitations
- Scope: Only nine consumer‑facing services; enterprise tiers and API‑only usage were not examined.
- Geography: Tests performed from Spain; regional variations may exist.
- Dynamic evasion: Some SDK behavior may be hidden under conditions not triggered in our scripted interactions.
- Server‑side opacity: Cannot fully observe backend data flows that do not manifest in network traffic.
Conclusion
The study demonstrates that mainstream conversational AI platforms have inherited the advertising and tracking ecosystem of the web and mobile, extending it with conversation‑specific data that can be linked to persistent identifiers. This creates a novel privacy attack surface that conflicts with EU data‑protection law and undermines user expectations of confidentiality. Stronger technical safeguards, transparent disclosures, and regulatory scrutiny are essential to protect users from inadvertent exposure of their most personal conversations.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch