Researchers at IMDEA Networks tested nine conversational AI services and found that every one of them contacts at least one advertising or tracking company, according to a paper that Wolfie Christl, a senior researcher at Cracked Labs and The Citizen Lab, flagged on LinkedIn today. Six of the nine web clients and three of the eight Android apps handed conversation URLs, auto-generated conversation names, prompts or screenshots to outside firms, often next to persistent user identifiers, according to the authors.
In Short
Researchers checked nine popular AI assistants and found that all of them send some data to advertising or analytics companies, and for several of them that data hints at what a conversation was about. It matters to anyone who types personal questions into an assistant, because the data can arrive next to identifiers, such as a hashed email address, that tie it to one person, and turning cookies off only cut the flow in part. What changes is the evidence base: privacy regulators and advertisers now have a service-by-service measurement, and Spain's data protection authority has already asked for EU-level review.
What the researchers tested
The paper is titled "Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents". Nine authors are listed, most at IMDEA Networks; one is also affiliated with UC3M and one is listed as independent. The PDF reviewed is formatted for Proceedings on Privacy Enhancing Technologies but carries a placeholder publication year and no date, and its latest dated entry is September 10, 2026. The document therefore does not say when it was released. Christl's post, which showed a timestamp of seven hours when captured, is the basis for placing its circulation on October 5.
The nine services are ChatGPT from OpenAI, Claude from Anthropic, Grok from xAI, DeepSeek, Perplexity, Gemini from Google, Microsoft Copilot, Mistral's Le Chat and Meta AI. The authors picked them by popularity proxies. On the Tranco web ranking, ChatGPT sits at 48, Claude at 617 and Grok at 956, with Meta AI lowest of the group at 13,701. Cumulative Google Play installs exceed 1 billion for ChatGPT and Gemini, and every app has at least 1 million.
All experiments ran in Spain during May 2026. Web sessions used Chrome v148.0.7778.167 with developer tools; Android sessions used an instrumented Pixel 3a running Android 12, complemented by static analysis of each app with Androguard. The team varied the cookie choice (ignore the banner, reject non-essential cookies, accept all), the account tier (guest, free, paid) and whether a conversation was shared. Prompts were health-related, modelled on a user asking about a medical condition, an approach the authors call persona-based auditing. Pilot runs showed deterministic behaviour, so each configuration was analysed once, and one co-author labelled the third-party domains by hand using what the paper calls conservative criteria. Advertising, analytics and telemetry endpoints owned by Google, Microsoft or Meta were counted as third-party services even where the AI product belongs to the same company.
Every service contacts an advertising or analytics firm
The measurements turned up 124 distinct third-party domains, which the authors attribute to 44 organisations; 34 of those are classed as advertising and tracking services. Eleven appear on both web and mobile clients, 15 only on the web (Google Tag Manager, TikTok and the consent platform OneTrust among them) and eight only on mobile, Braze included.
Google is the most widespread. Its Ads and Tag Manager products appear in the web clients of seven services, and Firebase turns up in all eight Android apps. Sentry, Meta, Datadog and Intercom follow.
The set also includes lesser-known names. DeepSeek's clients contact Fengkong Cloud on mobile and ShuMei on the web, which the paper says are rarely documented in academic literature and are absent from the WhoTracks.Me database; public reports, according to the paper, link them to device-fingerprinting and risk-scoring technologies.
In-app browsing matters on mobile. Some 71% of the distinct endpoints that mobile clients contact originate from WebViews rather than native code, ranging from 98% in Grok and 91% in Perplexity to 71% in ChatGPT, while most traffic from Copilot, Mistral and Meta AI comes from native code. Inside Grok's WebViews the team saw content from Google Ads, Google Tag Manager, TikTok Analytics, X/Twitter Analytics and the Meta Pixel. Separately, the authors found browser-fingerprinting APIs invoked in scripts from every provider; they stress that this shows capability, not proof of active fingerprinting.
What reaches third parties
The distinctive risk, in the authors' framing, is that conversational products generate artifacts that describe the interaction itself. Table 3 of the paper counts what leaked, across web and Android clients, in the course of ordinary use.
| Artifact | Web providers | Android providers | Web ad/tracking services | Android ad/tracking services |
|---|---|---|---|---|
| Conversation URL | 5 | 0 | 9 | 0 |
| Conversation ID only | 2 | 2 | 1 | 1 |
| User prompt | 1 | 1 | 2 | 1 |
| Conversation name | 3 | 0 | 9 | 0 |
| Screenshot | 1 | 0 | 1 | 0 |
Several table legends render ambiguously in the copy reviewed, so counts here follow the body text and the numeric tables.
Addresses and names
Five web clients disclose conversation URLs, or the identifiers behind them, to nine services. Three of the five do so by default; the remainder only after non-essential cookies are accepted. ChatGPT and Claude on the web also send the bare conversation ID to Datadog, which, the authors note, lets a recipient rebuild the public URL. No URL disclosure was seen on mobile.
Three web clients leaked the auto-generated conversation name: Gemini via Google Analytics, and Grok and Mistral through several services. Recipients included Meta, TikTok and Doubleclick, and eight of the nine recipient relationships occurred only after cookie acceptance. The researchers' own examples show why that matters. A prompt asking about the symptoms of early-stage Parkinson's disease became "Early-stage Parkinson's Symptoms" in ChatGPT and "Early Parkinson's Disease Symptoms" in Grok. A prompt giving an $85k salary and asking about a mortgage in New York became, in Grok, "$85k NYC Salary: $280k-$350k Mortgage".
One Grok conversation reached seven advertising and analytics services, according to the paper's Figure 7. The Meta Pixel alone collected the conversation ID, the generated name, the full URL, the share ID and the share URL, each tied to the user's Meta identity through the synced _fbp cookie.
Identifiers that travel with it
Conversation data matters more when it lands beside an identifier. Of the web identifier transmissions the team observed, 77.3% occurred only after a user accepted non-essential cookies. Third-party cookies appeared in four services, from nine third parties including Meta, TikTok, Twitter Analytics and Intercom; eight distinct cookies, such as _fbp, _ttp and _twpid, travelled alongside conversational artifacts.
Hashed email addresses feature too. According to the paper, Perplexity transmits users' hashed emails to the marketing and analytics company Singular. The authors cite the US Federal Trade Commission for the point that such hashes act as stable identifiers, a position PPC Land covered in its report on the FTC's warning that hashed data is not anonymous. Stable hash-based IDs also went to Intercom from Mistral and Claude, and from Grok to DoubleClick, Google Ads and Google Search.
Server-side forwarding
The team also observed cookie syncing and server-side tracking, both only after cookie acceptance. In Claude's case, the web client runs Segment Analytics through a first-party domain, a-cdn.anthropic.com, and loads a Conversions APIconfiguration that forwards user events from server to server to eleven services, among them Facebook, LinkedIn, TikTok, Reddit and Google Enhanced Conversions. According to the authors, this forwarding evades ad blockers and carries two shared IDs per event.
Grok embeds a Google Tag Manager container set to route events through a server-side instance. One custom event sends the conversation URL and chat topic to Meta's Conversions API and TikTok's Events API, with the _fbp and _ttp cookies in the same payload, which the paper says allows the two identities to be bridged. In the authors' words, these flows are invisible to browsers and unblockable by ad blockers.
Mobile identifiers
On Android, four clients disclose account-linked identifiers. Perplexity sends an email address to RevenueCat; Grok sends emails to TikTok and Twitter Analytics, and hashed emails to the same two. Three apps transmit the Android mobile advertising IDs: to Adjust from Copilot, to AppsFlyer from Grok and to Singular from Perplexity. Grok and Perplexity send those IDs together with user or installation IDs, a pairing the authors say defeats the point of a resettable identifier, since a new advertising ID can be re-bound to the old profile.
Claude's app sends geolocation coordinates and the Android ID to Sift Science, and Meta AI sends the Android ID to its own servers. The paper notes that such attributes are often collected for fraud prevention or analytics.
Consent and subscription tiers do little
Consent design differs by provider. Perplexity, Claude and Grok allow use without a banner decision, whereas Gemini and Meta AI require one, and Mistral requires acceptance of its terms and privacy policy before the service can be used.
Rejecting non-essential cookies did not end third-party contact. Perplexity, DeepSeek, Gemini, Copilot, ChatGPT and Claude all still connected to Google Ads in the reject-all scenario. Accepting everything switched on further services in Claude, Perplexity and Grok, mainly Meta, TikTok, Twitter Ads, DoubleClick and AppsFlyer. Ignoring the banner produced the same contacts as rejecting. In Claude, rejection prevented activation of the Meta Pixel, Datadog telemetry and the server-side forwarding to eleven advertising platforms.
Free and premium accounts showed nearly identical third-party sets. The one exception was Claude's mobile client, which contacted Intercom and Sentry on the free tier but not the premium one, a difference the authors attribute possibly to dynamically activated code paths.
Two figures in the paper describe how much survives rejection, and they are not reconciled. Section 6.1 says third-party services still collect data in 44.4% of free-tier services (four of nine) even when non-essential cookies are rejected. The Discussion says 80.8% of third-party services remain active in that scenario. The denominators differ, services in one case and individual third parties in the other, but the paper does not say so.
The discussion adds a structural point. On the web, only ChatGPT's paid tier supports persistent consent preferences, yet roughly 95% of ChatGPT users sit on free tiers, according to a Reuters report the authors cite, and OpenAI has announced advertising for free and low-cost users. Android apps offer no equivalent of a web cookie banner at all.
Shared conversations and public permalinks
A conversation address is only as sensitive as the access controls behind it. Testing each URL in a private browsing session without authentication, the team found that every service exposes shared conversations without login, and nine third parties were present on the sharing pages.
Defaults vary. Grok's free and premium conversation permalinks are accessible by default, with an opt-out. Perplexity always makes guest-tier conversations public, though on April 3, 2026 it stopped sending the URL to firms such as Meta, a change the authors link, with qualification, to the class action filed against Perplexity on March 31; that case was voluntarily dismissed on May 1 without a ruling on the merits.
Grok's sharing function showed the most data movement. When a shared conversation is opened, TikTok collects a screenshot of the conversation, while Meta and TikTok both collect the user's latest prompt. The example screenshot in the appendix shows a prompt asking for help interpreting a biopsy result. Claude's mobile app, for its part, sent Datadog a span for a conversation-share action that carried organisation, conversation and share UUIDs alongside account identifiers; the figure caption says the share resource is created without explicit user interaction.
To test whether disclosed resources are later fetched, the researchers planted canary URLs in prompts and uploaded documents. Only Grok showed repeated access: 70 activations over hours to days, from 70 IP addresses across 48 autonomous systems in 14 countries, 65.7% of them in the United States although the conversations took place in the EU. For DeepSeek, Copilot, Mistral and Claude, some canary URLs were activated once, at submission, from cloud providers such as Amazon Web Services and Google Cloud Platform. Perplexity repeatedly fetched canary URLs even when told not to, from its Perplexity-User crawler on AWS infrastructure in the United States. Absence of observed access, the authors add, is not evidence that no access occurs.
The legal reading
The authors are careful about scope. They write that their "objective is not to provide a definitive legal assessment or determine compliance", and instead identify provisions of the ePrivacy Directive and the GDPR that appear relevant.
On the ePrivacy side, Article 5(3) requires prior informed consent for cookies used for analytics, advertising or profiling. According to the paper, the European Data Protection Board's Guidelines 2/2023 extend the rule to URL and pixel tracking, to collecting unique persistent IDs and to client-side code that instructs the browser to send information to a server-side API.
Under the GDPR, the paper argues, both the disclosure to a third party and that party's later access are processing activities needing their own transparency and legal basis. It points to the terms used in provider privacy materials, "user content" at OpenAI, "conversations" at Anthropic, "service interaction info" at Perplexity and "user info" at xAI, and says the user's legitimate expectations would hardly include conversations being accessible to Meta or Google. It also calls it "highly debatable" whether training a service on user interactions is objectively necessary for providing it, noting the assurances major providers give about not training on enterprise-tier conversations. Because users commonly ask about health and other intimate matters, the authors add, special-category data rules come into play.
On pseudonymisation arguments drawn from recent Court of Justice case law, the authors say large platforms such as Meta or Google are unlikely to benefit, given their vast personal databases. They also cite the Court's Fashion ID judgment for the principle that enabling third-party access is itself a processing decision, whether or not the data is read.
Disclosure and responses
The disclosure timeline in the paper starts on March 23, 2026, when the team first noticed third-party activity in Perplexity and Grok traffic. Systematic testing began on April 6. On April 13, the researchers notified data protection authorities in the EU and the UK, and on April 17 they told xAI of what they describe as Grok's lack of access controls on conversation resources, the one issue the paper treats as exploitable by an outside party. Most other findings, according to the authors, "reflect intentional product-level practices rather than exploitable security vulnerabilities".
Part of the findings went public on May 4 on a project site called LeakyLM. On May 27, the Spanish authority AEPD cited the work in asking for the investigations to be raised to a plenary meeting of the European Data Protection Board planned for June 6, 2026. The paper does not report what the meeting decided.
Provider reactions are thinly recorded. OpenAI updated ChatGPT's privacy policy on August 15, 2026, explicitly mentioning third-party tracking, though the authors say they cannot confirm that their work influenced it. On September 10, Grok still used publicly accessible permalinks, and the paper states that "No official response has been received to date since our responsible disclosure." The timeline records no statements from the other providers.
Limits of the study
The authors call their results "a point-in-time lower bound". The study covers nine services, excludes enterprise and government tiers, and does not capture risks from long-lived conversations with persistent memory. Gemini's mobile traces could not be extracted. That limitation sits awkwardly with other parts of the paper: the introduction counts Android clients of the eight providers that offer one and reports 3 of 8, yet Table 1 lists Android package names for all nine services, Gemini's included.
The presence of a third party, the authors add, does not by itself show that data is sent only for advertising or tracking, since some software kits support attribution or in-app payments. Large providers can also track through their own first-party domains in ways the method cannot separate from functional traffic. The authors state they used generative AI tools to revise text and create LaTeX table structures, but not in the research methodology or analysis.
The LiveRamp connection
Christl's post links the paper to a second item. "Also, somehow related," he wrote, OpenAI's new integration with LiveRamp "enables profiling and targeting using population-scale identity records" linking names, postal, email, phone and digital IDs. He also pointed to his 2024 LiveRamp investigation. The IMDEA paper itself does not test LiveRamp or any ChatGPT audience feature; its only references to OpenAI's advertising are a Reuters-reported pilot with Criteo, the plan to show ads to free-tier users and a recent update to OpenAI's marketing-cookie controls.
The integration is the subject of a LiveRamp blog post dated September 28, 2026, by Travis Clinger. According to LiveRamp, it "is now a technology partner of ChatGPT Ads", so marketers can activate first-party audiences in ChatGPT Ads using RampID, drawing on CRM, loyalty, behavioural, web, app or other first-party data. LiveRamp lists the uses as targeting precision, reach, suppression of consumers and measurement. It describes RampID as "the most durable identifier for connecting data in the ecosystem", a vendor characterisation, and says the integration is live in 11 major markets without naming them. Allegiance Group & Pursuant is implementing it for a nonprofit client; its Megan Morris is quoted as saying the agency wants to "view the omnichannel impact of our ads across all of our investments". PPC Land covered the announcement the same day in its report on customer lists reaching ChatGPT ads via LiveRamp in 11 markets, which also notes that LiveRamp disputes critics' characterisations and points to contractual and technical safeguards. It followed LiveRamp's June 10 Conversions API Hub connection to ChatGPT ads.
What Christl refers to is a February 2024 report by Cracked Labs, written by him and Alan Toner and commissioned by Open Rights Group. It was based on research in September 2023 and on LiveRamp's own software documentation. According to the report, LiveRamp sells data about 700 million consumers worldwide from 150 data providers through its data marketplace, holds identity records on 45 million people in the UK and 25 million in France, and claims data on 14 billion devices. The report calls the company's identity graph systems "private population registers", and says that "Each time a company utilizes a RampID to link and match personal data, it processes a pseudonymous identifier that is tied to a person's partial or full identity record maintained by LiveRamp". Open Rights Group filed complaints with the UK Information Commissioner's Office and France's CNIL on February 28, 2024.
The report carries its own caveat: it relies on public documentation that "might be ambiguous and incomplete", and the authors say it remains largely unclear how clients implement the systems. LiveRamp is also changing hands. Publicis Groupe agreed in May 2026 to buy the company for $2.5 billion, as PPC Land's account of the SEC filing describes.
Why this matters for the marketing community
The channel at issue is young and growing fast. OpenAI began its ChatGPT ad pilot in the United States on February 9, 2026, and Criteo became its first ad tech partner on March 2. By the end of August, OpenAI said advertising inside ChatGPT had reached a $1 billion annualised run rate. In Europe, OpenAI told users on August 15 that ads would begin on the Free and Go plans later that month and that its privacy policy was being updated, as PPC Land reported in its piece on ads reaching European users.
The IMDEA paper widens an earlier line of evidence. PPC Land's report on a UC Davis study of 20 chatbots found that 17 shared data with at least one third party and that three leaked plaintext prompts through Microsoft Clarity session replay; the IMDEA authors cite that work and add mobile apps, consent states, subscription tiers and access controls. Litigation has run in parallel in the United States, from the Perplexity complaint to a May class action over ChatGPT.comalleging that queries reached Meta and Google. California's SB 690 narrows one route for such private suits, though the interception theory behind the AI complaints remains available.
Two distinctions keep the findings in proportion. The third parties in the paper sit on the AI services' own sites and apps, not inside ads that advertisers buy. The paper does not claim that conversation content flows into ChatGPT audiences, and PPC Land's coverage of a researcher's findings on a one-year OpenAI cookie records that the researcher has not shown it used for off-site tracking and that OpenAI says advertisers cannot access personal data or conversation history.
Yet the vocabulary overlaps. Meta Pixel, TikTok events, Google Tag Manager and server-side conversion pipes are the tools that performance marketers run on their own properties, and the paper shows them operating on conversational interfaces, with consent controls that, on the authors' numbers, reduce but rarely end third-party contact. Whether the identity systems now entering ChatGPT advertising will ever sit alongside the kind of conversational signals the paper documents is a question the paper leaves open, and neither OpenAI nor LiveRamp has addressed it in the materials reviewed. What regulators do with the AEPD's request is the other open item.
Timeline
- February 28, 2024 - Open Rights Group files complaints against LiveRamp with the UK ICO and France's CNIL; the Cracked Labs report is dated the same month
- February 9, 2026 - OpenAI formally begins testing ads in ChatGPT in the United States
- March 2, 2026 - Criteo becomes the first ad tech partner in the ChatGPT ad pilot
- March 23, 2026 - IMDEA researchers first notice tracking activity in Perplexity and Grok traffic
- March 31, 2026 - Class action filed against Perplexity AI, Meta and Google
- April 3, 2026 - Perplexity stops sending conversation URLs to third-party services such as Meta
- April 6, 2026 - Systematic testing of nine services begins
- April 13, 2026 - Researchers notify data protection authorities in the EU and UK
- April 17, 2026 - Researchers notify xAI about Grok's conversation access controls
- May 1, 2026 - Perplexity case voluntarily dismissed without prejudice
- May 4, 2026 - Part of the findings published on the LeakyLM project site
- May 13, 2026 - Class action alleges ChatGPT.com forwarded user queries to Meta and Google
- May 16, 2026 - PPC Land reports the UC Davis study of 20 chatbots
- May 27, 2026 - Spain's AEPD cites the work and asks for escalation to an EDPB plenary planned for June 6
- June 10, 2026 - LiveRamp connects its Conversions API Hub to ChatGPT ads
- August 15, 2026 - OpenAI tells European users that ads are coming and that its privacy policy is being updated
- August 31, 2026 - OpenAI reports a $1 billion annualised run rate for ChatGPT ads
- September 10, 2026 - Grok still uses publicly accessible permalinks; no official response from xAI
- September 28, 2026 - LiveRamp becomes a technology partner of ChatGPT Ads in 11 markets
- October 5, 2026 - Christl flags the IMDEA paper on LinkedIn
Related PPC Land coverage
- Brands' customer lists reach ChatGPT ads via LiveRamp in 11 markets - The September 28 report on LiveRamp's audience integration and how RampID matching works.
- LiveRamp CAPI Hub now connects ChatGPT ads - here's what it measures - The June server-side conversion connection that preceded the audience integration.
- Your AI chatbot may be sharing your prompts with ad networks - The UC Davis measurement of 20 chatbots on which the IMDEA paper builds.
- ChatGPT sued over secret data transfers to Meta and Google - The May complaint over the Facebook Pixel and Google Analytics on ChatGPT.com.
- Lawsuit claims Perplexity shared AI chat data with Google and Meta secretly - The March 31 federal complaint against Perplexity, Meta and Google.
- Perplexity data lawsuit dropped - but the privacy questions remain - How that case ended on May 1 without a ruling.
- ChatGPT Free and Go users in Europe face ads from later this month - OpenAI's August 15 notice to European users about ads and a privacy policy update.
- Criteo becomes first ad tech partner in OpenAI's ChatGPT ad pilot - The March 2 partnership the IMDEA paper's introduction points to.
- ChatGPT ads pass $1bn run rate as financial services spend triples - The scale of the advertising business now being built on ChatGPT.
- California lawmakers pass SB 690, cutting CIPA tracking lawsuits - The bill that narrows private wiretap suits over web tracking.
- LiveRamp-Publicis deal: what the SEC filing reveals about data, equity, and neutrality - The ownership change facing the RampID identity layer.
- Ad tech's most reliable numbers this week had to be pried out by courts - Includes a researcher's findings on a one-year OpenAI cookie and OpenAI's response.
Summary
Who: Researchers at IMDEA Networks (with UC3M and one independent co-author) behind the paper, with Wolfie Christl of Cracked Labs and The Citizen Lab drawing attention to it; the nine services tested (OpenAI, Anthropic, xAI, DeepSeek, Perplexity, Google, Microsoft, Mistral AI and Meta); LiveRamp and OpenAI in the related integration.
What: A measurement of third-party advertising and tracking in the web and Android clients of nine conversational AI services, finding that every service contacts at least one such firm, that 6 of 9 web clients and 3 of 8 Android clients disclose conversation artifacts, and that cookie rejection and subscription tier reduce but do not end the contact.
When: Experiments ran in May 2026; the disclosure timeline runs from March 23 to September 10, 2026; Christl flagged the paper on LinkedIn today, October 5, 2026. LiveRamp's related ChatGPT Ads announcement is dated September 28, 2026.
Where: Testing took place in Spain; the legal analysis addresses the EU ePrivacy Directive and the GDPR, and data protection authorities in the EU and UK were notified.
Why: AI assistants hold sensitive prompts, and ChatGPT now carries advertising, so the authors argue that conversational artifacts exposed to the ad-tech ecosystem, often next to persistent identifiers, call for stronger safeguards and regulatory scrutiny.
Discussion