Chatbot privacy: audit your chat widget for trackers

Chatbot privacy starts in the browser network tab. A 10-minute audit method for chat widget trackers, and what to demand from vendors.
TL;DR: Researchers who audited nine conversational AI services found that every one contacts at least one third-party advertising or tracking service, and that six of nine web clients passed conversation artifacts such as titles or prompts to third parties. Your support chat widget deserves the same scrutiny. You can audit it in about ten minutes with the browser network tab: load the page in three consent states, list every third-party domain, and search the requests for identifiers and conversation text. As of October 2026, that audit and a short vendor questionnaire are the most reliable privacy controls you have.
I audit data flows for a living in the sense that matters here. When someone hands me a chat widget, I do not read the privacy policy first. I open the browser, watch what the page sends and to whom, and only then read the policy to see whether it describes what I just saw.
That is the method of a privacy audit, and this article follows it. A policy tells you what a vendor intends. The network tab tells you what the code does, and the two disagree more often than anyone likes.
The GDPR guide for AI customer support covers the legal framework, and this article gives you the hands-on method to test whether your own deployment matches it.
A recent academic paper gave this method fresh urgency. It measured what nine popular AI chat services send to trackers, and the results should make anyone who runs a support widget check their own. I will walk through what the paper found, what it does and does not say about support widgets, and then the exact steps to run the same kind of test on your site.
Where I describe communicate.so, I say what I verified and what I did not.
What the research found about AI chat and trackers
The paper is titled Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents, by Guilherme Oliveira and eight co-authors at IMDEA Networks, and it is slated for Proceedings on Privacy Enhancing Technologies 2027(1). It examined the web clients of nine services and the Android clients of the eight that offer an Android app.
The headline numbers are specific. Across the services, the authors observed 124 distinct third-party domains, which they attribute to 44 organizations, and 34 of those organizations are advertising and tracking services. In their words, "every evaluated AI service integrates at least one third-party advertising or tracking service."
The more worrying finding concerns what gets sent. The authors report that 6 of 9 web clients and 3 of 8 Android clients "disclose conversation URLs, titles, prompts, and screenshots to third-party services, often alongside persistent user identifiers." A tracker that receives a conversation title and a hashed email address can connect a person to a topic.
They also found that consent changes the picture. Some third-party services activated only after users explicitly accepted non-essential cookies, which means the tracker list you see on first load is not the list a consenting visitor gets. A privacy test that stops at the first page view misses part of the surface.
One more finding carries a lesson for anyone who shares chat links. Several providers made conversation permalinks publicly readable to anyone who holds the URL, and the authors note that this allows trackers to read the entire conversation. The paper states the broader point plainly: "our results challenge the perception of conversational AI services as confidential exchanges between users and AI providers."
The authors were also careful about what the data cannot show, and I want to be equally careful. They write that "the presence of a third party does not by itself imply that data is transmitted under all conditions only for advertising or tracking purposes," since some SDKs support operational functions. A third-party domain in your network tab is a prompt to ask a question, and it is not a verdict.
What the paper does not say about your support widget
Precision matters here, because this research is easy to overstate. The nine services studied are consumer AI assistants that people use to ask their own questions. They are not customer support widgets embedded on company websites.
The authors also state the boundaries of their own work. They excluded enterprise and governmental tiers, which providers market with distinct data-handling commitments, and they could not extract traces for one service's mobile app. Their prompts were health-related and written to emulate a realistic user, so the exposure they found applies to that kind of content in those products.
So the paper does not say your chat widget leaks. It says something narrower and more useful. It shows that conversational AI products inherited the advertising and analytics plumbing of the wider web, that conversation-derived artifacts can travel through that plumbing, and that a black-box browser audit is enough to detect it.
Support teams should treat that as a reason to run the audit, and the PII redaction guide explains why the content of support chats makes the question sharper than it is for most websites.
There is one more reason to pay attention. The paper cites a Reuters report that OpenAI partnered with Criteo on an advertising pilot for ChatGPT free-tier users in the United States in early 2026. When the biggest names in the category adopt advertising models, the surrounding ecosystem of widget vendors feels the same commercial pull.
That pull is a reason for vigilance and not a reason to assume bad faith. Most widget vendors add a tracker for a mundane reason, such as a developer who wanted error reports. The audit tells you which kind of tracker you have.
Why a support widget carries more risk than a normal script
A typical analytics script sees page views. A support widget sees what customers type when they are confused, upset, or in a hurry, which is far more sensitive.
Customers paste order numbers, email addresses, phone numbers, and sometimes account details into chat. They describe a failed payment, a medical appointment, or a delivery to a home address. The widget also knows the page the customer was on and, when you pass it, the customer's identity.
The paper's list of third-party categories shows where this goes wrong. In its survey, the most common organizations after Google were Sentry, Meta, Datadog, and Intercom, which the authors associate with "error monitoring and observability, customer support and engagement, but also advertising and analytics." Error monitors capture page context when something breaks, and customer engagement tools capture messages by design, so both can end up holding chat content that you did not plan to share.
The risk compounds because the widget runs inside your page. Any script you load can read the page the visitor sees, and any request it makes carries the visitor's cookies for that domain. A widget that loads a second vendor's script has extended your trust boundary without a conversation about it.
For regulated teams, the stakes are higher still. If your support channel handles health, finance, or education data, the guardrails guide and the data retention guide describe controls that matter only if the data is not already leaving through a side door.
What counts as a tracker, in plain terms
Before you open any tools, agree on the vocabulary. People argue past each other because they use the same words for different things.
A first-party request goes to a domain you or your vendor controls as the site operator, such as your own domain or the vendor's chat API. A third-party request goes to any other organization. The paper uses the term advertising and tracking service, or ATS, for a third party whose business is advertising, attribution, analytics, or profiling.
- Analytics scripts count visitors and clicks. Google Analytics and similar tools are the usual examples.
- Advertising and attribution pixels tell an ad network that a visitor did something. Meta, TikTok, and Google Ads pixels fall here.
- Session replay tools record mouse movement and sometimes typed text. They are the most dangerous category for a chat widget.
- Error and performance monitors collect stack traces and page context when something fails. They are often benign and can still capture chat text in breadcrumbs.
- Consent management platforms record and enforce cookie choices. They are third parties, and they should be the only scripts running before consent.
- Identity and customer data platforms stitch visitors across sessions. Segment-style tools are the usual examples.
Two terms from the paper matter for the audit itself. A persistent identifier is any value that stays the same across visits, such as a cookie ID or a device ID. A hashed email address, which the paper abbreviates HEM, is a fingerprint of someone's email that lets two companies match the same person without sending the address in plain text.
The authors searched captured traffic for direct occurrences of identifiers and their common hash transformations (MD5, SHA-1, and SHA-256). You can copy that idea with a few lines of scripting, and I show how later.
The 10-minute browser audit, step by step
You need a desktop browser, the page that hosts your chat widget, and a throwaway test conversation with fake data. You do not need special tools. Chrome, Firefox, and Edge all include a network panel.
Chrome's documentation for its Network panel covers the controls I refer to below. The steps assume a fresh browser profile or a private window with no extensions, because an ad blocker will hide the very requests you are looking for.
- Open a private window with extensions disabled. Open developer tools, switch to the Network tab, and tick Preserve log and Disable cache.
- Load the page that hosts the widget. Do not click anything, and do not touch the cookie banner. Wait thirty seconds.
- Sort by Domain. Write down every domain that is not yours and not your chat vendor's API. This is your idle inventory.
- Open the widget and send three test messages that contain a fake name, a fake email address, and a fake order number. Watch which new requests appear.
- Click each new request and read the request URL, the headers, and the payload. Search the Network tab for your fake email, your fake order number, and the text of your messages.
- Repeat with the cookie banner set to reject, then with it set to accept, using a fresh private window each time. Compare the three domain lists.
- Export each session with Save all as HAR with content, and keep the files with the date and the page URL.
That is the whole method. Three page loads, one conversation each, and a written list of domains.
The fresh window between consent states matters. Cookies and local storage persist inside a session, so a tracker that set an identifier during the accept run will behave differently on a second load in the same window. Close the window, open a new one, and start clean.
Budget about three minutes per state. The remaining time goes to reading the payloads that look suspicious, which is where the findings are.
Read the requests like an auditor
A domain list is a start. The payload tells you whether a request matters. Here is how to triage what shows up.
| What you see in the request | What it usually means | Action |
|---|---|---|
| Your fake email or order number in a URL parameter | Personal data is leaving in a place that web servers log by default | Treat as a finding. Ask the vendor to remove it |
| Your message text in a request to a domain that is not your chat vendor | Conversation content is shared with another party | Treat as a serious finding. Block the script until explained |
| A 32, 40, or 64 character hex string near an identifier field | Possibly an MD5, SHA-1, or SHA-256 hash of an email | Hash your fake email and compare. A match means identity matching |
| Page URL and title sent to an ad or analytics domain | Standard page-view tracking, which may include chat-related page names | Check consent state. Fine only if the visitor allowed it |
| A long random value that repeats across requests | A persistent identifier or session ID | Note which domains receive it |
| Request to a consent platform domain only | Cookie banner doing its job | No action beyond confirming it runs first |
| Request to an error monitor with a large payload | Error report that may contain breadcrumbs of page content | Open the payload and look for chat text |
Two habits make this faster. First, use the filter box to search for your fake values across all requests at once. Second, right-click a suspicious request and copy it as cURL, so you can replay it later and show the vendor exactly what you saw.
When you find your fake email in a request, check how it is encoded. It may be URL-encoded, Base64-encoded, or hashed. If you see a long hex string in an identifier field, compute the hash of your test address and compare, because the paper's authors did the same thing at scale.
Test the consent states, not just the first load
The paper tested cookie consent in three states, which it labels ignore, reject, and accept, and it found that some third-party services appear only after acceptance. Your widget deserves the same treatment, because the banner is where legal promises meet code.
In the ignore state, the page should load only what is strictly needed to run the site. Any advertising pixel that fires here is a finding. In the reject state, the list should match the ignore state or be shorter.
In the accept state, additional domains are expected, and the question becomes whether the visitor was told about them. Compare the list to your cookie policy line by line. A domain that appears in the network tab and not in the policy is a gap you can close the same day.
| Consent state | Chat vendor API and storage | Disclosed analytics | Ad pixels or session replay | Message text to a non-chat domain |
|---|---|---|---|---|
| Ignored banner | ✓ Expected | ✗ Should not load | ✗ Red flag | ✗ Red flag |
| Rejected | ✓ Expected | ✗ Should not load | ✗ Red flag | ✗ Red flag |
| Accepted | ✓ Expected | ✓ Allowed if listed in your cookie policy | ✓ Allowed only if disclosed | ✗ Red flag |
| Chat opened, banner ignored | ✓ Needed for the conversation | ✗ Should not load | ✗ Red flag | ✗ Red flag |
| Chat opened, banner rejected | ✓ Needed for the conversation | ✗ No identifiers shared | ✗ Red flag | ✗ Red flag |
| Chat opened, banner accepted | ✓ Needed for the conversation | ✓ May see page context | ✓ Allowed only if disclosed | ✗ Red flag, including hashed email |
The legal reference point in Europe is Article 5(3) of the ePrivacy Directive 2002/58/EC. In plain terms, it requires consent before a site stores or reads information on a visitor's device, unless that access is strictly necessary to provide a service the visitor asked for. The paper's authors analyze their observations against the GDPR and this directive, and I recommend reading section 7 of their work if you want the full argument.
I am not a lawyer, and nothing here is legal advice. What an audit can show is that you cannot make a good-faith consent argument without knowing what loads, and the network tab is how you find out.
A chat widget raises a second, quieter question. Local storage that holds a conversation ID so the visitor can resume a chat is arguably necessary for the service they asked for. Local storage that holds an advertising ID is not, and the audit should tell you which you have.
Search for identifiers and hashed emails
The most serious pattern in the paper was conversation artifacts traveling beside persistent identifiers. You can look for the same pattern in a HAR file you exported.
Pick two fake test emails and compute their MD5, SHA-1, and SHA-256 hashes, in lowercase hexadecimal. Then search the HAR file for each value and for the plain address. A match in a request to a domain that is not your chat vendor means that your widget or your page is sharing an identity with a third party.
- Search for the plain test email and for its URL-encoded form, where the at sign appears as %40.
- Search for the three hash values, since the paper searched for direct occurrences and for MD5, SHA-1, and SHA-256 transformations.
- Search for your fake order number and for the first few words of each test message.
- Search for the visitor ID values you pass to the widget on purpose, such as a customer ID, and note every domain that receives them.
- Record the domain, the request URL, the field name, and the consent state for each match.
Do this on a test account, never on real customer data. The point is to learn how the code behaves, and fake values are enough.
If you find nothing, write that down too, with the date, the browser version, and the page. A dated clean result is useful evidence the next time a customer, auditor, or procurement team asks.
Public links and shared transcripts
The paper's second concern was not a tracker at all. It was the conversation permalink, a stable URL that identifies a chat, which some services leave readable by anyone who has the address.
Support tools have their own version of this. Many generate transcript links for emails, CSAT surveys, or agent handoffs, and some of those links work without a login. If a tracker or a log captures the link, whoever holds it can read the transcript.
Test your own transcript links. Open one in a private window where you are not logged in, and see whether it loads. If it does, ask how long the link lives, whether it can be revoked, and whether it appears in any third-party request, and compare your answers with the controls in the audit trail guide.
The paper gives a concrete example of the exposure. It points to a Grok conversation URL in which the path segment identifies the conversation, and it states that such a URL is by default publicly readable for any actor knowing it. A tracker that receives that URL in a page-view event receives a key to the conversation.
The fix on your side is boring and effective. Use links that expire, require authentication for transcripts, and strip query strings and fragments before page URLs reach analytics tools.
What to demand from a chat vendor
An audit finds problems after you have deployed. A questionnaire finds them before. Send these questions in writing and file the answers.
- List every third-party domain that your widget script contacts, with the purpose of each and whether it runs before consent.
- State whether any conversation text, title, URL, or screenshot is sent to a third party, and under what contract.
- State whether any session replay or heatmap tool is loaded inside the widget.
- Provide the current subprocessor list, with a notice period for changes.
- State which cookies and local storage keys the widget sets, their lifetimes, and which are strictly necessary.
- Confirm whether visitor email or ID values are hashed or shared with any advertising or identity platform.
- Describe how transcript links are secured, how long they live, and how they can be revoked.
- Confirm that the vendor will notify you before it adds a new third-party script to the widget bundle.
- Describe your error monitoring setup and whether it scrubs message content from reports.
- Describe the model provider arrangement, including retention and training terms.
The vendor questions checklist goes wider than privacy, and the SOC 2 guide for AI support explains what an audit report does and does not tell you about trackers. Neither replaces the network tab, because a report describes the controls an auditor tested and the widget you deploy can change after that.
Insist on written answers. A salesperson who says there are no trackers has told you something you cannot verify and cannot enforce. A written list of domains is something you can compare against your own capture.
Also decide in advance what a bad answer means. A vendor that cannot list its third-party domains has told you how well it knows its own product. I would treat that as a reason to keep looking.
Fixes you control on your own site
Even a well-behaved widget lives on your page, and your page decides what else runs next to it. These controls sit with you.
- Load the widget after consent only if it sets non-essential storage, and load it immediately only if you have confirmed it is limited to what the chat needs.
- Add a Content Security Policy that lists the domains your widget may contact. Anything outside the list is blocked and reported.
- Strip query strings, fragments, and email-like values from page URLs before they reach analytics.
- Turn off session replay on every page that hosts the chat widget, or mask the widget container.
- Pass the minimum visitor data to the widget. If the bot does not need an email address, do not send one.
- Review the network tab again after every widget or tag manager change.
Redaction is the last line of defense. Even if a tracker receives a request, it cannot read what was removed earlier, which is why the PII redaction guide belongs in the same project as this audit.
A Content Security Policy deserves extra attention because it turns an audit finding into a standing control. Once your allowed domain list exists, the browser enforces it for every visitor, and violation reports tell you when something new tries to load.
What I could verify about communicate.so
I work on communicate.so, so apply the usual discount to this section. I will separate what I checked from what I did not.
I read the source of the chat widget script, the embeddable widget, in a branch snapshot of the repository on October 8, 2026. The file is about 146 KB. It makes requests with the browser fetch API and an XMLHttpRequest call, it uses local storage and session storage, it contains no use of document.cookie, and the only absolute web address written into it is https://communicate.so.
That is a narrow finding, and I want to keep it narrow. It describes one file in one snapshot. It does not describe what is served in production today, what your site loads next to the widget, or what our marketing site and dashboard load.
It also says nothing about requests the script builds at runtime from configuration, which a source read cannot reveal. Only a capture in a browser, run by you on your page, settles that. For that reason I am not claiming that the widget has no trackers.
I am telling you how to find out.
What I can point to on the record is the security page. It names the subprocessors we use, says an up-to-date list is available on request, states that customer data is not used to train shared models, and lists the certifications we do not hold. Read it alongside your own capture, and ask us for the domain list in writing, as you would any vendor.
If you run this audit on a communicate.so widget and find a request you cannot explain, send it to the security contact on that page. A reproducible capture is more useful to us than a general concern.
Limits of a ten-minute audit
A short browser test is a screen, and it has gaps. Knowing them keeps you from trusting a clean result too much.
- It sees client-side requests only. Anything the vendor shares from its servers never passes through your browser.
- It tests the paths you take. A tracker that loads on a checkout page or after a login will not appear on a docs page.
- It cannot tell purpose from presence. As the paper's authors note, establishing purpose requires legal and contractual analysis.
- It is a snapshot. Tag managers, widget updates, and consent platform changes can alter the list overnight.
- It covers the web. Mobile SDKs need a different method, and the paper used an instrumented Android build for that.
Treat the result as one input. Combine it with the vendor's written answers, a contract review, and a repeat test on a schedule.
If the audit leads you to change vendors or settings, plan the switch carefully using the implementation guide. Moving a widget is quick, and moving the conversation history and the consent records is the part that takes time.
Frequently asked questions
How do I check if my chat widget has trackers?
Open a private browser window, open the Network tab in developer tools, load the page with the widget, and list every domain that is not yours or your chat vendor's. Then send a test message with fake data and search the requests for those values.
What is the difference between a first-party and a third-party request?
A first-party request goes to your own domain or to the vendor's API that runs the chat. A third-party request goes to another organization, such as an advertising network, an analytics provider, or a monitoring service.
Is every third-party domain in the network tab a privacy problem?
No. The paper's authors note that a third party's presence does not by itself imply that data is sent only for advertising or tracking, since some SDKs support operational functions. Treat each domain as a question to answer.
What did the research on AI chatbots and trackers actually find?
The authors found 124 third-party domains across nine services, attributed to 44 organizations. Every service integrated at least one advertising or tracking service, and 6 of 9 web clients passed conversation artifacts to third parties.
Does that research apply to customer support chat widgets?
Not directly. The nine services studied are consumer AI assistants. It shows that the same method works and that conversation artifacts can reach trackers, so it is a reason to test your own widget.
What is a hashed email address and why does it matter?
It is a fingerprint made from an email address, commonly with MD5, SHA-1, or SHA-256. Two companies can compare fingerprints to match the same person without exchanging the plain address, which is why a hashed email beside conversation data is a concern.
Should I test with real customer data?
No. Use a test account and fake values, such as a made-up name, email, and order number. Fake values are easy to search for and put no real person at risk.
Why test with the cookie banner rejected as well as accepted?
The paper found some third-party services that activated only after cookies were accepted. Testing ignore, reject, and accept shows whether your banner changes what loads and whether the changes match your policy.
What is a session replay tool and why is it risky in a chat widget?
It records what a visitor does on a page, and some tools capture typed text. Inside a chat widget, that can mean recording the messages themselves, so I would keep replay tools off any page that hosts a chat.
Can an error monitoring tool leak chat messages?
It can. Error reports often include page context and recent events, and a report triggered during a chat may carry fragments of the conversation. Check the payload of any monitor request for message text.
Do transcript links need a login?
They should, or they should expire quickly and be revocable. The research found providers whose conversation permalinks were readable by anyone with the URL, and the audit trail guide covers how to keep records without exposing them.
What does the ePrivacy Directive say about widgets?
Article 5(3) generally requires consent before storing or reading information on a visitor's device, unless the access is strictly necessary for a service the visitor requested. Ask counsel how that applies to your widget's storage.
How does GDPR apply to a chat vendor?
A vendor that processes chat messages for you is usually a processor, and you need a contract and a subprocessor list. The GDPR guide explains the obligations in more detail.
Can I block trackers with a Content Security Policy?
Yes. A policy that lists the domains your pages may contact makes the browser block everything else and report the attempt. It works best after an audit has produced the list of domains you actually need.
How often should I repeat the audit?
Repeat it after every widget upgrade, tag manager change, or consent platform change, and on a fixed schedule such as every quarter. Keep the dated HAR files so you can compare results over time.
What should I ask a vendor who says the widget has no trackers?
Ask for the list of third-party domains in writing, and compare it with your own capture. A claim with no list cannot be checked, and a list that disagrees with your capture is a finding.
Do mobile apps need a different audit?
Yes. The paper analyzed Android clients with an instrumented operating system build, because embedded SDKs run inside the app and are invisible to a desktop browser. If you ship a mobile chat SDK, ask for its dependency list.
Does redacting personal data help if trackers still load?
It helps a great deal, because a tracker cannot receive what was removed first. Redaction does not replace blocking unnecessary scripts, so use the two controls together.
Can I prove to a customer that my widget is clean?
You can show them a dated capture with the browser version, the page, and the domain list, along with your vendor's written answers. That is better evidence than a policy statement and far better than a verbal assurance.
Where can I see how communicate.so handles subprocessors?
The security page names the subprocessors we use and says an up-to-date list is available on request. The data sources page describes what content the agent reads.
Run the audit this week
Set aside ten minutes, a private window, and a fake customer. Capture the three consent states, search for your test values, and file the result with a date. If you want to test an AI agent and widget against your own pages, you can start on the pricing page and run this same audit on the sandbox before you go live.