What's firing on your site, and what you said was firing

A scan names the trackers it knows. The finding is the gap between what fires and the vendors your own privacy policy already named.

Last updated

You have already answered this question, in writing

Somewhere on your site sits a privacy policy, and behind your cookie banner sits a vendor list. Both make a specific factual claim about which companies receive data from the people who visit you. Somebody wrote that claim once, and nothing in a normal year prompts anyone to test whether it is still true. Privacy policies get revisited when counsel changes or a contract demands it, and neither of those events involves opening a browser and watching what loads.

That claim is the standard this audit should be graded against, and it is a better one than the arguments pixel guides make, data accuracy and regulation among them. The regulatory version says to run the audit because a regulator might come. Fair enough, except that a regulator’s timing is not yours to control, while the document you published is sitting in public being read by people who are not regulators at all — a client’s procurement questionnaire, a partner’s security review, an enterprise buyer’s legal team, anyone who thinks to look.

In the GCC that framing does more work than it would elsewhere, because the regional picture is genuinely uneven. Saudi Arabia’s Personal Data Protection Law came into force on 14 September 2023, with its implementing regulations published a week earlier, and a twelve-month grace period that ended in September 2024; consent is the general basis for processing personal data, subject to a defined set of exceptions. The UAE’s law has been in force since January 2022, but its executive regulations still have not been issued and no administrative penalties have been set under it. A business trading across both markets is therefore answering to a settled regime on one side of the border and an unfinished one on the other.

Waiting for that to resolve is the wrong move, and not because we think a fine is imminent. It’s that resolution changes nothing about the work. Whatever the executive regulations eventually say, the distance between your published vendor list and your real one is a gap today, discoverable by anyone with a browser, and closing it is the same afternoon either way.

The pages a scan never reaches

The quick version of this audit ends in one place: install an extension, load the site, read the list it hands you. That’s a good first ten minutes and a weak audit, because what it produces is a picture of one page, in one session, on one device — and it’s the page you’d have guessed anyway.

Sprawl doesn’t spread evenly, and where it clusters tells you where to walk.

Start at the bottom of the funnel. Cart, checkout, payment step and above all the thank-you page: every advertising platform wants its confirmation firing where the transaction completes, so that is where tags accumulate, and it is also the page nobody loads casually with developer tools open. Then there are the campaign landing pages, put together quickly, often outside the main template, carrying whatever a platform’s onboarding wizard suggested at the time, and never folded back into the main build once the campaign ended. A bilingual site can double the surface again. If your Arabic pages run on a separate template instead of sharing the English one, that is a second head block collecting its own snippets, and a scan of the English homepage tells you nothing about what loads on the Arabic side. And check a phone, because chat and support widgets in particular behave differently there.

So walk the site the way a customer does, and write down what fires as you go. Homepage to a category page to a product to cart to a completed transaction, on desktop and on a phone, and once down each language version you publish. For this pass, skip the extension and open the Network panel in your browser’s developer tools. Turn on Preserve log so the list survives each page change, right-click the column headers to add Domain, and sort by it. Note every domain that isn’t yours, and any request to your own domain you can’t account for, because server-side tagging can run there. An extension tells you which vendors it recognises. The Network panel shows every third party the browser contacted, including the ones no recognition database has heard of, which are disproportionately the interesting ones.

Then collect the two sources a live page can’t give you. Export your tag manager container, so you have every tag it holds and not just the ones your route happened to trigger. And read the page source, or the theme templates, for script tags that never went through the container at all. Three sources, deliberately. The overlap between them is the point.

Three ways a script gets on a page, and only one you fully govern

The tag manager container comes first. Consent settings in Google Tag Manager are configured per tag, and every tag sits in a Not set state until somebody configures it. The Consent Overview sorts the container into what has been configured and what hasn’t, which puts its whole consent posture on one screen. This is the door you want scripts coming through.

Then there are the scripts that never went near it — a snippet pasted into the site template, the theme header, a landing page’s head block, a plugin’s settings field. The container has no knowledge of these. Consent mode reaches only the Google tags among them, and only if a consent default sits above them on the page; every other hard-coded script ignores it. A consent management platform with automatic blocking can hold many of them back without per-script setup, but only the ones the platform’s scan has matched to a cookie category.

Hardest to see are the scripts you never added at all: a chat widget that pulls its own analytics library, an advertising tag that redirects into a partner’s network, a consent platform that fetches a font. These appear on no list you maintain, because the decision to load them was made inside another company’s product.

Put those together and the conclusion is uncomfortable and simple. A banner wired only into the container is a doorman on one entrance of a building with three. The fix is architectural: move every script you can through the first door, so the thing you configured is the thing that governs, and keep a written list of the ones that genuinely can’t move.

Price each row before you judge it

You now have the firing list. Each row needs a price before it can be reconciled against anything, because “remove unused pixels” is unactionable until you know what keeping one costs.

The performance case here gets argued badly, as though third-party code were a single substance you have too much of. Around nine in ten pages on the web load third parties at all, and the median page pulls roughly eighty third-party requests, so “fewer third parties” is close to meaningless as a goal. Cost per row is the useful unit.

Start with the cheapest thing on your list. A conversion tag on a thank-you page is a few requests on a page whose speed influences nothing — nobody abandons a purchase they have already completed — so its performance cost is a rounding error and you should judge it entirely on whether anyone uses the data.

Session recorders and heatmap tools carry a different kind of cost. They instrument every scroll, click and pointer movement for as long as the page is open. Microsoft reports no measurable load-time impact from Clarity for most sites, so the price worth weighing is what gets recorded. That is worth carrying during a research window and much harder to justify permanently, and the honest test is when somebody last opened a recording.

We’d look hardest at chat and support widgets, because they load a substantial bundle on every page, whether or not the visitor ever opens the chat. Swap the widget for a lightweight placeholder that fetches the real thing on click and you keep the function at a fraction of the cost.

Retargeting tags almost never look like a problem individually, and that’s the problem. They arrive one per campaign and leave on no schedule at all, so the count only goes up.

If you want a threshold to argue with rather than a feeling, Lighthouse failed pages whose third-party code blocked the browser’s main thread for more than 250 milliseconds in total, until version 13 turned that audit into an insight that always passes. The number still makes a fair line to measure yourself against. It won’t tell you which script to cut, but it will tell you whether the problem is yours or somebody else’s.

Reconcile the two lists, and the gap is the finding

Now set the priced inventory beside the published vendor list and work down it. Every row lands in one of three states.

Some rows come back matched. Treat that as a result: the vendor fires, your policy names it, a named person uses the data. Write the owner’s name next to it and move on.

The gap you came for is a vendor that fires and isn’t named — you are collecting data from visitors and passing it to a company you never told them about. There are two honest responses and choosing between them is a real decision. Remove the vendor, or amend the published statement to include it. Removal is usually the right call, for a structural reason: a vendor no named person will claim is a vendor whose output nobody is reading, which makes it dead weight whatever you decide about disclosure. Amending is right when the vendor is genuinely in use, and it isn’t a defeat. An accurate list naming twenty-four companies is a stronger position than an aspirational one naming nine.

Then there’s the reverse case, which gets skipped and shouldn’t: your policy names a vendor that fires nowhere you looked. That’s either a decommissioned tool nobody removed from the document, or a live campaign whose tracking has broken and which is currently optimising against nothing. The second one is costing money this week, and it’s the single finding in this whole audit with a same-day payback.

Removing things without breaking somebody’s month

Before anything comes off, ask who reads the output. That question resolves more rows than any technical check, and “nobody knows” is itself a resolution.

Sequence the rest so mistakes stay cheap. Anything inside the container can be paused rather than deleted, which stops it firing and leaves the configuration intact, so restoring it takes seconds. Anything hardcoded has no pause, so make its removal a tracked change in version control and record the exact snippet you pulled — otherwise “put it back” three weeks later becomes an archaeology project. Then let a full reporting cycle pass before deleting anything permanently. A month covers the point at which somebody who depended on the data notices that they did.

The one place we’d slow right down is advertising pixels on live campaigns, and in Saudi Arabia that caution has to be stronger than the general advice suggests. Snapchat’s advertising reach there was reported at 25.3 million people in late 2025, close to three-quarters of the population. At that reach, treat a Snapchat or TikTok pixel on a Saudi site as feeding a live plan until the people running paid media say otherwise. Switch one off mid-flight and nothing on the site errors, while the campaign’s optimisation stops receiving the signal it has been learning from.

So tell the media owner before rather than after, and give them the flight dates to work with. A pixel that survives this audit because an agency asked for two more weeks is a much better outcome than a clean list and an argument in November.

Stopping it from coming back

Sprawl is a governance failure with a technical symptom, which is why the cleanup keeps needing repeating at businesses that have already done it once. Nobody ever decided to run three dozen vendors. Three dozen separate reasonable decisions were made, and none of them carried an end date.

Three rules close the loop, and they’re cheap enough to actually hold. A new pixel gets a named owner before it gets added. It gets its line in the privacy policy the same week, not at the next annual review, which is the rule that keeps your two lists from drifting apart again. And it gets a review date, so somebody is scheduled to ask whether it still earns its place.

Then re-run the walk quarterly. It’s an hour once the journey is written down, and the second run is faster than the first, because most of what you find is already on the list.

Two adjacent audits are worth doing in the same stretch. Cleaning the container itself takes the tags coming through the first door and decides what to do with them, and the GA4 property those tags feed decides whether the data you kept is worth reading. If third-party weight turned out to be your real problem, that belongs with the wider site-speed work, and if your Arabic templates came back looking like a different site, the bilingual build is where that thread goes. Doing this in your first months in a job, on scripts you inherited and never chose, puts it inside the handover conversation instead.

Our Site Audit checks tracking as one of its four areas, which will tell you whether your setup is worth a closer look; it won’t replace walking the journeys yourself. And if the pixel list turned out to be the third broken thing you found this month, the rest of the reporting stack fails in a similar pattern.

Standing offer

The hard part is deciding what to do first

Bring us a URL, an audit report, or a proposal you're not sure about. We'll tell you which problems deserve budget this quarter, and you'll leave with an order of operations you can hand to whoever does the work, whether that's us or not.

Book twenty minutes