Skip to content
StackPatrol
How it works

How StackPatrol works

We believe a third-party scanner should be transparent. Here is exactly what StackPatrol does when you submit a URL.

1
Crawl
Headless Chromium loads your URL
2
Record
Every network request captured
3
Match
763 known vendor patterns
4
Score
EU Independence Score computed

1. Crawling

We launch a headless Chromium browser (via Playwright) with a desktop User-Agent and a 1366×800 viewport. We then load the URL you provided and wait for the page to reach networkidle for up to six seconds.

Paid scans improve coverage by following same-domain links up to the selected plan limit: five pages on Pro and Agency Starter, or twenty on Agency and Agency Pro. Auto-discovery prioritises useful internal routes, while paid users can also pin specific paths such as checkout, account or contact pages. Every path remains inside the total page budget.

Free scans load the front page once and record what the browser does without any interaction — no logins, no form submissions, no consent-banner clicks. Paid plans automatically run consent-aware scanning: we re-open the site in a clean browser context, click an “Accept all” button on the consent banner, then revisit the entry page plus up to four of the baseline subpages in the same context (so the consent cookies and localStorage flags persist). The report shows a side-by-side delta of services, cookies and third-party domains that only load after consent.

The consent clicker searches both the main page and likely CMP iframes (Sourcepoint, OneTrust, Cookiebot, Didomi, TrustArc, Quantcast, Usercentrics, Iubenda, Klaro, Osano, Borlabs, Complianz) using a curated list of vendor selectors and accept-all button labels in ~15 languages. When deterministic selectors miss, an optional LLM fallback (gpt-4o-mini, cached per host) identifies the accept button from a pruned list of visible elements. Only bounded control metadata from the scanned site (label text, tag, CSS classes, IDs and ARIA labels) is sent. Frame identifiers are opaque, and common email, phone, URL and long identifier patterns are redacted first. We do not intentionally send personal data, the scan URL, cookies or page content (see our privacy policy). Against an 18-site smoke test of major European publishers, the current hit rate is around 89%.

2. Recording requests

For every network request the browser makes we record the URL, the host and the resource type. A request is classified as third-party when its registrable domain is different from the registrable domain of the scanned site.

To determine the registrable domain we use the Public Suffix List — the same list browsers use to decide where one organisation’s domain ends and a shared registry begins. This correctly handles multi-part suffixes like .co.uk, .com.au and .co.jp, so shop.example.co.uk collapses to example.co.uk rather than the registry co.uk.

On paid plans we also tag each retained request with the page it fired on and whether it ran before a choice, after Accept all or after Reject all. Where Chromium exposes it, the evidence can also include request method, resource type, redirect hops and the initiating URL. This provenance lets the report show where and when a service was observed instead of only saying that it exists somewhere on the site.

Consent-control evidence is deliberately privacy-reduced. We retain bounded metadata about the activated control and may capture a small image of that control, never a full-page screenshot. Storage entries and TCF strings are represented by key names, byte sizes and SHA-256 hashes; raw storage values, raw cookie values and raw TC strings are never retained. Paid users can export a sanitized evidence manifest with stable evidence IDs and an integrity hash. Request evidence is bounded per service, so it is a reproducible sample rather than a claim that every request made by the site was retained.

3. Vendor matching

We maintain a curated database of 763 third-party vendors: browse the directory. Each vendor entry contains one or more domain patterns.

A request matches a vendor when:

  • its hostname equals the pattern, or
  • its hostname ends with . + pattern (suffix match), or
  • for the few non-domain patterns we use, the full URL contains the pattern.

When a request matches multiple patterns we pick the most specific (longest) one. Domains that don’t match any vendor are listed as unmatched.

4. Region classification

Each vendor is classified by ownership region. The classification is based on where the parent company is incorporated, not where data physically resides. A US-owned vendor with EU data centres is still classified as US because data-access requests (FISA 702, Cloud Act) are governed by ownership.

US
FISA 702 / Cloud Act
EU
GDPR home jurisdiction
EEA / UK
Adequacy decisions
China
Data law concerns

5. Loaded before consent and transfer review

Beyond the score, two parts of the report answer the questions a DPO actually asks: what fires before the visitor consents, and where does the data go. Both are derived directly from the recorded requests — they describe what we observed, not whether it is lawful.

Loaded before consent

The no-interaction load is our baseline. For multi-page scans, it consists of one clean no-interaction load per scanned URL, before any consent banner is clicked. We deduplicate the observed services across that baseline inventory, keep only genuine tracking-relevant categories (advertising, analytics, tag management, error tracking and similar), and separate first-party services from third-party ones. The headline count reflects only third-party trackers, because a site loading its own subdomain is not the concern.

This is a technical observation from the scan date. It is not a legal conclusion about consent validity.

Observed external services and transfer review worksheet

Technical observations that can support updates to processing records and vendor registers. Legal and organisational fields require customer verification. We group each observed external service by ownership and DPF signals, not by a concluded processing location or legal mechanism:

EU / EEA ownership
EU / EEA-owned services
Adequacy / DPF signal
UK or Swiss ownership, or a matched US DPF certification. Customer verification is still required.
Customer review required
No applicable mechanism can be inferred from the observed ownership and DPF signals.

Each service is also tagged with whether it fired before or after consent. The exposure matrix counts observed services per signal bucket and phase. It helps prioritise review, but it is not a GDPR risk rating. Paid plans list a possible mechanism or review prompt and export a CSV worksheet with customer-owned fields left empty.

How we verify DPF status: we don’t guess. A US-owned service only receives a DPF match signal when its legal entity is matched against the official participant list published by the U.S. Department of Commerce at dataprivacyframework.gov and that entity holds an active EU–US certification. When it matches, the report shows the exact certified entity and the date we last checked it (e.g. “Verified 8 Jul 2026 · Stripe, LLC”), so you can see how fresh the finding is. An automated job re-checks the full list every week: new certifications are picked up, and if a vendor’s certification lapses we’re alerted and the status is downgraded. Vendors we can’t match to an active certification are left as “status unknown” rather than implying a negative.

Observed ownership, DPF matches and request timing are technical signals, not conclusions about processor roles, processing locations or transfer legality. Customer verification is required before updating legal records.

6. The EU Independence Score

The score is an experimental signal, not a compliance rating. It starts at 100 and four penalty components are subtracted:

Score = 100 − P_vendor − P_mix − P_unknown − P_infra

First-party ownership override

The formula above applies to sites owned by an EU/EEA company. If the scanned site itself is owned by a US, Chinese or Russian company, that ownership takes precedence over the vendor maths: the score is capped and relabelled (for example US-owned site), because the third-party stack is then mostly the owner’s own first-party assets. In that case the score reflects the site’s ownership, not the EU-friendliness of its embedded vendors. This override is why an EU-independence figure and an ownership caveat can appear together on the same report.

P_vendor — jurisdiction risk × category
EU / EEA
GDPR home jurisdiction
0
Switzerland / UK
Adequacy decision in place
2–3
Global / Unknown owner
Jurisdiction unclear
5–6
US-owned
FISA 702 / Cloud Act jurisdiction
8
China-owned
PIPL / national security law jurisdiction
14
Each base risk is multiplied by a category weight. A US-owned tag manager (×1.8) is penalised more heavily than a US-owned font (×0.6). The score penalty is capped at 20 per ownership group and 60 total.
Data Privacy Framework nuance: for US-owned services we check the relevant legal entity against the official EU–US Data Privacy Framework participant list. An active entity match can be relevant to the customer’s mechanism review, but does not by itself establish that the DPF applies to the observed processing. A service with an active entity match receives a lower penalty (×0.85); a confirmed inactive match receives a higher one (×1.15). Unknown status leaves the penalty unchanged.
Fonts / static assets×0.6
JS library / CDN×0.7
Analytics / error tracking×1.2
Payments / auth×1.3
Advertising / retargeting×1.6
Tag management×1.8
P_mix — non-EU ratio penalty (max −20)

When significantly non-EU services (US, China, Global) make up a large share of the classified service inventory, an additional penalty applies, up to −20. The penalty is scaled by sample size so a single finding on a short scan does not over-fire.

Example: 3 US-owned services, 1 EU-owned service = 75% non-EU ratio → P_mix ≈ 15

P_unknown — unmatched domains (−4 each, max −25)
Domains not found in our vendor database reduce both the score and our confidence in the result. High confidence requires that under 15% of third-party domains are unclassified.
P_infra — non-EU infrastructure (max −22)

A site can run its trackers from the EU yet still host itself, its email or its DNS on a non-EU-owned provider. That is a first-order sovereignty concern (CLOUD Act / FISA 702) independent of the third-party stack, so we resolve the origin hosting, email (MX) and authoritative DNS (NS) providers and classify each by ownership region.

Hosting−8
Email (MX)−6 (max −10)
DNS (NS)−3 (max −6)

Only significantly non-EU providers (US, China, Global, Unknown) count. When infrastructure can’t be resolved, P_infra is 0 — it never penalises a scan for missing data.

Label guardrails

Labels are not derived purely from the numeric score. Hard rules prevent misleading labels even when the score is numerically high. For example, Mostly EU independent is blocked if non-EU services outnumber EU/EEA services, regardless of the score.

EU-first stack90–100
Mostly EU independent75–89
EU-leaning, with dependencies60–74
Mixed third-party stack40–59
High non-EU dependency20–39
Heavily non-EU dependent0–19

The score card in your report shows a full breakdown so you can see exactly how each component contributed, along with a confidence indicator.

7. Monitoring, history & policy tracking

Paid plans re-scan each monitored site on a weekly schedule and turn the results into a durable history rather than a single snapshot. Every scan is diffed against the last one, and three things are recorded.

Service timeline

For each service we keep a first-seen and last-seen date, whether it is currently active, and whether it loaded before or after consent. A change timeline logs every service added, service removed and score change between scans, so you can point to the exact date a tracker appeared — useful evidence at an audit. The full history is browsable per site and exportable to CSV or JSON.

Post-reject signal history

Every weekly scan repeats the Reject-all network test. When a service first appears in the post-reject observation, we log a timestamped post-reject signal on the timeline and email you. When that signal later disappears, we record the change. This gives you dated before-and-after technical evidence from StackPatrol’s scheduled tests; it does not prove the exact moment a wider consent issue started or was fully resolved.

Legal-document tracking

On the same run we discover the site’s privacy policy, cookie policy, DPA, terms and GDPR pages. Links are classified from both the URL and the anchor text across multiple languages, and we deliberately follow the very common case where a site hosts these documents on a parent-company or group domain (for example a newspaper whose privacy policy lives on the publisher’s domain), while excluding social-network and search-engine platform policies. This crawl runs with a lightweight, size- and time-bounded fetch outside the page budget, so it never reduces the service-inventory page allowance.

For each document we normalise the text, take a SHA-256 content fingerprint, and extract the date the page states it was last updated (cue-based, multilingual, rejecting impossible or future dates). The report retains the document title, sanitized source URL, stated update date, fetch timestamp and fingerprint so the comparison can be verified. On the next monitoring scan we compare fingerprints to detect when the content actually changed — independently of whatever date the page claims.

Documentation-gap flag

The signal we find most useful: when your service inventory changes but the privacy (or cookie/DPA) document stays frozen, the alert flags a possible documentation gap. It is a prompt to review, not a legal finding — a policy can be correct without changing, and a change does not prove anything. You decide what needs updating.

Disclosure gap (Consent Assurance)

We go one step further than tracking when a document changed: we compare the disclosure sources we can inspect against the services the scan actually observed. These sources include readable policies and, in paid reports when available, the configured vendor list exposed by the site’s consent manager. Each observed service receives one of three evidence states. Exact service match means its own domain or distinctive service name appears in a checked source. Provider or parent match means an explicitly curated provider alias appears, but the specific service name does not. No match found means no service, domain or approved provider alias was found in a checked source.

Consent-manager inspection runs in a separate clean browser context and only navigates to the vendor view. It never saves or changes consent. StackPatrol has dedicated extraction adapters for Sourcepoint and Cookiebot vendor views, plus structured vendor/provider views exposed by OneTrust and Cookie Information installations. When an installation exposes categories but no configured vendor names, the inspection remains incomplete and contributes no matches. Vendor names are extracted deterministically; AI may only choose a numbered visible navigation control when deterministic labels fail. The report retains the sanitized source URL, inspection time, configured names and a SHA-256 fingerprint. A CMP entry is technical disclosure evidence, not proof that the legal disclosure is sufficient.

Provider or parent matches are qualified evidence and never count as exact disclosure. They come from explicit catalogue aliases rather than automatic inheritance from a corporate parent. Short or ambiguous names (“Segment”, “Forms”) and shared infrastructure domains (amazonaws.com, googleapis.com) are ignored to reduce false matches. Linked sub-processor registers and inaccessible consent-manager views can still be missed, so “no match found” means “worth checking”, never proof of non-disclosure. Free scans show the policy-based count; paid reports show the combined named evidence when CMP inspection succeeds.

If auto-discovery misses a document, you can pin exact policy URLs per monitored site and we track those instead.

Cookie declaration duration checks

When a readable cookie policy contains a structured table, we can compare an exact cookie name with a persistent cookie observed by the browser. We report a possible mismatch only when the observed lifetime clearly exceeds one unambiguous numeric declaration: more than 25% beyond the declared duration, plus one day. Session cookies, missing expiry evidence, ambiguous declarations and free-form policy prose are skipped.

The comparison covers names and durations only. It does not infer whether a stated purpose, provider or legal basis is correct. A name such as consentUUID is safe to show because it is the cookie’s label; the UUID value itself is never retained or displayed. A duration mismatch is a technical prompt to compare the cookie configuration and policy, not a legal finding.

Reject-all verdict (Consent Assurance)

Disclosure asks what a policy names; the Reject-all test observes what happens after a restrictive control is activated. On paid plans we open the site in a fresh browser context, activate a “Reject all” or “Only necessary” control, then revisit the entry page plus up to four baseline subpages and record third-party requests seen after control activation. The clicker mirrors the Accept-all automation but is biased toward the most restrictive explicit choice and never selects Accept or a generic Save control.

We only count genuinely non-essential trackers against the banner: advertising, analytics, product analytics, session replay, A/B testing, customer-data platforms, marketing automation, email marketing, affiliate and social embeds. Functional, infrastructure and tag-management services are deliberately excluded so a green or red verdict stays defensible. A tracker that fired on the very first load, before the visitor could reject, is not counted here — that belongs to the “loaded before consent” finding. Only requests observed after the reject click count as a violation.

The report separates four things that should not be conflated: whether a control was found, whether the UI interaction completed, whether stored CMP or TCF state changed, and what network traffic followed. A completed interaction does not by itself prove a valid rejection state. The test may report no banner, an unconfirmed control, a technical scan error, or a completed interaction with either quiet or continued non-essential traffic. Incomplete states are never presented as a pass or fail.

Only requests observed after the evidence boundary count in this test. Continued traffic is a network observation to review, not proof that a service ignored a legally valid rejection. Likewise, quiet traffic shows what this automated run observed; it does not certify the banner or the site as compliant.

8. What we don't do

  • We don’t determine GDPR or DSA compliance.
  • We don’t click consent banners on free scans (paid plans run an automatic post-consent pass).
  • We don’t crawl the entire site by default (front page only on Free, up to 5 pages on Pro and Agency Starter, up to 20 on Agency and Agency Pro).
  • We don’t log in to authenticated areas or fill out forms.
  • We don’t store IP addresses in plaintext; they are salted-hashed daily.
  • We don’t retain raw cookie values, raw localStorage or sessionStorage values, or raw TC strings in report evidence.

9. Limitations

Consent-gated scripts are a significant blind spot on the free scan. Many tracking scripts only load after a user accepts a consent banner. Paid plans automatically add a second pass that clicks “Accept all” on the consent banner and then revisits the entry page plus up to four baseline subpages in the same browser context. This catches services that only fire post-consent on article or product pages. The click uses vendor-specific selectors, multilingual button labels, iframe traversal for Sourcepoint-style CMPs, and an optional LLM fallback. The click is best-effort: when no banner is detected, or detected but not actionable, the report flags this explicitly so you know whether the scan saw the full service inventory.

Geographic bias also matters: some vendors serve different scripts based on the visitor’s country. We currently scan from a European IP, so results approximate what European visitors see.

The methodology evolves as we improve coverage. If you find a wrong classification, please let us know.