๐ŸŽ‰ Premium Proxies ยท 24-Hour Free TrialClaim Now
Guides

How to Identify Which Anti-Bot Protection a Site Uses

Identify anti-bot protection fast: the cookie names, response headers, block pages and challenge scripts that name Cloudflare, Akamai, DataDome and Kasada.

S SparkProxy 4 19 min read
Share

Your scraper is returning 403, and the first instinct is to reach for a stealth plugin. Skip that. The most useful thing you can do is identify anti-bot protection by name first, because a technique that walks past Cloudflare does nothing against Kasada, and the fix for a plain IP deny rule is the opposite of the fix for a JavaScript challenge. Every major vendor leaves its name in the response: in a cookie, a header, a block-page string, or the path of the script it injects. This guide shows you where those fingerprints live, gives you a lookup table for the nine vendors you'll actually meet, and covers the case most guides skip, where two vendors are stacked on the same host.

If you aren't yet sure your requests are being blocked at all rather than simply failing, start with how to detect when your scraper is blocked and come back once you've confirmed it.

Why the vendor name comes before the bypass

Anti-bot products fail differently, so they have to be handled differently.

Cloudflare's managed challenge is mostly a browser-environment test: give it a real browser with a consistent TLS and JavaScript fingerprint and you're usually through. Kasada runs a client-side proof-of-work, so a real browser is mandatory and no amount of header tuning substitutes for the CPU cycles it costs. DataDome leans hard on IP reputation and behaviour, which means the browser can be perfect and a datacenter IP will still lose. Akamai scores a long-lived _abck cookie across a whole session, so a strategy of rotating your IP on every request actively hurts you there.

Pick the wrong model and you'll spend days tuning the one variable nobody is measuring. That's the real cost of guessing.

Here is the same point as a decision table. The middle column is the signal each product weights most heavily, and the right column is what to change first once you've named it:

VendorWeights most heavilyChange this first
CloudflareBrowser environment, TLS and JA3/JA4 consistencyRun a real browser with a matching TLS fingerprint
Akamai Bot ManagerSession continuity via the `_abck` cookieHold one IP and one session, stop rotating per request
DataDomeIP reputation and request behaviourMove to residential or ISP addresses, then pace requests
KasadaClient-side proof-of-work computeUse a real browser engine, no HTTP client will do
PerimeterX / HUMANBehavioural telemetry, mouse and scrollWarm the session with real interaction before deep crawling
Imperva IncapsulaHeader order and TLS fingerprintFix header ordering and use a browser-grade TLS stack
F5 / BIG-IP Bot DefenseObfuscated JS telemetry from the pageExecute the injected script in a real browser
AWS WAFRate and rule matching, token possessionSlow down and carry the `aws-waf-token` correctly
SucuriIP reputation and simple rule matchingChange IP range, no JavaScript work needed

Two rows in that table point in opposite directions, which is the clearest argument for identifying before acting. Against DataDome, rotating to a fresh IP is often the fix. Against Akamai, rotating is the thing that breaks you, because the _abck cookie is scored over the life of a session and a new IP mid-session reads as a hijack. The same action helps in one case and causes the block in the other.

There's a second reason to name the vendor, and it matters more than it sounds. A 403 with no vendor fingerprint at all usually isn't a managed anti-bot. It's an origin-level deny rule on your IP or your ASN, which is a different problem with a much simpler fix. Knowing the difference saves you from building a headless browser farm to solve something that only needed cleaner address space.

The three-minute triage

Four signals, ordered by how quickly they pay off:

  1. Response headers. Vendor-specific headers are the fastest tell and the hardest to strip without breaking the product.
  2. Cookie names. Nearly every vendor sets at least one uniquely named cookie, even on a block.
  3. Block-page body. Error pages carry incident IDs, vendor text and support URLs.
  4. Injected script path. The challenge script's URL names the vendor, and sometimes the tier.

Collect all four with one request and read the output instead of guessing:

curl -s -D - -o /tmp/body.html \
  -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36' \
  'https://target.example/some-page'

-D - prints response headers to stdout and -o keeps the body for inspection. Use a realistic User-Agent, because some vendors serve a generic 403 to obvious tooling and a named challenge to a browser-shaped request. A bare curl/8.x User-Agent can hide the very fingerprint you're trying to read.

If you check targets regularly, wrap the first four signals in one function and stop doing it by hand. This returns every vendor it can find rather than the first match, which is what you want for the stacked case further down:

import re
import requests

SIGNATURES = {
    "Cloudflare":        {"cookies": [r"^__cf_bm$", r"^cf_clearance$", r"^__cfruid$"],
                          "headers": ["cf-ray"],
                          "body":    [r"attention required", r"challenges\.cloudflare\.com"]},
    "Akamai":            {"cookies": [r"^_abck$", r"^ak_bmsc$", r"^bm_sz$", r"^bm_sv$"],
                          "headers": [], "body": [r"reference #[0-9a-f.]+"]},
    "DataDome":          {"cookies": [r"^datadome$", r"^dd_cookie_test$"],
                          "headers": ["x-datadome", "x-datadome-cid"],
                          "body":    [r"captcha-delivery\.com", r"\bdd\.js\b"]},
    "PerimeterX/HUMAN":  {"cookies": [r"^_px", r"^_pxhd$", r"^_pxvid$"],
                          "headers": [],
                          "body":    [r"access to this page has been denied", r"perimeterx"]},
    "Imperva Incapsula": {"cookies": [r"^incap_ses", r"^visid_incap", r"^nlbi_"],
                          "headers": ["x-iinfo"],
                          "body":    [r"incapsula incident id", r"_incapsula_resource"]},
    "Kasada":            {"cookies": [r"^KPSDK"],
                          "headers": ["x-kpsdk-ct", "x-kpsdk-cd", "x-kpsdk-r"],
                          "body":    []},
    "F5 BIG-IP":         {"cookies": [r"^TSPD_101", r"^TS[0-9a-f]{6,}$"],
                          "headers": [], "body": []},
    "AWS WAF":           {"cookies": [r"^aws-waf-token$"],
                          "headers": ["x-amzn-waf-action"],
                          "body":    [r"token\.awswaf\.com"]},
    "Sucuri":            {"cookies": [r"^sucuri_cloudproxy_uuid"],
                          "headers": ["x-sucuri-id", "x-sucuri-cache"],
                          "body":    [r"sucuri website firewall"]},
}

UA = ("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
      "(KHTML, like Gecko) Chrome/145.0.0.0 Safari/537.36")


def fingerprint(url, timeout=30):
    r = requests.get(url, headers={"User-Agent": UA}, timeout=timeout)
    cookie_names = {c.name for h in list(r.history) + [r] for c in h.cookies}
    header_keys = {k.lower() for k in r.headers}
    body = r.text[:200_000].lower()

    found = []
    for vendor, sig in SIGNATURES.items():
        if (any(re.search(p, n) for p in sig["cookies"] for n in cookie_names)
                or any(h in header_keys for h in sig["headers"])
                or any(re.search(p, body) for p in sig["body"])):
            found.append(vendor)
    return {"status": r.status_code, "vendors": found,
            "cookies": sorted(cookie_names)}


print(fingerprint("https://target.example/"))

A result of {"status": 403, "vendors": [], ...} is a finding, not a failure. That combination is the origin deny rule described above, and the sections below explain how to read each signal the function collects.

Free trial

Scraping at scale? Skip the blocks.

Fast, unblockable datacentre proxies with unlimited bandwidth.

Step 1: dump the response headers

Headers are where the vendor advertises itself. Pull out the ones that matter:

curl -sI 'https://target.example/' | grep -iE 'server|x-cdn|x-iinfo|x-datadome|x-kpsdk|x-amzn-waf|x-sucuri|cf-ray|akamai|via|set-cookie'

What you're looking for:

  • cf-ray and server: cloudflare put the request behind Cloudflare. On its own this only proves Cloudflare is the CDN, not that bot management is switched on.
  • x-iinfo, or x-cdn: Incapsula (sometimes x-cdn: Imperva), means Imperva.
  • x-datadome or x-datadome-cid means DataDome.
  • x-kpsdk-ct, x-kpsdk-cd or x-kpsdk-r means Kasada.
  • x-amzn-waf-action means AWS WAF.
  • x-sucuri-id and x-sucuri-cache mean Sucuri.
  • server: AkamaiGHost, or an akamai-prefixed header, points at Akamai.

One caveat on server headers. Cloudflare and Sucuri announce themselves reliably, Akamai often does, and Imperva frequently strips or rewrites it. Absence proves nothing. Treat server as a hint and the cookies as the evidence.

The vendor signature lookup table

VendorCookiesResponse headersChallenge host or script
Cloudflare`__cf_bm`, `cf_clearance`, `__cfruid``cf-ray`, `server: cloudflare``challenges.cloudflare.com/turnstile/v0/api.js`
Akamai Bot Manager`_abck`, `ak_bmsc`, `bm_sz`, `bm_sv``server: AkamaiGHost`inline sensor script, often under `/akam/`
DataDome`datadome`, `dd_cookie_test``x-datadome`, `x-datadome-cid``geo.captcha-delivery.com`, `ct.captcha-delivery.com`, `dd.js`
PerimeterX / HUMAN`_px3`, `_pxhd`, `_pxvid`, `_px2`varies, often none`client.perimeterx.net`, `/px/init.js`
Imperva Incapsula`incap_ses_*`, `visid_incap_*`, `nlbi_*``x-iinfo`, `x-cdn: Incapsula``_Incapsula_Resource` in script or iframe tags
Kasada`KPSDK`-prefixed`x-kpsdk-ct`, `x-kpsdk-cd`, `x-kpsdk-r`polymorphic `p.js` under a UUID-shaped path
F5 / BIG-IP Bot Defense`TS`, `TSPD_101*`often strippedinjected obfuscated JS, no stable host
AWS WAF`aws-waf-token``x-amzn-waf-action``*.token.awswaf.com/challenge.js`
Sucuri`sucuri_cloudproxy_uuid_*``x-sucuri-id`, `x-sucuri-cache`no JS challenge, serves a block page

Two rows need a caveat. F5's BIG-IP sets TS-prefixed cookies across several of its application-security modules, so TS01a2b3c4 alone only tells you BIG-IP is in front of the origin; the TSPD_101 variant is the one tied to Bot Defense specifically. And Kasada randomises its script path on purpose, so match the x-kpsdk-* headers rather than trying to pin a URL.

That leads to the thing most diagnoses get wrong, which is confusing presence with enforcement. Four false positives account for nearly all of it:

  1. Cloudflare is present but bot management is off. cf-ray and __cf_bm appear on a huge share of the web, because __cf_bm ships with Bot Fight Mode and basic bot handling on plans that have no managed challenge configured at all. If you're getting 200 responses with complete HTML, Cloudflare is proxying the site and nothing more. Don't build a browser farm for a site that isn't challenging you.
  2. A CDN banner is not a bot product. server: AkamaiGHost means Akamai delivers the content. Akamai Bot Manager is a separate product on top of that. The _abck cookie, not the server header, is what tells you Bot Manager is switched on.
  3. TS cookies from plain load balancing. BIG-IP uses them for session persistence as well as for security modules, so a generic TS cookie on an enterprise site often means nothing more than a hardware load balancer.
  4. A CAPTCHA on one route only. Turnstile or hCaptcha on /login or a contact form says nothing about /products. Fingerprint the exact path you intend to scrape, not the homepage, because protection is routinely applied per route and per tier.

The reliable test for enforcement is behavioural rather than structural: request the path you actually want, from the IP range you actually plan to use, and see whether the content arrives. Fingerprints tell you who is standing at the door. Only a real request tells you whether the door is locked.

Step 3: read the block page

When headers are stripped, the body usually still talks. Grep it for vendor strings:

grep -ioE 'incapsula incident id|_incapsula_resource|powered by imperva|access to this page has been denied|perimeterx|sucuri website firewall|attention required|datadome|captcha-delivery|reference #[0-9a-f.]+' /tmp/body.html | sort -u

What the distinctive strings mean:

  • Incapsula incident ID, usually inside an iframe, plus subject=WAF Block Page in the HTML: Imperva.
  • Attention Required! | Cloudflare with a ray ID in the footer: Cloudflare.
  • Access to this page has been denied: the classic PerimeterX wording.
  • Access Denied - Sucuri Website Firewall: Sucuri.
  • Reference # followed by a dotted hex string: Akamai's error format.

Keep those incident and reference IDs. If you're scraping a site you have a relationship with, quoting the incident ID to their support team gets a block reviewed far faster than describing the symptom does.

Step 4: find the challenge script

If the page rendered but the content is missing, the challenge is running in JavaScript. List the script sources:

grep -oE '<script[^>]+src="[^"]+"' /tmp/body.html | sed -E 's/.*src="([^"]+)".*/\1/' | sort -u

CAPTCHA widgets are unambiguous, and they tell you which pathway applies:

  • challenges.cloudflare.com/turnstile/v0/api.js: Cloudflare Turnstile.
  • www.google.com/recaptcha/api.js: reCAPTCHA.
  • js.hcaptcha.com/1/api.js: hCaptcha.
  • client-api.arkoselabs.com: Arkose Labs FunCaptcha.
  • static.geetest.com: GeeTest.

A CAPTCHA widget and a bot-management product sit at different layers, and plenty of sites run both. Turnstile on the login form tells you nothing about what guards the product pages.

What the status code narrows down

The code alone won't name a vendor, but it rules options out.

CodeUsual meaningWhat it suggests
`403`Request refused outrightVendor block, or an origin IP/ASN deny rule if no fingerprints are present
`429`Rate or challenge responseKasada uses this for its challenge, so check `x-kpsdk-*` before assuming rate limiting
`503`Interstitial or challenge pageHistorically Cloudflare's challenge code; read the body before calling it an outage
`200`Page served, content missingA JavaScript challenge replaced the content, go to step 4
`401`Authentication requiredUsually a real auth wall, not an anti-bot

That 429 row catches people out. A 429 carrying x-kpsdk-ct is a Kasada proof-of-work challenge, not a request to slow down, and backing off your rate won't clear it.

Step 5: confirm with a real browser

Some vendors only reveal themselves once JavaScript executes. Load the page in a real browser and watch where it talks:

from playwright.sync_api import sync_playwright

VENDOR_HOSTS = ("perimeterx", "captcha-delivery", "datadome", "awswaf",
                "kasada", "challenges.cloudflare.com", "arkoselabs",
                "hcaptcha", "recaptcha", "geetest", "incapsula")

def log_vendor_call(request):
    if any(host in request.url for host in VENDOR_HOSTS):
        print("vendor call:", request.url[:120])

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.on("request", log_vendor_call)
    page.goto("https://target.example/", wait_until="networkidle")
    print("cookies:", sorted({c["name"] for c in page.context.cookies()}))
    browser.close()

The cookie set after a real page load is the complete picture. A bare HTTP request often sees only the edge layer, because the application-layer product doesn't engage until its script runs.

Bear in mind that an unpatched headless browser is itself detectable, so a vendor may serve you a challenge it wouldn't serve a normal visitor. See headless browser detection for what gives automation away.

You can run the same check through an API that renders the page for you, which helps when your own IPs are already burned and you want a clean read. Mirroring the target's status code and asking for the JSON envelope gives you headers alongside the body:

import requests

r = requests.get(
    "https://scrape.sparkproxy.io/api/v1",
    headers={"X-API-Key": "YOUR_API_KEY"},
    params={
        "url": "https://target.example/",
        "render_js": "true",
        "transparent_status_code": "true",
        "json_response": "true",
    },
    timeout=120,
)
print(r.json())

transparent_status_code passes the target's real status through instead of the API's own, so a 403 stays a 403 and you can see what the site actually returned. The full parameter list is in the Scraping API docs.

When two vendors are stacked

Most guides assume one vendor per site. In practice a CDN-layer product and an application-layer product on the same host is common, and the combination changes what you have to solve.

A typical stack sets __cf_bm from Cloudflare at the edge and _px3 from PerimeterX at the application layer. Both cookies show up in a merged cookie jar, which tells you nothing about which one refused you. The per-hop output from step 2 does:

  • Vendor cookies on the first hop, then a redirect, then a challenge on a later hop: the edge let you through and the application layer stopped you. Solve the application-layer product.
  • A challenge on the first response, before any redirect: the edge stopped you and the application layer never ran. Nothing behind it matters yet.

There's one more distinction, and it's the one that should change your fix. Look at whether the blocking response sets a fresh challenge cookie or rejects a token you were already carrying:

  • A fresh Set-Cookie on the block means you were never admitted. That's a cold-start problem: fingerprint, IP reputation, or a challenge you didn't solve.
  • A rejection while you held a valid token means your session was retired mid-flight. That's a behavioural or rate signal, and rotating identity makes it worse rather than better. Slow down and keep the session.

Finally, the negative result, which is just as informative. If you see a 403 with no vendor cookies, no vendor headers, no challenge script and a short generic body, you're almost certainly looking at an origin or firewall deny rule keyed to your IP or ASN rather than a bot-management product. Browser automation will not fix that. Cleaner address space will. Datacenter ranges with poor reputation get caught by exactly this kind of blunt rule, which is covered in how to bypass IP fingerprinting with clean datacenter subnets.

Where to go once you know the vendor

Each product has its own guide:

Two cross-cutting pieces apply whatever the vendor, because nearly every product listed scores both: TLS fingerprinting and browser fingerprinting.

Re-run the triage after any change you make. Vendors get swapped, tiers get upgraded, and a site that was plain Cloudflare last quarter may have an application-layer product in front of its pricing pages today. The fingerprint is cheap to re-read. The stale assumption is expensive to carry.

Frequently asked questions

FAQ

Request a harmless page such as the homepage or /robots.txt with a realistic User-Agent, then read the Set-Cookie names and response headers. Most products set their cookie on the first successful response, so you can identify anti-bot protection before you ever trigger a block.

Yes, and it's common. A CDN-layer product like Cloudflare and an application-layer product like PerimeterX or DataDome frequently coexist. Print cookies per redirect hop rather than from a merged jar to see which layer actually refused the request.

That pattern usually points to an origin or firewall deny rule on your IP or ASN rather than a bot-management product. It's the one case where a better browser won't help and cleaner IP space will.

No. Kasada serves its proof-of-work challenge as a 429, so check for the x-kpsdk-ct and x-kpsdk-cd headers first. If they're present, slowing your request rate won't clear the block, because it isn't a rate problem.

Not on its own. Cloudflare and Sucuri announce themselves there consistently, but Imperva often strips or rewrites it, and a CDN banner only proves the CDN is present, not that its bot management is enabled. Cookie names are the stronger evidence.

Usually not. Headers, cookie names and the block-page body answer it for most sites. A real browser matters when the page returns 200 with content missing, because the application-layer product only engages once its JavaScript executes.

Special Discount ยท 20% off

Get 20% off your first month

Premium datacentre proxies with unlimited bandwidth. Use the code at checkout.

Save up to 15% more on quarterly, half-yearly and yearly plans

Claim Discount

About the Author

The SparkProxy Technical Team builds and operates SparkProxy's datacenter and residential proxy networks and its Scraping API. We spend a lot of time reading block pages, both from the scraping side and from running infrastructure that gets scraped, and the fingerprints above come from traffic we handle daily. Our guides document what the request and response actually contain, not what a vendor's marketing page claims.

Keep reading

Related articles