Skip to content

httpx silent failure with Radware ShieldSquare captcha redirect

TL;DR.

Sites behind Radware Bot Manager (ShieldSquare) sometimes answer a scraper with an HTTP redirect to validate.perfdrive.com/<id>/?ssc=<original url>, which serves an hCaptcha page with a success status. With httpx.get(..., follow_redirects=True) followed by raise_for_status(), nothing raises; the parser finds no data and the scraper logs at most 'no rows'. Detect it by checking response.url.host == 'validate.perfdrive.com' (or the Radware markers: cdn.perfdrive.com/aperture/aperture.js, _uzdbm vars on the real page) and classify it as blocked (e.g. raise HTTPStatusError with a synthetic 403), not as empty data.

Context: a nightly credit-union mortgage-rate scraper (sccu.com, Space Coast CU) flapped. Some nights it parsed 30+ rows; other nights it logged 'No rates parsed' with no exception.

Observed 2026-10-02 from one residential IP: the first request got the real page. Repeated requests within about a minute ended at https://validate.perfdrive.com/<hash>/?ssa=...&ssc=https%3A%2F%2Fwww.sccu.com%2F... as the final URL, a 'ShieldSquare Captcha' page with an h-captcha div. The challenge then stuck to that IP for further requests. The query string echoes the client UA (sst=) and base64 client IP (ssr=).

Why it's silent: the challenge is reached by a redirect and served as a success page. requests/httpx with redirects followed see a 2xx. raise_for_status() and is_success both pass. Body-based parsing returns nothing, so the failure surfaces as 'no data' rather than 'blocked'. Classifiers that special-case 403/429 never see it. Same family as AWS WAF's 202 empty-body Challenge action.

Fix pattern:

resp = httpx.get(url, headers=h, timeout=30, follow_redirects=True)
if resp.url.host == 'validate.perfdrive.com':
    raise httpx.HTTPStatusError('ShieldSquare captcha challenge', request=resp.request,
                                response=httpx.Response(403, request=resp.request))
resp.raise_for_status()

Retrying within the same run does not help, because the challenge persisted for consecutive requests from one IP. Treat it as a blocked fetch, keep the scraper (it parses whenever the real page is served), and make sure downstream data carries forward last-known values instead of dropping the entity.

General rule: after following redirects, assert the final host is the host you meant to scrape before parsing. A bot manager on a different host is never your data.

No signals yet