GSC reported 81 pages "Crawled – currently not indexed" and "2 server errors" on a SvelteKit 2.69.1 / adapter-node 5.2.12 app behind Render. Uptime monitoring was green and every page loaded fine by hand. Crawling the site's own sitemap revealed the app sheds ~2% of requests with 5xx under even a gentle concurrent crawl.
The technique
Fetch your sitemap, then request every URL concurrently and tally status codes:
import urllib.request, ssl, re, concurrent.futures as cf
from collections import Counter
ctx = ssl.create_default_context()
xml = urllib.request.urlopen('https://example.com/sitemap.xml', context=ctx).read().decode()
urls = re.findall(r'<loc>([^<]+)</loc>', xml)
def check(u):
req = urllib.request.Request(u, method='GET',
headers={'User-Agent': 'Mozilla/5.0 (compatible; SEOCheck/1.0)'})
try:
with urllib.request.urlopen(req, context=ctx, timeout=30) as r:
r.read(1) # touch body; don't download it all
return r.status
except urllib.error.HTTPError as e:
return e.code
except Exception as e:
return f'ERR:{type(e).__name__}'
for workers in (4, 16):
with cf.ThreadPoolExecutor(max_workers=workers) as ex:
print(workers, Counter(str(c) for c in ex.map(check, urls)))Two traps that cause misdiagnosis
1. Don't audit with HEAD. My first pass used HEAD and returned six 500s. I nearly filed them as six broken pages. All six returned 200 on an isolated GET. Some frameworks and proxies handle HEAD differently, and more importantly a single failure sample tells you nothing about whether the URL or the moment was broken.
2. One pass is not a measurement. Results across four passes on an unchanged 362-URL sitemap:
| 1 | HEAD | 16 | 6 | 6x 500 |
| 2 | GET | 4 | 7 | 5x 500, 2x 504 |
| 3 | GET | 16 | 10 | 9x 500, 1x 502 |
| 4 | GET | 4 | 2 | 2x 500 |
Different URLs failed each pass and every failing URL returned 200 on retry. That signature — random victims, retry-clean, rate scaling with concurrency — means load shedding, not broken pages. If the same URLs failed every pass you would have a data bug instead.
Why it matters beyond latency
This class of bug is usually filed as a performance nit because a human never sees it. But Google throttles crawl rate and withholds indexing from hosts that intermittently 5xx, so a 2% shed rate surfaces in Search Console as "Crawled – currently not indexed" — a bucket most people go triage URL-by-URL. Check the host-level 5xx rate before triaging individual pages; the 81 pages may be one problem, not 81.
Localize it before fixing
Bucket failures by route shape and compare against the sitemap's composition. In our case one route family was 18% of the sitemap but produced nearly all the failures, which pointed at the heaviest SSR route rather than at infrastructure sizing. The upstream cause was a known duplicate-fetch N+1 in the layout chain ([[SvelteKit SSR page hangs (2-3s) caused by child +layout.ts re-fetching the same API endpoint already fetched by a parent +layout.ts. SvelteKit's handleFetch forwards cookies to the API, so authenticat]]): a child layout awaited event.parent() and then re-fetched two endpoints the parent had already fetched, so the route ran 8 API calls in 2 serial waves with 25% pure waste. The latency bug and the indexing bug were the same bug.