Block, error or bug: how production teams classify failed scrapes

A scraper that trusts the status code saves the wrong things. A block page can return a 200, and so can a page that doesn’t exist. At volume, each wrong label leads a team to the wrong fix.

2 talks at Prague Crawl 2026 showed how production teams handle this. I tested the core problem on 371 top sites, then built a classifier the way the talks described and measured it on the same sites.

What production teams do

Logan Harless (opens in a new tab) works on a team that collects “4 billion rows of data per week”. His point was that “the status codes, the content, and everything else you can get from servers can be a lie”. So every response is labeled with one of several terminal states, because each one needs a different action.

StateWhat the team does next
SuccessSend the data to the next step
BlockRetry with a different method, or escalate
Page moved or goneStop retrying
Server errorRetry later
Site format changedFix the parser

The classifier runs in steps, cheapest first. Simple rules on the status code, headers and page content come first. When they can’t decide, a more complex but still fast method is used. Formats that nothing recognizes go to a heavier model, and the team uses them to update the rules. In his words, “simple is generally best”, and “a feedback loop is going to beat some clever complex approach”.

Juan Manuel Pérez (opens in a new tab) leads a team of 36 engineers that runs more than 5,000 spiders on retail sites. At that size, even 1% of spiders failing means many bugs, and the team had a backlog of 300 tickets. Failures are triaged first. AI compares the run with earlier runs to find where it failed, whether that’s a proxy issue, the homepage, the navigation, product-page parsing or a block. Engineers then test a fix on a small run before a full one.

Google’s crawler works the same way. Its documentation (opens in a new tab) says a page with a 200 is still reported as a soft 404 when “the content suggests an error”, like an empty page or an error message.

What top sites return for a page that doesn’t exist

The simplest test of a classifier is a page that can’t exist. I took the 371 homepages from my test of the top 1,000 websites that returned content with a 200. For each one, I requested a random path like /doccraft-test-4f39159f8545. I ran it from Mumbai on October 9, 2026, with httpx 0.28.

Response to the missing pageSites
404, the correct answer317 (85%)
200, as if the page existed43 (12%)
Another status (403, 406, 400, 429, 202)11 (3%)

A scraper that checks only the status code records those 43 missing pages as successes. They weren’t all the same kind of page.

  • 14 redirected to the homepage. Steam and Dell sent the missing URL to their main page.
  • 6 said “not found” with a 200, like Kick’s “Channel Not Found”.
  • 4 redirected to a login page, like Microsoft and Threads.
  • 2 were challenge pages. Walmart returned “Robot or human?” with a 200.
  • 13 returned a page that matched the homepage. Instagram, Twitch and Telegram send the same HTML for every URL. On Instagram and Twitch, the page’s scripts decide what to show.
  • 4 others returned their own generic page.

A classifier built the same way

This classifier follows the order from the talks. It checks the status code first, then the page content and where the request landed. Last, it compares the page with the site’s own “not found” page, learned once from a random URL.

classify.py
import re
from urllib.parse import urlparse
NOT_FOUND_TEXT = re.compile(
r"not found|\b404\b|doesn.t exist|does not exist|no longer (exists|available)|nothing (was )?found",
re.I,
)
BLOCK_TEXT = re.compile(
r"captcha|robot or human|are you a robot|verify (that )?you are human|access denied|client challenge|just a moment",
re.I,
)
LOGIN_PATH = re.compile(r"/(login|signin|sign-in|auth|oauth2?)\b", re.I)
def path(url):
return urlparse(url).path.rstrip("/")
def looks_alike(a, b):
"""Same title and a similar amount of text: the same page template."""
return a["title"] == b["title"] and abs(a["words"] - b["words"]) <= max(5, 0.15 * max(a["words"], b["words"]))
def classify(resp, requested_url, home=None, missing=None):
"""resp, home and missing are page summaries: status, final_url, title, words, text.
home is the site's homepage; missing is its answer to a random URL that can't exist."""
status, top = resp["status"], resp["title"] + " " + resp["text"][:600]
if status in (404, 410):
return "not_found"
if status >= 500:
return "server_error"
if status in (403, 429) or BLOCK_TEXT.search(top) or (status == 202 and resp["words"] < 50):
return "block" # a 202 with an almost empty page is a challenge (AWS WAF answers bots this way)
if 400 <= status < 500:
return "client_error"
if LOGIN_PATH.search(path(resp["final_url"])) and not LOGIN_PATH.search(path(requested_url)):
return "login_required"
landed = path(resp["final_url"])
if home and landed != path(requested_url) and landed in ("", path(home["final_url"])):
return "not_found" # a deep URL that redirected to the homepage
if NOT_FOUND_TEXT.search(top):
return "not_found"
# The site's "not found" page, learned from one random URL. Skip it when that URL just
# redirected to the homepage: the redirect check above already covers those sites.
learned = missing and missing["status"] == 200 and not (home and path(missing["final_url"]) == path(home["final_url"]))
if learned and looks_alike(resp, missing):
# If the homepage looks the same too, the site serves the same HTML for every URL,
# and the HTML alone can't say whether this page exists.
return "unclear" if home and looks_alike(home, missing) else "not_found"
return "success"

To use it, fetch the site’s homepage and one random missing URL once, then classify each page you scrape against them. The full script does this for 4 real URLs, and all 4 returned a 200.

Output
200 -> success https://stripe.com/pricing
200 -> not_found https://store.steampowered.com/app/9999999999/
200 -> not_found https://kick.com/no-such-channel-4821
200 -> unclear https://www.instagram.com/no-such-profile-4821/

How well it worked

I ran the classifier on all 371 homepages and all 371 missing pages. The “not found” fingerprint for each site came from a second random URL, not the one being classified.

Status code onlyClassifier
Missing pages labeled as a success431
Missing pages labeled “unclear”013
Homepages labeled as a success371357

Of the 43 missing pages that returned a 200, the classifier labeled 23 as not found, 4 as a login redirect and 2 as a block. Only 1 was still labeled a success, a generic page with no error message.

The 13 “unclear” sites are the honest limit of reading HTML. Their homepages and missing pages are the same HTML, so the classifier marks both as unclear instead of guessing. For those sites, render the page in a browser, or check for the data you expected, like a product price or a channel name. That’s the job of the heavier model in Logan’s talk.

One homepage was labeled a block by mistake. hCaptcha’s own homepage contains the word “CAPTCHA”. Text rules fail on sites that are about the thing they detect, which is why production teams keep a review loop.

Limits of this test

  • Homepages and one made-up path per site. Real crawls see more states, like changed layouts and partial pages, which need checks for the specific data you collect.
  • Text rules in English. Sites that say “not found” in another language need their own patterns, or the fingerprint check.
  • One location and one client. Some sites answer differently by country, or to browsers and proxies.

Data and code

Doccraft writes research and tutorials like this for web data and AI infrastructure companies.