Block, error or bug: how production teams classify failed scrapes
A scraper that trusts the status code saves the wrong things. A block page can return a 200, and so can a page that doesn’t exist. At volume, each wrong label leads a team to the wrong fix.
2 talks at Prague Crawl 2026 showed how production teams handle this. I tested the core problem on 371 top sites, then built a classifier the way the talks described and measured it on the same sites.
What production teams do
Logan Harless (opens in a new tab) works on a team that collects “4 billion rows of data per week”. His point was that “the status codes, the content, and everything else you can get from servers can be a lie”. So every response is labeled with one of several terminal states, because each one needs a different action.
| State | What the team does next |
|---|---|
| Success | Send the data to the next step |
| Block | Retry with a different method, or escalate |
| Page moved or gone | Stop retrying |
| Server error | Retry later |
| Site format changed | Fix the parser |
The classifier runs in steps, cheapest first. Simple rules on the status code, headers and page content come first. When they can’t decide, a more complex but still fast method is used. Formats that nothing recognizes go to a heavier model, and the team uses them to update the rules. In his words, “simple is generally best”, and “a feedback loop is going to beat some clever complex approach”.
Juan Manuel Pérez (opens in a new tab) leads a team of 36 engineers that runs more than 5,000 spiders on retail sites. At that size, even 1% of spiders failing means many bugs, and the team had a backlog of 300 tickets. Failures are triaged first. AI compares the run with earlier runs to find where it failed, whether that’s a proxy issue, the homepage, the navigation, product-page parsing or a block. Engineers then test a fix on a small run before a full one.
Google’s crawler works the same way. Its documentation (opens in a new tab) says a page with a 200 is still reported as a soft 404 when “the content suggests an error”, like an empty page or an error message.
What top sites return for a page that doesn’t exist
The simplest test of a classifier is a page that can’t exist. I took the 371 homepages from my test of the top 1,000 websites that returned content with a 200. For each one, I requested a random path like /doccraft-test-4f39159f8545. I ran it from Mumbai on October 9, 2026, with httpx 0.28.
| Response to the missing page | Sites |
|---|---|
| 404, the correct answer | 317 (85%) |
| 200, as if the page existed | 43 (12%) |
| Another status (403, 406, 400, 429, 202) | 11 (3%) |
A scraper that checks only the status code records those 43 missing pages as successes. They weren’t all the same kind of page.
- 14 redirected to the homepage. Steam and Dell sent the missing URL to their main page.
- 6 said “not found” with a 200, like Kick’s “Channel Not Found”.
- 4 redirected to a login page, like Microsoft and Threads.
- 2 were challenge pages. Walmart returned “Robot or human?” with a 200.
- 13 returned a page that matched the homepage. Instagram, Twitch and Telegram send the same HTML for every URL. On Instagram and Twitch, the page’s scripts decide what to show.
- 4 others returned their own generic page.
A classifier built the same way
This classifier follows the order from the talks. It checks the status code first, then the page content and where the request landed. Last, it compares the page with the site’s own “not found” page, learned once from a random URL.
import refrom urllib.parse import urlparse
NOT_FOUND_TEXT = re.compile( r"not found|\b404\b|doesn.t exist|does not exist|no longer (exists|available)|nothing (was )?found", re.I,)BLOCK_TEXT = re.compile( r"captcha|robot or human|are you a robot|verify (that )?you are human|access denied|client challenge|just a moment", re.I,)LOGIN_PATH = re.compile(r"/(login|signin|sign-in|auth|oauth2?)\b", re.I)
def path(url): return urlparse(url).path.rstrip("/")
def looks_alike(a, b): """Same title and a similar amount of text: the same page template.""" return a["title"] == b["title"] and abs(a["words"] - b["words"]) <= max(5, 0.15 * max(a["words"], b["words"]))
def classify(resp, requested_url, home=None, missing=None): """resp, home and missing are page summaries: status, final_url, title, words, text. home is the site's homepage; missing is its answer to a random URL that can't exist.""" status, top = resp["status"], resp["title"] + " " + resp["text"][:600] if status in (404, 410): return "not_found" if status >= 500: return "server_error" if status in (403, 429) or BLOCK_TEXT.search(top) or (status == 202 and resp["words"] < 50): return "block" # a 202 with an almost empty page is a challenge (AWS WAF answers bots this way) if 400 <= status < 500: return "client_error" if LOGIN_PATH.search(path(resp["final_url"])) and not LOGIN_PATH.search(path(requested_url)): return "login_required" landed = path(resp["final_url"]) if home and landed != path(requested_url) and landed in ("", path(home["final_url"])): return "not_found" # a deep URL that redirected to the homepage if NOT_FOUND_TEXT.search(top): return "not_found" # The site's "not found" page, learned from one random URL. Skip it when that URL just # redirected to the homepage: the redirect check above already covers those sites. learned = missing and missing["status"] == 200 and not (home and path(missing["final_url"]) == path(home["final_url"])) if learned and looks_alike(resp, missing): # If the homepage looks the same too, the site serves the same HTML for every URL, # and the HTML alone can't say whether this page exists. return "unclear" if home and looks_alike(home, missing) else "not_found" return "success"To use it, fetch the site’s homepage and one random missing URL once, then classify each page you scrape against them. The full script does this for 4 real URLs, and all 4 returned a 200.
200 -> success https://stripe.com/pricing200 -> not_found https://store.steampowered.com/app/9999999999/200 -> not_found https://kick.com/no-such-channel-4821200 -> unclear https://www.instagram.com/no-such-profile-4821/How well it worked
I ran the classifier on all 371 homepages and all 371 missing pages. The “not found” fingerprint for each site came from a second random URL, not the one being classified.
| Status code only | Classifier | |
|---|---|---|
| Missing pages labeled as a success | 43 | 1 |
| Missing pages labeled “unclear” | 0 | 13 |
| Homepages labeled as a success | 371 | 357 |
Of the 43 missing pages that returned a 200, the classifier labeled 23 as not found, 4 as a login redirect and 2 as a block. Only 1 was still labeled a success, a generic page with no error message.
The 13 “unclear” sites are the honest limit of reading HTML. Their homepages and missing pages are the same HTML, so the classifier marks both as unclear instead of guessing. For those sites, render the page in a browser, or check for the data you expected, like a product price or a channel name. That’s the job of the heavier model in Logan’s talk.
One homepage was labeled a block by mistake. hCaptcha’s own homepage contains the word “CAPTCHA”. Text rules fail on sites that are about the thing they detect, which is why production teams keep a review loop.
Limits of this test
- Homepages and one made-up path per site. Real crawls see more states, like changed layouts and partial pages, which need checks for the specific data you collect.
- Text rules in English. Sites that say “not found” in another language need their own patterns, or the fingerprint check.
- One location and one client. Some sites answer differently by country, or to browsers and proxies.