Why a 200 response can still be a block
A status code of 200 tells you the server answered. It doesn’t tell you the answer holds the data you asked for. Many blocked requests come back with a 200, and a scraper that trusts the status code saves them as if they were real pages.
This post shows the forms these hidden blocks take, and a short check that catches them.
A real example
I requested Amazon’s homepage with Python’s httpx library and a descriptive user agent. The response had a status of 200, but it was 3,781 bytes long and its only text was this.
Amazon.com Click the button below to continue shopping Continue shopping Conditions of Use Privacy PolicyA real Amazon homepage is hundreds of kilobytes. This page is a gate that asks a person to click a button, and it contains none of the words that usually signal a block, like “CAPTCHA” or “robot”.
The same homepage behaved differently in my 1,000-site test a day earlier. With a Chrome user agent, several Amazon stores returned a 2,183-byte challenge page with a status of 200. Today, the same request gets a 202 with an empty body. The block changes shape from day to day, so a scraper can’t rely on one signal to spot it.
The forms a hidden block takes
A challenge page with a status of 200. The page asks the visitor to prove they’re human, and it often needs JavaScript to continue. Cloudflare, AWS WAF, DataDome and HUMAN each have their own version.
A success code that isn’t 200. A 202 means “accepted”, and a scraper may treat it as success. Amazon’s current block uses it, with an empty body.
A real page with less data. The Scrape.do team described this in a Reddit AMA. A response “looks successful but the content is quietly degraded”, for example “if a page returns 5 records instead of 10”.
A real page with wrong data. At ZyteCon 2026, Zyte’s data team described honeypots, where “you get a 200 but it’s the wrong data”. Some sites serve false prices or details to clients they think are bots.
A block from a scraping API. At Prague Crawl 2026, Logan Harless said that well-known unblocking services sometimes return a 200 “when really what they’re serving you is a CAPTCHA”. A 200 from a vendor needs the same check as a 200 from a site.
A check that catches them
The most reliable check asks whether the response holds the data you expected. Known block markers catch the common challenge pages, and a count of the items you came for catches the rest.
import httpxfrom bs4 import BeautifulSoup
BLOCK_SIGNS = [ "just a moment", # Cloudflare challenge "challenge-platform", # Cloudflare challenge script "verify that you're not a robot", # AWS WAF challenge "captcha-delivery", # DataDome "px-captcha", # HUMAN (PerimeterX)]
def check(response, selector, expected): """Say whether a response holds the data you asked for, not just a status code.""" if response.status_code != 200: return f"failed: status {response.status_code}, {len(response.content)} bytes" text = response.text.lower() for sign in BLOCK_SIGNS: if sign in text: return f"blocked: challenge page ({sign!r}) with status 200" found = len(BeautifulSoup(response.text, "html.parser").select(selector)) if found < expected: return f"incomplete: {found} of {expected} expected items" return f"ok: {found} items"
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}pages = [ ("https://quotes.toscrape.com/", "div.quote", 10), ("https://quotes.toscrape.com/js/", "div.quote", 10), ("https://www.amazon.com/", "#nav-logo", 1),]for url, selector, expected in pages: response = httpx.get(url, headers=headers, follow_redirects=True) print(url, "->", check(response, selector, expected))https://quotes.toscrape.com/ -> ok: 10 itemshttps://quotes.toscrape.com/js/ -> incomplete: 0 of 10 expected itemshttps://www.amazon.com/ -> incomplete: 0 of 1 expected itemsAll 3 pages returned a status of 200. Only the first one held the data.
The second page isn’t a block. Its quotes load with JavaScript, so the plain HTML has none. The check can’t tell you why the data is missing, but it stops you from saving an empty page as a result.
The Amazon page passed the block-marker test and failed the item count. That’s the reason for the highlighted lines. Counting the items you expect catches blocks that no list of markers knows about.
What to check in production
- Count what you came for. Check for the number of items, a price, or a field you know the page always has.
- Watch the size. A product page that’s suddenly 4 KB instead of 400 KB is rarely a real page.
- Keep the blocked pages. Save a sample of failed responses. They show which system blocked you and how its page changes over time.
- Compare values over time. Honeypot data can pass every structural check. Prices that change sharply, or values that differ from a second source, are worth a closer look.
- Track success by data, not status. Report the share of requests that returned usable data. At ZyteCon 2026, Zyte’s CTO said that “one in 10 requests never delivers usable data at all” in their experience.
- Treat a block as information. It tells you how the site handles automated traffic. Slow down, check whether the site offers an API, and change your approach instead of retrying the same request.