Skip unchanged pages on a re-crawl: how many top sites return 304

A re-crawl usually downloads every page again, including pages that haven’t changed. HTTP has a standard way to avoid that. Your scraper sends what it knows about its last copy, and the site answers with a 304 status, “not modified”, and no page body. Googlebot uses it when it re-crawls pages for Google Search.

I tested how often this works on real sites. I used the homepages from my test of the top 1,000 websites that returned content, plus one inner page from each site’s sitemap. I ran it from Mumbai on October 8, 2026, with httpx 0.28. I ran the full test 3 times, and the shares below changed by 1 point at most.

How a conditional request works

A site can send 2 headers with a page, called validators.

  • ETag is an ID for this version of the page, like W/"6eced526...".
  • Last-Modified is the date the page last changed.

On the next crawl, you send them back. If-None-Match contains the ETag, and If-Modified-Since contains the date. If the page hasn’t changed, the site returns a 304 with headers only. If it has, you get a normal 200 with the new page.

recrawl.py
import httpx
headers = {"User-Agent": "MyScraperBot/1.0 (https://example.com/bot; you@example.com) httpx/0.28"}
url = "https://www.datadoghq.com/product/experiments/"
with httpx.Client(headers=headers, follow_redirects=True, timeout=30) as client:
first = client.get(url)
etag, last_modified = first.headers.get("etag"), first.headers.get("last-modified")
print("First crawl |", first.status_code, "|", first.num_bytes_downloaded, "bytes | ETag", etag)
# On the next crawl, send the validators from the last response.
conditional = {}
if etag:
conditional["If-None-Match"] = etag
if last_modified:
conditional["If-Modified-Since"] = last_modified
again = client.get(url, headers=conditional)
print("Re-crawl |", again.status_code, "|", again.num_bytes_downloaded, "bytes")
Output
First crawl | 200 | 36470 bytes | ETag W/"6eced526946936a77f45c9467c323d0f"
Re-crawl | 304 | 0 bytes

In a real crawler, you store the validators with each page and send them on the next run. When the answer is a 304, you keep your stored copy and skip parsing, chunking and embedding.

How many top sites support it

For each page, I sent 2 normal requests a second apart, then a request with each validator.

HomepagesInner pages
Pages that returned a 200368224
Sent an ETag125 (34%)74 (33%)
Sent a Last-Modified date123 (33%)71 (32%)
Sent neither181 (49%)110 (49%)
Returned a 304 to a conditional request151 (41%)90 (40%)

About 40% of pages, homepages or not, can skip the download on a re-crawl. On those pages, the full page body had a median size of about 45 KB, compressed. The 304 had no body at all.

That matters most when you pay for bandwidth. Residential proxies are usually billed per GB, so a 304 costs almost nothing compared with a full page.

What production crawlers do

Google’s crawlers use the same mechanism. Google’s crawler documentation (opens in a new tab) says they support both validators, and when a page sends both, they use the ETag.

Google has also published how rarely it works across the web. In a December 2024 post (opens in a new tab), Google asked sites to support caching, because “the number of requests that can be returned from local caches has decreased”. About 0.026% of its fetches were cacheable 10 years earlier, and 0.017% at the time of the post.

That’s far below the 40% in my test, and the 2 numbers measure different things. Google counts all of its fetches, across the whole web and over long gaps between crawls, when pages are more likely to have changed. My test sent the conditional request seconds after the first one. Together, they show the limit of the method. It saves a lot on the sites that support it, and it saves nothing on the many that don’t.

When the validators don’t help

Some validators don’t do their job. These are the cases I found.

  • The ETag changes on every request. On 17 homepages and 6 inner pages, the ETag was different on 2 requests a second apart. A conditional request can never match, so every re-crawl downloads the page.
  • The site ignores the ETag. 38 homepages and 21 inner pages sent an ETag but returned a full 200 to If-None-Match. 15 of those homepages and 5 of those inner pages returned a 304 to If-Modified-Since instead, so send both.
  • Last-Modified is the time of your request. On 13 homepages and 5 inner pages, the date was within a minute of the response time, so it says nothing about when the page changed.
  • A 304 when the text did change. On 3 homepages and 2 inner pages, the visible text was different on 2 normal requests, but the site still returned a 304. That’s rare, but a crawler that trusts every 304 will keep an old version of these pages.

To catch an ETag that changes on every request, request a page twice when you first add it. If the ETag changes, don’t use it for that site.

In Scrapy

Scrapy has this built in. Its HTTP cache stores each response on disk, and the RFC2616Policy sends the validators when it checks a stored page again.

settings.py
HTTPCACHE_ENABLED = True
HTTPCACHE_POLICY = "scrapy.extensions.httpcache.RFC2616Policy"

I ran a spider on the same Datadog page twice, with Scrapy 2.19. On the second run, the stats showed httpcache/revalidate, which Scrapy counts when the site returns a 304. It then used its stored copy.

When a site sends no validators

About half the pages sent neither header. For those, you have to download the page again. You can still skip the expensive work after the download. Hash the page’s main text, compare it with the hash from the last crawl, and process the page only when the hash changed.

This saves processing, not bandwidth. Hash the extracted text, not the raw HTML. On 28 homepages, the visible text changed between 2 requests a second apart. The raw HTML changed on 150, in parts a reader doesn’t see, like scripts and attributes.

Limits of this test

  • One request pattern. I sent the conditional requests seconds after the first ones. Over days, more pages will change and return a 200, which is correct.
  • One inner page per site, the first page in its sitemap. It may be older or less visited than a typical page.
  • One location and one client. Some sites answer differently by country, or to browsers, proxies and CDNs.

Data and code

Doccraft writes research and tutorials like this for web data and AI infrastructure companies.