Does copying Chrome's user agent help a scraper? A 472-site test
A common first fix for a blocked scraper is to copy Chrome’s user agent into the request. I tested that against 4 other user agents, first on Wikipedia and then on 472 top homepages.
Replacing the library’s default user agent helped. Copying Chrome’s didn’t beat an honest one. It mostly changed which sites worked, and on some sites it changed which version of the page came back.
The Wikipedia test
I started with Wikipedia, because my 1,000-site test got a 403 there with a copied Chrome user agent. This script requests the same page 5 times, with a different user agent each time.
import httpx
url = "https://en.wikipedia.org/wiki/Web_scraping"chrome = ( "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 " "(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36")user_agents = { "library default": None, "copied from Chrome": chrome, "made-up name": "MyScraperBot/1.0", "Chrome + email": f"{chrome} (you@example.com)", "descriptive": "MyScraperBot/1.0 (https://example.com/bot; you@example.com) httpx/0.28",}
for name, user_agent in user_agents.items(): headers = {"User-Agent": user_agent} if user_agent else {} response = httpx.get(url, headers=headers, follow_redirects=True) print(f"{name:20} {response.status_code} {response.text[:60]!r}")library default 403 'Please set a user-agent and respect our robot policy https:/'copied from Chrome 403 'Please respect our robot policy https://w.wiki/4wJS when cra'made-up name 403 'Please respect our robot policy https://w.wiki/4wJS when cra'Chrome + email 200 '<!DOCTYPE html>\n<html class="client-nojs vector-feature-lang'descriptive 200 '<!DOCTYPE html>\n<html class="client-nojs vector-feature-lang'Every user agent without contact details was blocked, including the copied Chrome one. The 2 with an email address got the page. Wikipedia’s User-Agent policy explains why. It says “Do not copy a browser’s user agent for your bot, as bot-like behavior with a browser’s user agent will be assumed malicious”.
The same test on 472 sites
One site doesn’t show a pattern, so I ran the same 5 user agents against the homepages from the Tranco top 1,000 that returned a normal page in an earlier check. That’s 472 sites, with one request per user agent, 1 second apart. I ran it once, from Mumbai, on October 3, 2026, with httpx 0.28.
A user agent counted as successful on a site when it got a status of 200 and at least half the text of the best page that site sent to any of the 5. 402 sites sent a full page to at least one of them.
| User agent | Full page | Share |
|---|---|---|
Library default, python-httpx/0.28 | 353 | 88% |
| Copied from Chrome | 375 | 93% |
Made-up name, MyScraper/1.0 | 377 | 94% |
| Descriptive, with “bot” and contact details | 378 | 94% |
| Chrome with an email added | 380 | 95% |
The library default did worst. The other 4 each got 22 to 27 more sites. Copying Chrome’s user agent came out level with an honest, descriptive one.
Chrome’s user agent changed which sites worked
The totals hide a split. 23 sites gave the full page to the descriptive user agent but not to Chrome’s, and 20 did the opposite.
- Better with the descriptive one. Amazon’s stores, Wikipedia and Wikimedia, Facebook and WhatsApp, Uber, W3C, SourceForge, MySQL and LG. Most of them answered Chrome’s user agent with an error status or a short challenge page.
- Better with Chrome’s. Google’s sites, Yahoo Japan and Mail.ru, plus Lowe’s and DigiCert. Lowe’s and DigiCert sent other clients a short page instead of the full homepage.
Most of the Google cases weren’t blocks. Google’s homepage sent both Chrome user agents a fuller version, with 71 words of text, and the other 3 a simpler one with 33.
The user agent changes the page you get
Some sites don’t just allow or block a user agent. They send a different version of the page. Le Monde sent the descriptive user agent a homepage with 11,736 words of text, and Chrome’s user agent one with 1,547.
Amazon’s homepage reacted to the word “bot”. The first 3 rows come from one run of requests, a few seconds apart, and the Chrome row from the 472-site test.
| User agent | Bytes | Words |
|---|---|---|
MyScraperBot/1.0 | 1,074,330 | 558 |
MyScraper/1.0 | 3,781 | 22 |
MyScraper/1.0 (https://example.com/bot; you@example.com) httpx/0.28 | 1,074,330 | 558 |
| Copied from Chrome | 2,185 | 0 |
With “bot” anywhere in the user agent, Amazon sent the full homepage. Without it, the client got a short page that asks the visitor to click a button, and Chrome’s user agent got a challenge. That explains the short Amazon page in my post on 200 responses that are really blocks. Its user agent didn’t include “bot”.
What a good user agent looks like
Wikipedia’s policy gives a format, and in the 472-site test it did as well as Chrome’s user agent.
MyScraperBot/1.0 (https://example.com/bot; you@example.com) httpx/0.28It has 3 parts.
- A name and version for your scraper, with “bot” in it, not the library’s default.
- Contact details, like a web page about the bot or an email address, so a site owner can reach you instead of blocking you.
- The library you use, which helps with debugging.
A user agent is only one of many signals a request carries. Saksham Solanki, creator of the httpcloak library, described the general rule in a Reddit AMA as “the handshake is a gate, the session is a score”.
Where bot identity is going
A user agent is a claim that anyone can copy. The IETF has a working group, Web Bot Auth, on a standard for bots to sign their requests with a key that sites can check. Its architecture draft was last revised in March 2026. Once sites can verify who a bot is, an honest identity becomes worth more than a copied one.
When identifying yourself isn’t enough
Many sites block automated traffic whatever the user agent says. In my 1,000-site test, 14% of top homepages blocked a plain request but allowed a headless browser. A better user agent won’t change that.
On those sites, production teams change the rest of the setup, not the user agent. They use a real browser, better IPs or an unblocking API, which is the service most web data companies sell. For Wikipedia, none of that is needed. Its official APIs and database dumps are faster and more complete than scraping pages.
A short checklist
- Replace the library’s default user agent. It was the worst choice in this test.
- Name your scraper, include “bot” and add contact details. It did as well as Chrome’s user agent, and it’s what sites like Wikipedia ask for.
- Don’t copy a browser’s user agent into a script. It didn’t raise the success rate, and Wikipedia’s policy treats it as a sign of a malicious bot.
- Check which version of the page you got. Compare the text you received with what a browser shows, because some sites send different pages to different clients.
- Keep your request rate low, and wait longer between retries when you get a 429 or 503.