AI in web scraping: what 2026 conferences agree on

Teams that scrape billions of pages use AI to build their scrapers, not to run them. That’s the clearest pattern in more than 60 talks, AMAs and podcasts I watched from 2025 and 2026. It covers OxyCon, Prague Crawl, ZyteCon, Extract Summit, AI Engineer and RAISE. I also read 10 research papers on web agents, bot detection and LLM extraction.

Stan Sadokov, founder of the proxy provider NodeMaven, said it most directly in a Reddit AMA. AI is “excellent at building the scraper and repairing it when a layout changes, and terrible as a component inside the request path”.

This post covers where AI pays off, where it still fails, and the new risks that come with it.

Why AI stays out of the request path

The reason is cost. At OxyCon 2026, Oxylabs showed the math for its own traffic. The company scrapes about 7 billion pages a day. Parsing only 10% of them with the cheapest LLM, at about 1 cent a page, would cost “7 million dollars a day”.

Zyte reached the same conclusion from a different direction. At ZyteCon 2026, its R&D team said per-page LLM extraction had limited use because of its cost, so they moved the LLM to writing extraction code instead. Zyte’s CTO said that “generic LLMs right now aren’t yet suited to production scraping at scale”.

Accuracy is the second reason. Scraped data often feeds pricing, finance and research products, where a small error rate is a failure. As Zyte’s R&D speaker put it at ZyteCon, “we don’t want like 85% solution here. In many many cases you need like 99% solution”.

A 2026 study on embeddings shows the same cost gap in a related task. Across 37 tasks, the best LLMs scored about the same as the best embedding models overall, at 1,431 times the cost. The authors’ advice fits scraping too. Use the cheap model by default, and save the LLM for the hard cases.

The case for LLM extraction

Not everyone agrees. At Extract Summit 2025, Jerome Choo of Diffbot argued that recent models are now accurate and cheap enough. His runs cost “roughly a tenth of a cent to 20th of a cent” per page, and his advice was to “write schemas, not rules”.

His inputs were cleaned article text, not raw HTML, and that detail matters. The NEXT-EVAL paper tested one LLM on 164 real pages with 3 input formats. The F1 score was 0.10 on slimmed HTML and 0.96 on a flat JSON format. The input format changed the result more than anything else tested.

Where AI pays off

AI pays off at build time and in operations. These are the uses that production teams described with numbers.

UseExampleResult reported
Writing scrapers from ticketsAn agent pipeline shown at Prague Crawl62 of 70 tickets correctly detailed on the first try
Generating a full dataset pipelineA demo at OxyCon 20261,000 products in about 36 minutes, from an empty project
Diagnosing broken scrapersOxylabs support agentMedian time to fix a degraded scraper fell from 5 hours to 2
Reading obfuscated anti-bot codeA speaker at Extract Summit DublinWork that took “three months” now takes about 15 minutes
Maintaining scrapersRafael Levi, Bright Data“I don’t ask the LLM to go do. I always ask, build a script”

All of these numbers come from the companies themselves, and none has been tested independently.

The agent that engineers stopped using

The most useful story came from the Oxylabs talk, because it was a failure. The team built an agent that changed parsers on its own. It made 251 changes in under a year and cost $847.

The results were weaker than the engineers’ own work.

  • 59% of its merge requests reached production, compared with 87% for engineers.
  • 76% passed CI tests, compared with 91%.
  • Engineers made 3.6 extra changes to each agent change.

The agent’s share of parser changes fell from 51% to 14%, while the models kept improving. As the speaker described it, “They just quietly stopped calling the agent”.

The version that worked supported engineers instead of replacing them. It had command-line access, written instructions and guardrails, such as not drawing conclusions from fewer than 30 jobs.

Where AI still fails

Fully autonomous scraping doesn’t exist yet. Ivan Sanchez of Zyte opened his Extract Summit talk with the line “an end to end fully autonomous scraping web scraping agent still does not exist”.

He gave 2 reasons. A CAPTCHA ends the agent’s run, and a single Amazon product page can exceed 1 million tokens of HTML.

Long web tasks are also slow and expensive. The Odysseys benchmark tested agents on 200 multi-site tasks on the live web. The best model completed 44.5% of them perfectly. Each 100-step run took about 30 minutes and cost about $2.50.

Coding agents hit a limit on the hardest targets. At Prague Crawl, the team behind the ticket-to-code pipeline said that on tricky sites or at high volume, Claude Code “just cannot do it”.

Another OxyCon speaker made the same point about general coding agents. They can scrape some sites, but for most of them “it’s not LLM helping you, but you helping it”.

Too much page content in the context makes agents worse. Hostinger’s team said at OxyCon that its agents received up to 500,000 characters from each web search. The team switched to showing page titles first and fetching full pages on request. Token use fell by 95%, and the agent’s answers improved.

Use the browser last

Most scraping teams now treat a full browser as the last option, not the default. The reasons are memory, speed and cost.

At OxyCon 2026, ScrapingBee compared running a page’s JavaScript in a lightweight DOM with loading it in a full browser. The lightweight method used 207 MB of memory against 470 MB for the browser, on one test page. The speaker’s point was cost, because “the memory becomes density… And density becomes cost”.

ScrapingBee ranks the options from cheapest to most expensive.

  1. Parse the HTML the server returns.
  2. Call the API the page uses.
  3. Run the page’s JavaScript without a browser.
  4. Load the page in a full browser.

The speaker’s advice was to “reach the browser last, not first”.

Other teams report the same thing at a larger scale. Batuhan Özyön, founder of Scrape.do, said “more than 99% of our requests work without a browser”. Geonode’s team said “Half the sites that look like they need a browser are just hitting a JSON endpoint with a signature”.

The other side of this argument comes from anti-bot design. As Sarah McKenna of Sequentum explained at Extract Summit, “The only way to effectively block a bot is to make it more expensive and slower. And the way you do that is by forcing the use of browsers”.

AI agents are different. They usually need a browser to act on a page. Paul Klein IV of Browserbase said at AI Engineer that the most reliable browser agents in production “are often writing code alongside using the browser”.

The new risks

AI brings new ways for scrapers to fail quietly. Three came up again and again.

Successful requests with the wrong data

A status code of 200 doesn’t prove you got the data. At Prague Crawl, Logan Harless said “vendors with serious mindshare are returning 200 when really what they’re serving you is a CAPTCHA”.

Zyte’s data team described honeypots at ZyteCon, where “you get a 200 but it’s the wrong data”. Scrape.do described partial pages, where “the content is quietly degraded”, for example 5 records instead of 10.

Ian Kerins, CEO of ScrapeOps, expects this to get worse. In his AMA, he said the next generation of anti-scraping “may make you collect data you cannot trust”. Checking the content of each response, not only its status, is becoming part of the job.

Stealth tools that make detection easier

A 2026 study from the University of Lille and Inria tested 12 scraping tools and web agents against 9 bot defenses. Its main finding was that “stealth and anti-detection mechanisms often increase detectability rather than decrease it”.

The researchers could identify every tool they tested, and they reached 99.3% accuracy by combining IP, TLS and browser fingerprints. Any one of those signals alone was far weaker. Saksham Solanki, creator of the httpcloak library, described the same idea in his AMA, saying “the handshake is a gate, the session is a score”.

Prompt injections aimed at scrapers

LLM-based scrapers read text that site owners control. A 2026 study scanned about half of a Common Crawl snapshot and found 15,300 prompt injections on 11,722 pages. Crawlers and scrapers were the main targets.

About 87% of the injections were invisible to human visitors. Many sat in HTTP headers and JSON-LD blocks, which are exactly where scrapers look for clean data. The authors’ conclusion was measured. In their words, the threat “is not yet a dominant threat, but it is already sufficiently real, structured, and widespread to deserve attention”.

Is the web closing, or getting more expensive?

Speakers disagreed on this question more than on any other. Most positions fall into 4 groups.

PositionWho said itKey quote
Repriced, not closedIan Kerins, ScrapeOps“It is being repriced. Access remains possible, but fewer datasets are profitable to collect at scale”
Splitting in twoStan Sadokov, NodeMaven“open machine paths for the commodity stuff and harder walls around the valuable stuff”
Closing, for some usersKrishna, a lawyer at Black Forest Labs, OxyCon 2026 panelMore data sits behind logins, paywalls and platforms, which creates “an asymmetry in data access”
Opening, for new usersPaul, journalism educator, OxyCon 2026 panel“for the vast majority of people, it was always closed. And actually, it’s opening up now”

The cost numbers support the “repriced” view. As Kerins put it, “Proxies themselves are cheaper than they used to be, while the cost of producing one genuinely successful scrape has increased”. At ZyteCon, Zyte shared findings from a cost study with its clients. People’s time was about 70% of the cost of running scrapers. Anti-ban tools were roughly 15%. And 20% to 30% of what teams counted as bans were malformed URLs.

How to decide where to use AI

Based on these talks, a practical order looks like this.

  1. Use AI to write and repair scrapers. This is where production teams report the clearest gains.
  2. Keep extraction deterministic where accuracy matters. If you use an LLM, give it clean, structured input, and test the output against known data.
  3. Try the cheapest access method first. Check the HTML, then the page’s own API, and use a browser only when the data needs one.
  4. Check the content, not only the status code. A 200 response can hold a CAPTCHA, a partial page or planted data.
  5. Treat page content as untrusted input. Hidden instructions can reach any LLM that reads a page.

Most of the evidence here comes from vendors describing their own systems, so treat each number as a report, not a benchmark. Even so, companies that compete with each other every day described the same pattern. AI is now part of how scrapers get built. The scrapers themselves still run as regular code.

Sources

Talks I watched online, in the order they appear.

Papers I cited are NEXT-EVAL, The Embedder’s Dilemma, Odysseys, Unmasking Web Agents with Multi-Layer Fingerprinting and Indirect Prompt Injection in the Wild.