Scraping JavaScript pages without a full browser, retested
At OxyCon 2026, ScrapingBee gave a talk on “web hydration” (opens in a new tab), which means running a page’s JavaScript in a lightweight DOM instead of a full browser. The claim was that this uses much less memory for the same result. I reran the test to check.
The memory saving was confirmed on the talk’s test page. I also tried a third engine, Lightpanda, and ran all the methods on 29 real sites that need JavaScript. On real sites, the lightweight methods worked far less often.
What the talk proposed
Many servers return pages as an empty shell. The data appears only after the page’s JavaScript runs and builds the page. The talk calls this step hydration.
The speaker’s point was that hydration needs only 2 things, a JavaScript engine and a DOM. It doesn’t need the parts of a browser that draw pixels on a screen. So a library like jsdom, running on Node.js, can hydrate a page without launching Chromium.
The talk compared the 2 methods on the JavaScript version of Quotes to Scrape, a practice site for scraping. The hydration method took 2.2 seconds and 207 MB of memory at its peak. The browser took 2.7 seconds and 470 MB.
The speaker was careful about the speed. “If you came here expecting 10 times faster, this is not the real finding here”. The real gain was memory, which limits how many scrapers can run on one machine. In the speaker’s words, “the memory becomes density… And density becomes cost”.
The retest
I ran jsdom, headless Chromium and Lightpanda on the same page, 5 times each, on an Apple M3 Mac. The memory figure is the peak for all the processes each method started, including Node.js.
This is the hydration script, with jsdom 29.
// Hydrate the JavaScript version of Quotes to Scrape without a browser.import { JSDOM } from 'jsdom';
const url = 'https://quotes.toscrape.com/js/';const html = await (await fetch(url)).text();
const dom = new JSDOM(html, { url, // so relative script URLs resolve runScripts: 'dangerously', // run the page's own JavaScript resources: 'usable', // and load its external scripts (jQuery here)});await new Promise((resolve) => dom.window.addEventListener('load', resolve));
const quotes = [...dom.window.document.querySelectorAll('.quote .text')];console.log(quotes.length, 'quotes');console.log(quotes[0].textContent.slice(0, 60));dom.window.close();The browser script loads the same page in headless Chromium through Playwright, and reads the same elements.
Lightpanda is an open-source headless browser built for automation, not a DOM library. Version 1.0.0 was released on October 2, 2026. I ran lightpanda fetch --dump html on the same page 5 times, on October 3.
| Method | Median time | Median peak memory |
|---|---|---|
| jsdom | 2.92 s | 252 MB |
| Headless Chromium | 2.88 s | 468 MB |
| Lightpanda | 2.66 s | 35 MB |
The browser used about 1.9 times as much memory as jsdom, close to the talk’s 2.3 times. The times were close, as the talk said. On one small page, every method is mostly waiting for the network.
Lightpanda found all 10 quotes, with a median peak of 35 MB, about a seventh of jsdom’s.
A difference in the output
jsdom and the browser both found all 10 quotes. But they returned them in a different order.
jsdom : Steve Martin, Eleanor Roosevelt, Thomas A. Edison, André Gide, Albert Einstein, Marilyn Monroe, Jane Austen, Albert Einstein, J.K. Rowling, Albert Einsteinbrowser: Albert Einstein, J.K. Rowling, Albert Einstein, Jane Austen, Marilyn Monroe, Albert Einstein, André Gide, Thomas A. Edison, Eleanor Roosevelt, Steve MartinThe jsdom order is exactly reversed. This page builds its quotes with document.write(), and jsdom handles that call differently from a real browser.
The data is complete, so for many uses this doesn’t matter. It would matter for rankings, search results or anything where position is part of the data. jsdom is a close imitation of a browser, not a copy, so check its output against a real browser before you rely on it. Lightpanda returned the quotes in the browser’s order.
On real JavaScript sites
The test page is small and simple. Real sites that need JavaScript are larger apps, so I tried each method on the 29 homepages that needed JavaScript in my 1,000-site test. I ran each one once, from Mumbai, on October 3, 2026, with jsdom 29.1, Lightpanda 1.0.0 and headless Chromium through Playwright. I also added happy-dom 20.14, another lightweight DOM library.
Headless Chromium got real content from 26 of the 29. On those 26, I counted a method as working when it got at least half of Chromium’s text.
| Method | Worked on |
|---|---|
| jsdom | 3 of 26 |
| happy-dom | 1 of 26 |
| Lightpanda | 12 of 26 |
jsdom built most of these pages only partly, or not at all. happy-dom did no better. Lightpanda did much better, with 12 sites, including Spotify, Roblox, Steam Community, Intel and Font Awesome. Its median peak was 96 MB across all 29 sites.
So the memory saving is real, but these methods built only some of the real pages. The DOM libraries ran the test page’s simple script, but built few of the real apps. Lightpanda built 12 of the 26, and jsdom or happy-dom built 2 more. On the other 12, only the full browser worked.
Where hydration fits
The talk ended with a list of methods ordered by cost. Here it is, with my results added. Each step costs more than the one before.
- Parse the HTML, if the data is already there. In my 1,000-site test, this worked for 62% of top homepages.
- Call the page’s API directly, if the page loads its data from one.
- Hydrate the page with a lightweight DOM or a lightweight browser like Lightpanda, if the data only appears after JavaScript runs and there’s no clean API. Test it on your target pages first, because even Lightpanda worked on only 12 of the 26 real sites I tried.
- Use a full browser, if the page needs features the DOM library doesn’t have, like canvas and WebGL, or the site uses anti-bot scripts that inspect the browser.
The talk’s last line was this advice. “Please reach the browser last, not first”.
What the talk said about blocking
The talk also named the main limit. Without a browser, your requests no longer look like a browser’s at the network level. A plain HTTP client’s TLS handshake doesn’t look like Chrome’s, and anti-bot systems can block the request before any JavaScript runs.
A page also makes several requests as it loads, and they need to come from the same IP and fingerprint. The speaker suggested ScrapingBee’s API as one way to handle this, and said openly that “I’m biased because I work there”.
Hydration saves memory on pages you can access. It doesn’t help you access pages that block you.
Data and code
- Results for the 29 sites
- The Chromium, jsdom and happy-dom test, with its jsdom and happy-dom scripts
- The Lightpanda test