How to find the JSON data hidden in a page's source
Many pages that look like they need a browser already contain their data as JSON inside the HTML. JavaScript reads that JSON and builds the page. A scraper can read the same JSON directly, which is faster and more stable than parsing the page’s HTML.
In my test of the top 1,000 websites, about 55% of homepages with readable content included JSON-LD, and 62 were built with Next.js, which sends page data as JSON. This post shows how to find that data and read it with Python. I ran every script from Mumbai on October 3, 2026, with httpx 0.28.
How to spot embedded data
Start with a value you can see on the page, like a product name, a price or a headline.
- Open the page in your browser, then right-click and choose View page source. This shows the HTML the server sent, before any JavaScript runs.
- Search the source for the value you picked.
- If the value appears inside a
<script>tag, the page sends its data as JSON or JavaScript.
If the value isn’t in the source at all, the page loads it later from an API. Open DevTools, go to the Network tab, filter by Fetch/XHR, and reload to find that request.
The patterns to look for
Most embedded data follows one of a few patterns. The counts come from the same 1,000-site test, out of 390 homepages with readable content.
| Pattern | What it looks like | Homepages |
|---|---|---|
| JSON-LD | <script type="application/ld+json"> | 215 |
| Next.js (App Router) | self.__next_f.push(...) | 33 |
| Next.js (Pages Router) | <script id="__NEXT_DATA__"> | 29 |
| Nuxt | <script id="__NUXT_DATA__"> (Nuxt 3) or window.__NUXT__ (Nuxt 2) | 13 |
| Redux and similar | window.__PRELOADED_STATE__ or window.__INITIAL_STATE__ | 9 |
JSON-LD is the easiest to use, because its format is standard. The Next.js Pages Router format is the next easiest, because the whole page’s data sits in one tag. The App Router format is harder, since the data is split into many small pieces in a React-specific format.
Nuxt changed its format in version 3. On a Nuxt 3 site, window.__NUXT__ holds only settings, and the page data sits in the __NUXT_DATA__ tag. That data is a flat list where values point to other positions in the list, so you rebuild objects by following those positions.
Reading __NEXT_DATA__
TED’s homepage is built with Next.js, and its recommended talks are in the __NEXT_DATA__ tag. This script reads the newest talks without running any JavaScript.
import json
import httpxfrom bs4 import BeautifulSoup
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}html = httpx.get("https://www.ted.com/", headers=headers, follow_redirects=True).text
tag = BeautifulSoup(html, "html.parser").select_one("script#__NEXT_DATA__")data = json.loads(tag.string)
talks = data["props"]["pageProps"]["sailthruDataContent"]["RecommendationsNewest"]for talk in talks[:3]: minutes = talk["duration"] // 60 print(f'{talk["title"]} | {talk["presenterDisplayName"]} | {minutes} min')AI and the end of loneliness | Paul Bloom | 12 minWhy we should design cities like Disney | Zach DeBoer | 16 minWhy are AI data centers using so much electricity? | Sajan Saini | 7 minYour output will list different talks, because TED updates its homepage. To find a path like props.pageProps.sailthruDataContent, save the JSON to a file and open it in an editor, or search it for a value you saw on the page.
A shortcut, the _next/data route
Pages Router sites often serve the same data as plain JSON, without the HTML around it. The address uses the site’s build ID, which __NEXT_DATA__ holds.
import json
import httpxfrom bs4 import BeautifulSoup
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}html = httpx.get("https://www.ted.com/", headers=headers, follow_redirects=True).texttag = BeautifulSoup(html, "html.parser").select_one("script#__NEXT_DATA__")build_id = json.loads(tag.string)["buildId"]
for page in ["index", "talks"]: url = f"https://www.ted.com/_next/data/{build_id}/{page}.json" response = httpx.get(url, headers=headers) props = list(response.json()["pageProps"]) print(f"{page}.json | {response.status_code} | {len(response.content):,} bytes | {props}")index.json | 200 | 288,751 bytes | ['programmerRibbons', 'prismicPage', 'sailthruDataContent']talks.json | 200 | 44,079 bytes | ['talks', 'absoluteUrl']The homepage data came back at 288,751 bytes, against 589,257 for the full HTML, with nothing to parse. The same route served the data for TED’s talks page.
It doesn’t work everywhere. I tried it on the 26 Pages Router homepages from my 1,000-site test that still had __NEXT_DATA__, and 9 returned JSON. The build ID also changes with every deploy, so read it again when a request returns 404.
Reading App Router data
Newer Next.js sites use the App Router. They send the page’s data in many self.__next_f.push(...) script tags, and joining them gives you the full payload. DeepL’s pricing page works this way.
import jsonimport re
import httpx
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}html = httpx.get("https://www.deepl.com/en/pro", headers=headers, follow_redirects=True).text
# App Router pages send their data in many small script tagschunks = re.findall(r'self\.__next_f\.push\(\[1,"(.*?)"\]\)</script>', html, re.DOTALL)payload = "".join(json.loads(f'"{chunk}"') for chunk in chunks)
prices = re.findall(r'"packageId":"[^"]+","price":\{"basePrice":\{"monthly":(?:\d+|null)', payload)currencies = sorted(set(re.findall(r'"currency":"(\w+)"', payload)))print(len(chunks), "chunks,", f"{len(payload):,}", "characters")print(len(prices), "price entries in", len(currencies), "currencies:", currencies)33 chunks, 239,913 characters23 price entries in 4 currencies: ['cad', 'eur', 'jpy', 'usd']The page itself showed 4 prices, all in euros. Its source held 23 price entries in 4 currencies. It happens because frameworks send the data the page might need, not only what it shows.
The payload isn’t plain JSON. It’s a React format, made of lines that each start with an ID, so a regular expression on the known field names is often the practical way in. A Python library for this format, njsparser, failed on both App Router pages I tried, which shows how often the format changes.
Reading a JavaScript variable
Some pages store their data in a JavaScript variable instead. Genius, the lyrics site, keeps its homepage data in window.__PRELOADED_STATE__. It shows 2 things real pages do that tidy examples don’t.
- The JSON is inside a JavaScript string. The page runs
JSON.parse('...')on it, so the text is escaped twice. You undo the string escaping first, then parse the JSON. - The data is normalized. The chart holds only song IDs. The song details sit in a separate
entitiestable, so you look each one up by ID. Redux-style sites often store data this way.
import jsonimport re
import httpx
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}html = httpx.get("https://genius.com/", headers=headers, follow_redirects=True).text
# The state is JSON inside a JavaScript string: JSON.parse('...')match = re.search(r"window\.__PRELOADED_STATE__ = JSON\.parse\('(.*?)'\);", html, re.DOTALL)text = json.loads('"' + match.group(1).replace("\\'", "'") + '"') # undo the string escapingstate = json.loads(text)
chart = state["home"]["chartSection"]["chartItems"]songs = state["entities"]["songs"]for entry in chart[:3]: song = songs[str(entry["item"]["id"])] print(f'{song["title"]} | {song["artistNames"]} | {song["stats"].get("pageviews")} views')Patient Zero | Taylor Swift | 1989869 viewsCleveland! | Taylor Swift | 1641714 viewsNicole Kidman | ADÉLA | 596234 viewsYour output will list different songs, because the chart changes every day. The second json.loads() call also keeps characters like the “É” in ADÉLA intact, which a quick unicode_escape decode would break.
Other pages assign a plain JavaScript object, like window.__INITIAL_STATE__ = {...}. That often isn’t valid JSON, because JavaScript allows single quotes, comments and unquoted keys. For those, a JavaScript parser like chompjs for Python handles the conversion.
Things to watch
The JSON can hold more than the page shows. As the DeepL example shows, frameworks often send data the page never displays, including internal IDs, and sometimes personal details. Take only the public fields you need, and leave personal data out.
The structure changes without notice. Paths like sailthruDataContent are internal names. A site can rename them in any release, so check that each field exists and fail clearly when it doesn’t.
Not every value comes from the first load. Prices, stock and search results often load later from an API. If a value isn’t in the source, look for that request in the Network tab.
This is the second option, not the first. If the data is already in the HTML, parse the HTML. If it’s in embedded JSON, read the JSON. Use a browser only when the data needs one.