How to find the JSON data hidden in a page's source

Many pages that look like they need a browser already contain their data as JSON inside the HTML. JavaScript reads that JSON and builds the page. A scraper can read the same JSON directly, which is faster and more stable than parsing the page’s HTML.

In my test of the top 1,000 websites, about 55% of homepages with readable content included JSON-LD, and 62 were built with Next.js, which sends page data as JSON. This post shows how to find that data and read it with Python. I ran every script from Mumbai on October 3, 2026, with httpx 0.28.

How to spot embedded data

Start with a value you can see on the page, like a product name, a price or a headline.

  1. Open the page in your browser, then right-click and choose View page source. This shows the HTML the server sent, before any JavaScript runs.
  2. Search the source for the value you picked.
  3. If the value appears inside a <script> tag, the page sends its data as JSON or JavaScript.

If the value isn’t in the source at all, the page loads it later from an API. Open DevTools, go to the Network tab, filter by Fetch/XHR, and reload to find that request.

The patterns to look for

Most embedded data follows one of a few patterns. The counts come from the same 1,000-site test, out of 390 homepages with readable content.

PatternWhat it looks likeHomepages
JSON-LD<script type="application/ld+json">215
Next.js (App Router)self.__next_f.push(...)33
Next.js (Pages Router)<script id="__NEXT_DATA__">29
Nuxt<script id="__NUXT_DATA__"> (Nuxt 3) or window.__NUXT__ (Nuxt 2)13
Redux and similarwindow.__PRELOADED_STATE__ or window.__INITIAL_STATE__9

JSON-LD is the easiest to use, because its format is standard. The Next.js Pages Router format is the next easiest, because the whole page’s data sits in one tag. The App Router format is harder, since the data is split into many small pieces in a React-specific format.

Nuxt changed its format in version 3. On a Nuxt 3 site, window.__NUXT__ holds only settings, and the page data sits in the __NUXT_DATA__ tag. That data is a flat list where values point to other positions in the list, so you rebuild objects by following those positions.

Reading __NEXT_DATA__

TED’s homepage is built with Next.js, and its recommended talks are in the __NEXT_DATA__ tag. This script reads the newest talks without running any JavaScript.

ted_next_data.py
import json
import httpx
from bs4 import BeautifulSoup
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}
html = httpx.get("https://www.ted.com/", headers=headers, follow_redirects=True).text
tag = BeautifulSoup(html, "html.parser").select_one("script#__NEXT_DATA__")
data = json.loads(tag.string)
talks = data["props"]["pageProps"]["sailthruDataContent"]["RecommendationsNewest"]
for talk in talks[:3]:
minutes = talk["duration"] // 60
print(f'{talk["title"]} | {talk["presenterDisplayName"]} | {minutes} min')
Output
AI and the end of loneliness | Paul Bloom | 12 min
Why we should design cities like Disney | Zach DeBoer | 16 min
Why are AI data centers using so much electricity? | Sajan Saini | 7 min

Your output will list different talks, because TED updates its homepage. To find a path like props.pageProps.sailthruDataContent, save the JSON to a file and open it in an editor, or search it for a value you saw on the page.

A shortcut, the _next/data route

Pages Router sites often serve the same data as plain JSON, without the HTML around it. The address uses the site’s build ID, which __NEXT_DATA__ holds.

ted_next_route.py
import json
import httpx
from bs4 import BeautifulSoup
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}
html = httpx.get("https://www.ted.com/", headers=headers, follow_redirects=True).text
tag = BeautifulSoup(html, "html.parser").select_one("script#__NEXT_DATA__")
build_id = json.loads(tag.string)["buildId"]
for page in ["index", "talks"]:
url = f"https://www.ted.com/_next/data/{build_id}/{page}.json"
response = httpx.get(url, headers=headers)
props = list(response.json()["pageProps"])
print(f"{page}.json | {response.status_code} | {len(response.content):,} bytes | {props}")
Output
index.json | 200 | 288,751 bytes | ['programmerRibbons', 'prismicPage', 'sailthruDataContent']
talks.json | 200 | 44,079 bytes | ['talks', 'absoluteUrl']

The homepage data came back at 288,751 bytes, against 589,257 for the full HTML, with nothing to parse. The same route served the data for TED’s talks page.

It doesn’t work everywhere. I tried it on the 26 Pages Router homepages from my 1,000-site test that still had __NEXT_DATA__, and 9 returned JSON. The build ID also changes with every deploy, so read it again when a request returns 404.

Reading App Router data

Newer Next.js sites use the App Router. They send the page’s data in many self.__next_f.push(...) script tags, and joining them gives you the full payload. DeepL’s pricing page works this way.

deepl_rsc.py
import json
import re
import httpx
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}
html = httpx.get("https://www.deepl.com/en/pro", headers=headers, follow_redirects=True).text
# App Router pages send their data in many small script tags
chunks = re.findall(r'self\.__next_f\.push\(\[1,"(.*?)"\]\)</script>', html, re.DOTALL)
payload = "".join(json.loads(f'"{chunk}"') for chunk in chunks)
prices = re.findall(r'"packageId":"[^"]+","price":\{"basePrice":\{"monthly":(?:\d+|null)', payload)
currencies = sorted(set(re.findall(r'"currency":"(\w+)"', payload)))
print(len(chunks), "chunks,", f"{len(payload):,}", "characters")
print(len(prices), "price entries in", len(currencies), "currencies:", currencies)
Output
33 chunks, 239,913 characters
23 price entries in 4 currencies: ['cad', 'eur', 'jpy', 'usd']

The page itself showed 4 prices, all in euros. Its source held 23 price entries in 4 currencies. It happens because frameworks send the data the page might need, not only what it shows.

The payload isn’t plain JSON. It’s a React format, made of lines that each start with an ID, so a regular expression on the known field names is often the practical way in. A Python library for this format, njsparser, failed on both App Router pages I tried, which shows how often the format changes.

Reading a JavaScript variable

Some pages store their data in a JavaScript variable instead. Genius, the lyrics site, keeps its homepage data in window.__PRELOADED_STATE__. It shows 2 things real pages do that tidy examples don’t.

  • The JSON is inside a JavaScript string. The page runs JSON.parse('...') on it, so the text is escaped twice. You undo the string escaping first, then parse the JSON.
  • The data is normalized. The chart holds only song IDs. The song details sit in a separate entities table, so you look each one up by ID. Redux-style sites often store data this way.
genius_state.py
import json
import re
import httpx
headers = {"User-Agent": "MyScraper/1.0 (you@example.com) python-httpx"}
html = httpx.get("https://genius.com/", headers=headers, follow_redirects=True).text
# The state is JSON inside a JavaScript string: JSON.parse('...')
match = re.search(r"window\.__PRELOADED_STATE__ = JSON\.parse\('(.*?)'\);", html, re.DOTALL)
text = json.loads('"' + match.group(1).replace("\\'", "'") + '"') # undo the string escaping
state = json.loads(text)
chart = state["home"]["chartSection"]["chartItems"]
songs = state["entities"]["songs"]
for entry in chart[:3]:
song = songs[str(entry["item"]["id"])]
print(f'{song["title"]} | {song["artistNames"]} | {song["stats"].get("pageviews")} views')
Output
Patient Zero | Taylor Swift | 1989869 views
Cleveland! | Taylor Swift | 1641714 views
Nicole Kidman | ADÉLA | 596234 views

Your output will list different songs, because the chart changes every day. The second json.loads() call also keeps characters like the “É” in ADÉLA intact, which a quick unicode_escape decode would break.

Other pages assign a plain JavaScript object, like window.__INITIAL_STATE__ = {...}. That often isn’t valid JSON, because JavaScript allows single quotes, comments and unquoted keys. For those, a JavaScript parser like chompjs for Python handles the conversion.

Things to watch

The JSON can hold more than the page shows. As the DeepL example shows, frameworks often send data the page never displays, including internal IDs, and sometimes personal details. Take only the public fields you need, and leave personal data out.

The structure changes without notice. Paths like sailthruDataContent are internal names. A site can rename them in any release, so check that each field exists and fail clearly when it doesn’t.

Not every value comes from the first load. Prices, stock and search results often load later from an API. If a value isn’t in the source, look for that request in the Network tab.

This is the second option, not the first. If the data is already in the HTML, parse the HTML. If it’s in embedded JSON, read the JSON. Use a browser only when the data needs one.