JSON-LD for scrapers: what it contains and what it misses
JSON-LD is structured data that websites add to their pages for search engines. In my test of the top 1,000 websites, about 55% of the homepages that loaded with real content included it. For a scraper, it’s often the cleanest source of names, dates, prices and ratings on a page.
This post explains what JSON-LD contains, how to read it with Python, and how well it worked on 8 real product pages. I ran every script from Mumbai on October 3, 2026, with httpx 0.28.
What JSON-LD is
JSON-LD is stored in a <script type="application/ld+json"> tag in the page’s HTML. It describes the page with the shared vocabulary from schema.org (opens in a new tab), which Google, Bing and other search engines read.
Because the format is standard, the same code can read it on most sites, and the field names are shared. A product has name and offers. An article has headline, datePublished and author.
Common types include these.
| Type | Typical fields |
|---|---|
Product | name, sku, brand, offers (price, currency, stock) |
Article, NewsArticle, BlogPosting | headline, author, datePublished, dateModified |
Organization | name, logo, sameAs (social profiles) |
Event | name, startDate, location |
Recipe | recipeIngredient, cookTime, nutrition |
BreadcrumbList | the page’s position in the site |
Reading it with Python
This function collects every JSON-LD item on a page. Some sites put several items in an @graph list, so the function flattens that list too.
import json
import httpxfrom bs4 import BeautifulSoup
headers = {"User-Agent": "MyScraperBot/1.0 (https://example.com/bot; you@example.com) httpx/0.28"}
def json_ld(url): """Return every JSON-LD item on a page, with @graph lists flattened.""" html = httpx.get(url, headers=headers, follow_redirects=True).text items = [] for tag in BeautifulSoup(html, "html.parser").select('script[type="application/ld+json"]'): try: data = json.loads(tag.string) except (TypeError, json.JSONDecodeError): continue # skip empty or broken blocks for item in data if isinstance(data, list) else [data]: items.extend(item.get("@graph", [item])) return items
for item in json_ld("https://en.wikipedia.org/wiki/Web_scraping"): if item.get("@type") == "Article": print("Article |", item["headline"], "| created", item["datePublished"][:10])
for item in json_ld("https://www.allbirds.com/products/womens-allbirds-flip-flop-dusty-pink"): if item.get("@type") == "ProductGroup": # one product, with its sizes and colors as variants offer = item["offers"] variants = item["hasVariant"] print("ProductGroup |", item["name"], "|", offer["price"], offer["priceCurrency"], "|", len(variants), "variants") print(" first variant |", variants[0]["name"], "| SKU", variants[0]["sku"])Article | data scraping used for extracting data from websites | created 2005-09-17ProductGroup | Women's Allbirds Flip Flop | 25 USD | 35 variants first variant | Women's Allbirds Flip Flop - Dusty Pink - Size 5 | SKU A12513W050The Wikipedia result shows a useful detail. The article page doesn’t show when it was first created, but its JSON-LD does. It was created on September 17, 2005.
The Allbirds result shows a type many tutorials miss. A product with sizes and colors is often a ProductGroup, with each variant under hasVariant. Code that only looks for Product, as many tutorials do, finds nothing on these pages.
The try block matters in practice. Some sites leave JSON-LD tags empty or add JavaScript comments that aren’t valid JSON, and one broken block shouldn’t stop the rest.
On real product pages
Product data is where JSON-LD should help most, so I checked 8 product pages on well-known Shopify stores. Each store’s robots.txt allowed the page. I read the JSON-LD and OpenGraph tags with extruct (opens in a new tab), a library from Zyte that reads several kinds of structured data at once. Then I compared them with the product data Shopify itself serves (its product JSON, covered below). I ran it twice, and both runs matched.
| Store | JSON-LD product (blocks, price) | Price in OpenGraph | Shopify product JSON |
|---|---|---|---|
| Allbirds | 1 block, 25 | No | 25.00, 7 variants |
| ColourPop | 2 blocks, “5.0” and “5.00” | No | 5.00 |
| Kylie Cosmetics | None | 1.00 | 1.00 |
| Rothy’s | 1 block, no price | No | 0.00, 2 variants |
| Brooklinen | None | No | 21.00 |
| Death Wish Coffee | 1 block, 25.0 | No | 25.00 |
| Beardbrand | 1 block, 50.0 | 50.00 | 50.00 |
| Tentree | None | No | 88.00, 5 variants |
3 of the 8 product pages had no JSON-LD product at all. On one of them, Kylie Cosmetics, the price was in the OpenGraph tags instead, which is why a library that reads several formats helps. Where JSON-LD did have a price, it matched Shopify’s. Allbirds’ JSON-LD lists all 5 colors in every size, 35 variants, but gives full details only for the color on the page. Its Shopify JSON covers that one color, in 7 sizes.
Shopify’s product JSON
Shopify stores serve each product as JSON at the product page’s address plus .json. It had a price field for all 8 products, with every variant, though Rothy’s showed 0.00.
import httpx
headers = {"User-Agent": "MyScraperBot/1.0 (https://example.com/bot; you@example.com) httpx/0.28"}url = "https://www.tentree.com/products/morrell-sweater-warm-oak-nep.json"
product = httpx.get(url, headers=headers, follow_redirects=True).json()["product"]print(product["title"], "|", len(product["variants"]), "variants")for variant in product["variants"][:3]: print(" ", variant["title"], "|", variant["price"], "| SKU", variant["sku"])Morrell Sweater | 5 variants WARM OAK NEP / XS | 88.00 | SKU TCW6065-5634-XS WARM OAK NEP / S | 88.00 | SKU TCW6065-5634-S WARM OAK NEP / M | 88.00 | SKU TCW6065-5634-MTentree’s product page has no JSON-LD product, but this endpoint has its price and every size. Not every store allows it. Some stores I tried blocked these JSON endpoints with a 403 or 404, so check the store’s robots.txt and the response before you rely on it.
Where JSON-LD can mislead you
It’s written for search engines, not for you. Sites add the fields that help them in search results. A product page may include a price but leave out stock or shipping details that are visible on the page.
It can be missing, or repeated. 3 of the 8 product pages had no JSON-LD product, and ColourPop’s had 2, with the price written differently in each. Decide which block you trust, and compare a few pages against what a person sees.
Different sites use it differently. On one site author is a string, on another it’s an object, and on a third it’s a list. Prices came as 25, “25.0” or “5.00” in this test. The headline on Wikipedia is a short description, not the article title. Check the actual values before you map them to your own fields.
It can contain instructions for AI. A 2026 study found 1,996 prompt injections hidden in structured data, mostly in JSON-LD. If you pass JSON-LD to an LLM, treat it as untrusted text.
When to use it
JSON-LD is a good first check on any page. It’s quick to read, usually stays the same when a site is redesigned, and uses the same vocabulary across sites.
Use it for the fields it covers well, like names, dates, prices and ratings. When it’s missing, check in this order.
- OpenGraph tags, which a library like extruct reads in the same call.
- The platform’s own JSON, like Shopify’s product JSON endpoint.
- Other JSON in the page source.
- The HTML, before you use a browser.