Focus area
Data extraction and structured data
The data a scraper needs is often already in the page, as JSON, JSON-LD or clean HTML, before any browser or LLM is involved. Doccraft shows where to find it, how to extract it reliably, and when an LLM is worth the cost.
Research and guides
- How to find the JSON data hidden in a page's source
Many pages send their data as JSON inside the HTML. How to find it, the patterns to look for, and working Python code to read it without a browser.
- JSON-LD for scrapers: what it contains and what it misses
More than half of the top homepages that returned content in my test include JSON-LD, structured data written for search engines. What it contains, how to read it in Python, and what it missed on 8 real product pages.
- Common Crawl's WET text vs trafilatura vs Resiliparse, retested
Hugging Face's FineWeb team found Common Crawl's ready-made text too noisy and extracted their own. I retested that choice on 2,000 pages from the September 2026 crawl, with FineWeb's own quality filter.
- What Oxylabs learned from its AI parser agent
At OxyCon 2026, Oxylabs openly shared why its engineers stopped using an AI parser agent, and what its second agent did differently.
Published for clients
- Auto Parts and Tire Data: Pricing, Fitment, and Matching (opens in a new tab)brightdata.com
- Zero-Shot E-Commerce Scraping: Call the LLM Last (opens in a new tab)scrapingbee.com
- Cursor + Bright Data vs a default coding agent setup: building a real price tracker (opens in a new tab)brightdata.com
- Monitor website changes in Python without false alerts (opens in a new tab)evomi.com
- Web Scraping with Scrapling: A Python Tutorial (2026) (opens in a new tab)brightdata.com