Focus area
RAG and retrieval pipelines
A RAG app is only as good as the web data inside it, and that data changes. Doccraft covers how to extract clean text, keep an index fresh on every re-crawl, and protect a pipeline from what pages hide.
Research and guides
- Common Crawl's WET text vs trafilatura vs Resiliparse, retested
Hugging Face's FineWeb team found Common Crawl's ready-made text too noisy and extracted their own. I retested that choice on 2,000 pages from the September 2026 crawl, with FineWeb's own quality filter.
- Hidden prompt injections that target scrapers: what a 2026 study found
A 2026 study found 15,387 prompt injections hidden in real web pages, aimed mostly at crawlers and scrapers. What they say, where they hide, what the top 1,000 sites do, and how to handle them.
Published for clients
- How to ground a LlamaIndex RAG app in fresh web data (opens in a new tab)blog.apify.com
- How to feed Amazon Bedrock Knowledge Bases with live web data using Bright Data (opens in a new tab)brightdata.com
- How to Build a Data Retriever With Dify AI Agents (opens in a new tab)scrapingbee.com
- How to Build a RAG Pipeline with Bright Data and Weaviate (opens in a new tab)brightdata.com
- Monitor website changes in Python without false alerts (opens in a new tab)evomi.com