Common Crawl's WET text vs trafilatura vs Resiliparse, retested
At Prague Crawl 2026, Alejandro AO, a developer advocate at Hugging Face, presented how the FineWeb dataset was built (opens in a new tab). FineWeb is an open dataset for training LLMs, made from Common Crawl. The first decision in its pipeline was where the text comes from.
Common Crawl provides ready-made text for the pages it crawls. The FineWeb team didn’t use it. In the talk’s words, “the wet format was actually suboptimal for training large language models”. They extracted the text from the raw HTML with trafilatura instead, and the talk called that step “one of the most expensive parts of the pipeline”.
I retested that choice on 2,000 pages from Common Crawl’s September 2026 crawl. I also added Resiliparse, a faster extractor that another open dataset, DataComp-LM, chose for the same job.
WARC and WET files
Common Crawl publishes each crawl in 2 main formats.
- WARC files contain the raw response for each page, including its full HTML.
- WET files contain plain text that Common Crawl extracted from the same pages, with every HTML tag removed.
WET is the easy option, because the text is ready to use. The FineWeb paper (opens in a new tab) explains the problem. The team “found that WET files retained too much boilerplate and menu text”. A model trained on their own trafilatura text performed better than one trained on WET text.
How I tested it
I took the first 2,000 HTML pages with a 200 status from one WARC file of the crawl, and the WET text for each of them. 1,972 of the pages had WET text. I used the 664 that Common Crawl’s language detection marked as English, because FineWeb is an English dataset.
Each page then had 3 versions of its text.
- WET, the text Common Crawl provides.
- trafilatura 2.3.1 on the page’s HTML, with the settings FineWeb’s pipeline uses.
- Resiliparse 1.0.9 on the same HTML, extracting the main content only.
To judge them, I didn’t invent a score. I ran each version through FineWeb’s own quality filter, from Hugging Face’s datatrove (opens in a new tab) library. It removes a page when too few lines end with punctuation, too many lines are short, or lines repeat.
This script reads one page from Common Crawl and extracts it both ways. The index tells you which WARC file contains the page and where, so you download only that record.
import gzipimport json
import httpximport trafilaturafrom resiliparse.extract.html2text import extract_plain_textfrom resiliparse.parse.encoding import bytes_to_str, detect_encoding
CRAWL = "CC-MAIN-2026-39"url = "airshipdaily.com/blog/on-the-fashion-fence-canvas-doc-martens"
# 1. Ask the Common Crawl index where the page is stored.index = httpx.get(f"https://index.commoncrawl.org/{CRAWL}-index", params={"url": url, "output": "json"}, timeout=60)hit = json.loads(index.text.splitlines()[0])
# 2. Download only that record from the WARC file, with a range request.start, length = int(hit["offset"]), int(hit["length"])record = httpx.get( f"https://data.commoncrawl.org/{hit['filename']}", headers={"Range": f"bytes={start}-{start + length - 1}"}, timeout=60,).contenthtml = gzip.decompress(record).split(b"\r\n\r\n", 2)[2] # skip the WARC and HTTP headers
# 3. Extract the main text both ways.traf = trafilatura.extract(html, favor_precision=True, include_comments=False, deduplicate=True)resi = extract_plain_text(bytes_to_str(html, detect_encoding(html)), main_content=True)for name, text in [("trafilatura", traf), ("resiliparse", resi)]: print(f"{name:12} {len(text.split()):4} words | {text.strip().splitlines()[0][:60]}")trafilatura 532 words | As someone who came of age when grunge came a’ knockin’, I hresiliparse 581 words | On The Fashion Fence: Canvas Doc MartensThe WET text for this page has 732 words. The extra words are the site’s menu, its tagline and links to other posts.
The results
| Text source | English pages that passed FineWeb’s filter | Words on the pages that passed | Time to extract all 1,972 pages |
|---|---|---|---|
| WET | 68 of 664 (10%) | 85,002 | Ready-made |
| trafilatura | 336 of 664 (51%) | 170,495 | 30.0 s |
| Resiliparse | 224 of 664 (34%) | 170,387 | 1.8 s |
WET text is the longest, with a median of 392 words a page, but most WET pages didn’t pass the filter. After filtering, it kept half as many words as either extractor. The most common reason was lines without punctuation, which is what menus and link lists look like.
Trafilatura passed the most pages. Resiliparse passed fewer, but kept almost exactly as many words, and it was about 17 times faster, running in a single process on an Apple M3.
What WET text looks like
A business directory page from the American Chamber of Commerce in Thailand is a typical case. Its WET text starts like this.
Local Market Advisory & Business Research Category | The American Chamber of Commerce in Thailand - AMCHAM Member DirectoryFacebookLinkedInAbout UsCommitteesContact UsSubscribeHomeEventsThe menu continues for dozens of lines, and parts of it appear twice. FineWeb’s filter removed the page for having too many short lines. Trafilatura returned 129 words from the page without the menu, and that version passed.
Where each extractor gets it wrong
Neither extractor was right every time. I read the outputs for the pages where they disagreed.
- Trafilatura returned nothing for 86 English pages. Most had no real main text, like product grids, login pages, error pages and spam. About 8 were real posts or articles, and 6 of those were on Blogger blogs. FineWeb’s settings tell trafilatura to favor precision. With trafilatura’s default settings, 6 of the 8 came back, and the other 2 were empty either way.
- Trafilatura sometimes keeps a sidebar. On a comic shop’s archive page, its text started with the shop’s newsletter box, before the actual post.
- Resiliparse keeps more of the page around the content, like titles, dates and bylines. In the pages I checked, the title often appeared twice, as the heading and again in the text. FineWeb’s repeated-line check removes a page when repeated lines are more than 1% of its characters, so a repeated title alone can remove it.
- A filter isn’t a spam detector. One spam page full of scrambled sentences passed the filter in its Resiliparse version. Datasets like FineWeb use more steps after this one, including deduplication and quality classifiers.
Which one to use
The DataComp-LM paper (opens in a new tab) tested the same 3 options by training models on each. Resiliparse and trafilatura both did better than WET text, and “have similar downstream performance”. Resiliparse was about 8 times faster, so the DataComp-LM team used it.
My results agree.
- Don’t use WET text as it is for training data, RAG or analysis. In my sample, it had more than twice as many words as trafilatura’s text, and the extra words I read were mostly menus and links.
- Use Resiliparse at Common Crawl scale. It keeps about the same amount of usable text, and on billions of pages, the speed difference has a big effect on compute cost.
- Use trafilatura when each page matters more than speed, like a few thousand pages for a RAG index, because more of its pages pass the filter as they are.
The talk suggested that teams without the budget for custom extraction “can skip that”. With Resiliparse, the extraction step needs much less compute, so there’s less reason to skip it.
Limits of this test
- One WARC file from one crawl. The pages are the first 2,000 HTML pages with a 200 status in that file, not a random sample of the web.
- English only, using Common Crawl’s own language tags.
- The filter may favor trafilatura. FineWeb designed its filter on trafilatura text, so a higher pass rate doesn’t prove trafilatura is better. DataComp-LM’s training results are the stronger evidence.
- Passing a filter isn’t model quality. This test measures how much text is kept, not how good a model trained on it would be.
Data and code
- Results for the 1,972 pages
- The sampling script, which downloads the pages again
- The comparison script
- The single-page example