Clean HTML for AI Training Data
Building a fine-tuning dataset or training corpus from web content? Getting the HTML is the easy part; cleaning it is the hard one. Every page comes wrapped in navigation, ads, cookie dialogs, related-article widgets, and layout markup that ends up in your training data.
This tool uses density-based content extraction to find and pull the main content from each page. The output is Markdown that matches what a person would call "the article."
Why clean data matters for AI training
- Quality in, quality out: models trained on noisy data learn to reproduce the noise (navigation text, cookie-consent phrases, "click here" patterns)
- Deduplication: boilerplate text (headers, footers, nav) repeats across thousands of pages, filling your corpus with unintended duplicates
- Token efficiency: with content-only text, every token teaches the domain rather than the page chrome
- Better evaluation: clean test sets give you honest metrics on how the model does with real content
The density extraction algorithm
The converter scores each container element (<div>, <section>,<article>, <main>) on three things:
- Text density: ratio of visible text to HTML markup inside the container
- Structural signals: headings, paragraphs, and semantic elements
- Noise indicators: navigation links, form elements, and the repetitive patterns typical of boilerplate
The highest-scoring container becomes the content block. Everything else (nav, footer, sidebar, ads) is dropped before conversion.
How it works
- Collect your HTML: crawl, scrape, or export the pages you want in your training set
- Batch convert: ZIP up to 200 HTML files and upload them (up to 100 MB per archive)
- Download clean Markdown: each file is processed on its own, so one bad page doesn't block the rest
- Use in your pipeline: feed the
.mdfiles into your tokenizer, chunker, or fine-tuning script
Compared to other approaches
| Approach | Pros | Cons |
|---|---|---|
| Regex stripping | Fast | Brittle, site-specific, misses semantic structure |
| BeautifulSoup + heuristics | Flexible | Requires per-site tuning, slow on large batches |
| Readability.js | Good for articles | Designed for single pages, not batch processing |
| This tool (density extraction) | Works on any site without tuning, batch processing, GFM output | May not handle very unusual layouts |
Practical tips for training data
- Convert in batches by domain or topic to keep each set thematically consistent
- Review a random sample of outputs before feeding them to your training pipeline
- Filter out empty results (pages with no extractable content); these are usually index or listing pages
- Combine with metadata (source URL, date, category) for richer datasets
Privacy
Processing happens entirely in your browser via WebAssembly. Your training data is never uploaded to any server. This is especially important when working with proprietary content, licensed datasets, or data subject to agreements.