Convert Scraped HTML to Markdown for AI Pipelines
Web scraping gives you raw HTML: navigation, ads, cookie-consent dialogs, and layout markup that your AI pipeline doesn't need. Cleaning scraped pages by hand doesn't scale once you're processing hundreds or thousands of them.
The converter scores each part of a page by text density to find the real content. It drops the navigation bars, footers, sidebars, and promotional blocks, then converts what's left to GitHub-Flavored Markdown.
The problem with raw scraped HTML
- Noise ratio: a typical web page is 70-80% boilerplate and only 20-30% actual content
- Inconsistent structure: every site uses a different layout, so regex-based extraction breaks easily
- Token waste: feeding raw HTML to an LLM spends tokens on
<nav>,<footer>, and inline styles - Embedding quality: vector embeddings of noisy HTML retrieve poorly
How it works
- Scrape your target: use wget, Scrapy, Puppeteer, or any crawler to download pages as
.htmlfiles - ZIP the output: bundle the scraped HTML into a ZIP archive (up to 200 files, 100 MB)
- Convert in-browser: the WASM engine scores each container element by text density, picks the content block, and outputs GFM Markdown
- Download results: get clean
.mdfiles that keep the original directory structure
Works with any scraper
wget -r, recursive site mirroring- Scrapy, the Python crawling framework
- Puppeteer / Playwright, headless browser scraping (save as HTML)
- curl + scripts, custom URL list downloads
- HTTrack, the website copier
Good for
- Building RAG (Retrieval-Augmented Generation) knowledge bases from public documentation
- Creating training datasets for domain-specific models
- Migrating web content to Markdown-based systems (Obsidian, Notion, MkDocs)
- Archiving websites in a readable, searchable format
Privacy guarantee
All processing runs locally in your browser through WebAssembly. Your scraped content is never uploaded to any server, which makes it safe for proprietary data, internal wikis, and sensitive material.