HTML→Markdown

Convert Scraped HTML to Markdown for AI Pipelines

Web scraping gives you raw HTML: navigation, ads, cookie-consent dialogs, and layout markup that your AI pipeline doesn't need. Cleaning scraped pages by hand doesn't scale once you're processing hundreds or thousands of them.

The converter scores each part of a page by text density to find the real content. It drops the navigation bars, footers, sidebars, and promotional blocks, then converts what's left to GitHub-Flavored Markdown.

The problem with raw scraped HTML

  • Noise ratio: a typical web page is 70-80% boilerplate and only 20-30% actual content
  • Inconsistent structure: every site uses a different layout, so regex-based extraction breaks easily
  • Token waste: feeding raw HTML to an LLM spends tokens on <nav>, <footer>, and inline styles
  • Embedding quality: vector embeddings of noisy HTML retrieve poorly

How it works

  1. Scrape your target: use wget, Scrapy, Puppeteer, or any crawler to download pages as .html files
  2. ZIP the output: bundle the scraped HTML into a ZIP archive (up to 200 files, 100 MB)
  3. Convert in-browser: the WASM engine scores each container element by text density, picks the content block, and outputs GFM Markdown
  4. Download results: get clean .md files that keep the original directory structure

Works with any scraper

  • wget -r, recursive site mirroring
  • Scrapy, the Python crawling framework
  • Puppeteer / Playwright, headless browser scraping (save as HTML)
  • curl + scripts, custom URL list downloads
  • HTTrack, the website copier

Good for

  • Building RAG (Retrieval-Augmented Generation) knowledge bases from public documentation
  • Creating training datasets for domain-specific models
  • Migrating web content to Markdown-based systems (Obsidian, Notion, MkDocs)
  • Archiving websites in a readable, searchable format

Privacy guarantee

All processing runs locally in your browser through WebAssembly. Your scraped content is never uploaded to any server, which makes it safe for proprietary data, internal wikis, and sensitive material.

Related