HTML→Markdown

Free · Private · Nothing is uploaded

Any document to Markdown for AI

Drop PDFs, Word files, slides, spreadsheets or whole website exports. Get clean, LLM-ready Markdown in seconds, right in your browser.

PDFDOCXHTMLXLSXPPTXEPUBCSVRTF

Initializing conversion engine…

Messy in. Clean out.

Navigation, cookie banners and div soup disappear. Headings, lists and tables stay.

Source pageExample
<nav>Home · Pricing · Login</nav>
<div class="cookie">We use cookies…</div>
<h1>Quarterly report</h1>
<div><span>Revenue grew 12%</span></div>
<ul><li>EMEA</li><li>APAC</li></ul>
<footer>© Acme · Imprint · Privacy</footer>
Markdown.md
# Quarterly report
 
Revenue grew 12%
 
- EMEA
- APAC

Fewer tokens, same meaning

Boilerplate, scripts and layout markup are stripped, so your context window holds content and not clutter.

Structure that survives

Headings, lists, links and tables come out as GitHub-Flavored Markdown that models read well.

Private by design

WebAssembly does the work on your machine. No account, no upload, no server copy of your files.

Popular use cases

View all →

Convert the next batch

There is no limit on how many rounds you run. Every result stays on this page until you clear it.

Back to the drop zone

Got a feature request, found a bug, or just want to share feedback?

html-to-markdown-ai - Bulk-convert HTML to Markdown for AI - private, in-browser | Product Hunt

Also listed on Twelve

How to download HTML pages and create a ZIP package

How to download HTML pages and create a ZIP package

Have a documentation site you want to feed to an LLM? Mirror it withwget, zip the result, and convert the whole set in one pass.

1. Mirror the pages with wget

Downloads a documentation section recursively, staying inside the starting path and rewriting links so the copies work offline.

wget -r -np -k -E \
  --reject-regex '\.(png|jpe?g|gif|svg|webp|woff2?|ttf|css|js|zip|pdf)$' \
  -w 0.5 --random-wait \
  https://community.denodo.com/docs/html/browse/latest/en/vdp/developer/index

wget writes into a folder named after the host, e.g. community.denodo.com/. On macOS install it first with brew install wget.

2. Pack the HTML files into a ZIP

Bundle only the .html files. Nested folders are fine; the converter reads them at any depth and mirrors the paths in the output.

cd community.denodo.com

# every .html file, directory structure preserved
find . -name '*.html' -print | zip ../docs-html.zip -@

Prefer everything as-is? zip -r ../docs-html.zip . works too; non-HTML entries are ignored on upload.

3. Drop the ZIP above

Upload docs-html.zip into the converter. Each page becomes a .md file with the same name and path, ready to download as a single archive.

# check the archive stays inside the limits first
unzip -l docs-html.zip | grep -c '\.html$'   # ≤ 200 files
du -h docs-html.zip                          # ≤ 100 MB

What the wget flags do

-r
Follow links recursively
-np
Never ascend above the starting path
-k
Rewrite links to point at the local copies
-E
Append .html so every page is detected
-p
Optional: also grab CSS and images, not needed for Markdown
-l 2
Limit crawl depth, useful for large sites
-w 0.5 --random-wait
Throttle requests to stay polite

On Windows

Install wget with winget install JernejSimoncic.wget, then zip with PowerShell instead of the zip command.

Compress-Archive -Path .\community.denodo.com\* -DestinationPath docs-html.zip

Only crawl sites whose terms permit it, and keep the request rate low. Large manuals can exceed the 200-file limit per archive; narrow the starting URL or lower-l and convert in batches.

Read the full guide →It also covers curl, browser saving, Puppeteer, HTTrack, and more.