How to Download HTML Pages for Conversion
Before you can convert HTML to Markdown, you need the files on your machine. There are several ways to get them, from a one-click browser save to a full recursive site mirror. Pick the one that fits your situation.
Method 1: recursive site mirroring with wget
Best for: documentation sites, product manuals, and developer references, any site with a clear URL hierarchy.
Basic command
wget -r -np -k -E \
-w 0.5 --random-wait \
https://docs.example.com/en/latest/What the flags do
| Flag | Purpose |
|---|---|
-r | Follow links recursively |
-np | Never ascend above the starting path (stay inside the docs section) |
-k | Rewrite links to point at local files |
-E | Append .html to extensionless URLs so the converter can detect them |
-w 0.5 | Wait 0.5 seconds between requests |
--random-wait | Randomize the delay so it looks less automated |
Skip images and assets (recommended)
You only need the HTML. Reject binary files to save bandwidth and disk space:
wget -r -np -k -E \
--reject-regex '\.(png|jpe?g|gif|svg|webp|woff2?|ttf|css|js|zip|pdf)$' \
-w 0.5 --random-wait \
https://community.denodo.com/docs/html/browse/latest/en/vdp/developer/indexLimit crawl depth
For large sites, limit how deep wget follows links:
# Only 2 levels deep from the start URL
wget -r -np -k -E -l 2 \
-w 0.5 --random-wait \
https://docs.example.com/en/latest/Installing wget
| System | Command |
|---|---|
| macOS | brew install wget |
| Ubuntu/Debian | sudo apt install wget |
| Windows | winget install JernejSimoncic.wget |
Method 2: download a list of URLs with curl
Best for: when you already have a specific list of pages (from a sitemap, a table of contents, or a URL file).
Download from a URL list
# urls.txt: one URL per line
https://docs.example.com/page-1.html
https://docs.example.com/page-2.html
https://docs.example.com/api/reference.html# Download all pages, preserve directory structure
while IFS= read -r url; do
path=$(echo "$url" | sed 's|https\?://[^/]*/||')
mkdir -p "$(dirname "downloaded/$path")"
curl -s -o "downloaded/$path" "$url"
sleep 0.5
done < urls.txtExtract URLs from a sitemap
# Pull URLs from sitemap.xml
curl -s https://docs.example.com/sitemap.xml \
| grep -oP '<loc>\K[^<]+' \
| grep '\.html' \
> urls.txtMethod 3: save individual pages from the browser
Best for: a handful of pages, or sites that require a login.
- Open the page in Chrome, Firefox, or Edge
- Press Ctrl+S (Windows/Linux) or ⌘+S (macOS)
- Choose "Web Page, HTML Only" (not "Complete"; you don't need images or CSS)
- Save to a folder
- Repeat for each page, or use a browser extension for batch saving
Browser extensions for batch saving
- SingleFile (Chrome/Firefox): saves pages as self-contained HTML files
- Save All Resources (Chrome): bulk-saves open tabs as HTML
- DownThemAll (Firefox): a download manager that can grab every link on a page
Method 4: HTTrack, a GUI website copier
Best for: people who prefer a graphical interface over the command line.
- Download HTTrack from
httrack.com(Windows/Linux/macOS) - Enter the starting URL and set filters to HTML only
- Let it mirror the site
- Find the HTML files in the output folder
Method 5: Puppeteer or Playwright for JavaScript-rendered pages
Best for: single-page apps and sites that render content with JavaScript. wget and curl only fetch the initial HTML; if the content loads dynamically, you need a headless browser.
// save-pages.mjs: Node.js + Playwright
import { chromium } from 'playwright';
import { writeFileSync, mkdirSync } from 'fs';
const urls = [
'https://spa-docs.example.com/getting-started',
'https://spa-docs.example.com/api-reference',
];
const browser = await chromium.launch();
const page = await browser.newPage();
for (const url of urls) {
await page.goto(url, { waitUntil: 'networkidle' });
const html = await page.content();
const filename = url.split('/').pop() + '.html';
mkdirSync('pages', { recursive: true });
writeFileSync(`pages/${filename}`, html);
console.log(`Saved: ${filename}`);
}
await browser.close();Creating the ZIP package
Once you have your HTML files downloaded, bundle them into a ZIP for the converter:
macOS and Linux
# From inside the downloaded folder:
cd community.denodo.com
# ZIP only HTML files, preserving directory structure
find . -name '*.html' -print | zip ../docs.zip -@Or zip everything (non-HTML files are ignored by the converter):
zip -r ../docs.zip .Windows (PowerShell)
Compress-Archive -Path .\downloaded\* -DestinationPath docs.zipWindows (File Explorer)
- Select the folder with your HTML files
- Right-click → Compress to ZIP file
Checking your ZIP before upload
The converter accepts ZIPs up to 100 MB with up to 200 HTML files:
# Count HTML files
unzip -l docs.zip | grep -c '\.html$'
# Check total size
du -h docs.zipIf your archive exceeds the limits, split it by subdirectory and convert in batches.
Tips and practices
- Respect robots.txt: check the site's robots.txt before crawling. wget honors it by default (
-e robots=on) - Throttle requests: always put a delay between requests (
-w 0.5at a minimum) so you don't overload the server - Check the terms of service: some sites prohibit automated downloading
- Use
-Ewith wget: it saves extensionless URLs as.htmlfiles that the converter can detect - Test a small section first: add
-l 1to limit depth and check the output before a full crawl - Login-protected sites: use
wget --load-cookieswith exported browser cookies, or save the pages by hand in the browser
Next step
Once your ZIP file is ready, open the converter and drop it in. The conversion happens entirely in your browser; your files are never uploaded to any server.