Fewer tokens, same meaning
Boilerplate, scripts and layout markup are stripped, so your context window holds content and not clutter.
Free · Private · Nothing is uploaded
Drop PDFs, Word files, slides, spreadsheets or whole website exports. Get clean, LLM-ready Markdown in seconds, right in your browser.
Initializing conversion engine…
Navigation, cookie banners and div soup disappear. Headings, lists and tables stay.
Boilerplate, scripts and layout markup are stripped, so your context window holds content and not clutter.
Headings, lists, links and tables come out as GitHub-Flavored Markdown that models read well.
WebAssembly does the work on your machine. No account, no upload, no server copy of your files.
There is no limit on how many rounds you run. Every result stays on this page until you clear it.
Back to the drop zoneGot a feature request, found a bug, or just want to share feedback?
Also listed on Twelve
Have a documentation site you want to feed to an LLM? Mirror it withwget, zip the result, and convert the whole set in one pass.
Downloads a documentation section recursively, staying inside the starting path and rewriting links so the copies work offline.
wget -r -np -k -E \
--reject-regex '\.(png|jpe?g|gif|svg|webp|woff2?|ttf|css|js|zip|pdf)$' \
-w 0.5 --random-wait \
https://community.denodo.com/docs/html/browse/latest/en/vdp/developer/indexwget writes into a folder named after the host, e.g. community.denodo.com/. On macOS install it first with brew install wget.
Bundle only the .html files. Nested folders are fine; the converter reads them at any depth and mirrors the paths in the output.
cd community.denodo.com
# every .html file, directory structure preserved
find . -name '*.html' -print | zip ../docs-html.zip -@Prefer everything as-is? zip -r ../docs-html.zip . works too; non-HTML entries are ignored on upload.
Upload docs-html.zip into the converter. Each page becomes a .md file with the same name and path, ready to download as a single archive.
# check the archive stays inside the limits first
unzip -l docs-html.zip | grep -c '\.html$' # ≤ 200 files
du -h docs-html.zip # ≤ 100 MBInstall wget with winget install JernejSimoncic.wget, then zip with PowerShell instead of the zip command.
Compress-Archive -Path .\community.denodo.com\* -DestinationPath docs-html.zipOnly crawl sites whose terms permit it, and keep the request rate low. Large manuals can exceed the 200-file limit per archive; narrow the starting URL or lower-l and convert in batches.
Read the full guide →It also covers curl, browser saving, Puppeteer, HTTrack, and more.