About this project
Hey, I am Christoph. I am a Cloud Systems Engineer. When I am not deploying AWS infrastructure or doing heavy DIY work around the house, I build small tools. I like tools that solve exactly one problem and just work. That is how this converter started.
The problem
I needed clean Markdown from existing documentation. The workflow I wanted was simple: mirror a docs site with wget, drop the folder into a converter, and get structured Markdown for a RAG pipeline. No signup, no API keys. And I did not want to push internal company documents to third-party servers, which is not an option with client data.
Existing tools always had catches.
- Privacy: Almost all tools upload your files to the cloud. You cannot do that with internal docs or NDAs.
- Quality: Normal converters keep all the garbage. Navbars, cookie banners, and sidebars stick around. That noise burns LLM tokens for zero benefit.
- Scale: Clicking 500 HTML files one by one is not a workflow. Documentation lives in directory trees.
So I built what I needed. It runs entirely in your browser, filters out structural noise on its own, and chews through whole ZIP archives at once. The output is clean enough to paste straight into a context window.
How it works
Drop HTML files, PDFs, Word documents, or Excel spreadsheets into the window. The converter strips navigation, scripts, and boilerplate. You are left with clean Markdown for ChatGPT, Claude, or RAG pipelines. Everything happens directly in your browser's memory. No file ever leaves your machine.
Supported formats
- HTML: .html, .htm, and extensionless scrapes
- Word: .doc, .docx, .docm
- PowerPoint: .ppt, .pptx, .pptm, .pps, .ppsx
- Excel: .xls, .xlsx, .xlsm, .xlsb
- OpenDocument: .odt, .ods, .odp
- PDF: Text-based PDFs, plus built-in OCR for scanned pages
- Other: .rtf, .epub, .csv
Technology
- HTML conversion: Custom Rust WASM engine (52KB). Filters noise based on text density.
- Documents: anydoc by Firecrawl. Built from source and optimized for file size.
- OCR: PaddleOCR models running on ONNX Runtime Web. PDF.js renders the scanned pages.
- Frontend: Astro, Preact, and Tailwind CSS.
- ZIP handling: fflate for fast local extraction.
Privacy by architecture
After the initial page load, converting files makes no network requests: no uploads, no telemetry, no analytics, no cookies. The one exception is OCR. When you start it, your browser loads the model files from this site, and your PDF still stays on your machine. You do not have to take my word for it. Open the network tab in your browser's developer tools and watch.
Support this project
If this little tool saves you time, you can support the project in two ways.
Option 1: Cover hosting costs. You buy me a virtual coffee. This directly helps with the monthly server bills and keeps development going.
Option 2: Do something meaningful. I honestly prefer this option. As a father, projects that help children are very close to my heart. If you like this tool, please consider donating a small amount to Kinderhaus Weimar.
Other projects
- How To Lose Money Fast: a lottery simulator and analysis tool that shows why playing the lottery burns money.
- DruckZugPro: a free chimney cross-section pre-check calculator per DIN EN 13384-1.
- PatioPlanner: plan paved surfaces and calculate stones precisely.
- My Blog: homelab, self-hosting, and other technical projects.
- FusePlan: free browser-based planner for electrical sub-distribution boards. Drag breakers and RCDs onto DIN rails and export a parts list.