MarkItDown
A privacy-first, serverless client-side document converter that instantly transforms local files into clean, LLM-ready Markdown text without external cloud uploads.
Our Role
Lead Developer, UI/UX Designer, Full Stack Developer
Technologies & Frameworks
Project Overview
MarkItDown is a privacy-first web application designed to convert document formats like PDF, DOCX, PPTX, XLSX, CSV, HTML, XML, JSON, and EPUB into clean, structured Markdown. Built to satisfy the data input needs of modern Large Language Models (LLMs), it focuses on utility, data privacy, and a bold Neo-Brutalist design. All core parsing algorithms execute 100% locally in the user's browser, bypassing the need to upload files to external servers. An optional integration with Supabase provides persistent history tracking for premium accounts, while a seamless postMessage bridge securely connects the web app with the Companion Extension.
The Challenge & Problem
- Cloud-based File Exposure: Uploading corporate documents or proprietary PDFs to external servers for conversion risks data leaks, violating compliance guidelines.
- LLM Input Noise (Token Waste): Raw document conversions often inject verbose HTML styles, empty table cells, layout grids, or duplicate line breaks, causing LLMs to waste valuable context window tokens.
- Complex File Formats (EPUB, PPTX, XLSX): Reading slides, multi-sheet spreadsheets, or packaged book volumes directly in the browser is traditionally slow and lacks structural mapping.
The Engineering Solution
- Client-Side Local Processing: Implemented document parsing libraries (pdfjs-dist for PDF, mammoth for DOCX, xlsx for Excel, jszip for slide/EPUB contents) to execute entirely inside the client browser sandbox using Web Workers and local array buffers, ensuring zero server-side exposure.
- Custom HTML-to-Markdown and Text Sanitization Pipeline: Developed a rule-based parser that cleans metadata, merges fragmented layout lines, strips inline styling, formats multi-sheet tables with correct alignment, and outputs standard GFM (GitHub Flavored Markdown) to save up to 93% on token consumption.
- Binary Archive and DOM Walking Parsers: Created a slide extractor that reads the internal XML structure of slide objects using jszip, matching slide titles to headings, and an EPUB parser that traverses the container.xml spine list to compile chapters sequentially in Markdown.
Key Architecture & Features
100% Browser-Native Conversion Suite
- Direct drag-and-drop parsing of PDF, DOCX, XLSX, PPTX, CSV, JSON, XML, HTML, EPUB, and TXT files locally.
Advanced PDF Table Detection
- Evaluates Y-coordinate positions of text strings to dynamically reconstruct structured tables from PDF documents.
Interactive Token Estimation Sandbox
- A calculator on the landing page that estimates raw token sizes, clean Markdown token sizes, and dollar savings based on pricing models.
Supabase-Backed Conversion History
- Optional account sync that saves processed metadata (file size, page counts, processing speeds) and clean Markdown outputs in a PostgreSQL database under Row-Level Security (RLS).
One-Click LLM Export & Copy
- Actions to instantly copy Markdown code or deep-link redirect to popular AI interfaces (Claude, ChatGPT, Gemini) pre-filled with the converted text.
Neo-Brutalist Dashboard
- A bold, high-contrast, utility-first workspace characterized by stark borders, bright accents, and sharp corners.