# run-llama/liteparse > A local Rust document parser that extracts text with bounding boxes from PDF, DOCX, XLSX, PPTX and images, with OCR via Tesseract or HTTP servers. Runs without cloud or LLMs, with bindings for Python, Node.js, WASM and a CLI. - Magnitude: 3.7 out of 10 — Breakout - Stars: 12,627 total · +10 stars today, ≈ 21 by evening - Star trust: star growth looks organic - Category: Data & analytics · Language: Rust · License: Apache-2.0 · Created: 2026-02-09 · Last push: 2026-09-22 - GitHub: https://github.com/run-llama/liteparse · Homepage: https://developers.llamaindex.ai/liteparse/ · Page: https://gitnova.dev/en/r/run-llama/liteparse ## Useful for - Extract text and bounding boxes from PDFs for downstream processing in Python or Node.js - Run scanned pages through Tesseract OCR and get structured Markdown for RAG - Detect document complexity before a full parse to decide whether OCR is needed ## Why it’s here - 10 stars so far today, about 21 expected by the end of the day. The usual pace is 7 per day, so that's 3× as much. - Over the last two days the pace is 3.5× that of the previous week and a half. - The spike has held for 3 days in a row — not a one-off blip. - GitHub Trending Rust today: #5, +44 stars. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 863 - Issues and pull requests: 466 - Watchers: 40 - Average over the last week: 45 per day - Usual pace: 7 per day - Stars in the last hour (measured): 3 - Latest release: node-v2.14.7 (2026-09-22) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-08-27 … 2026-09-25: 10, 7, 11, 9, 3, 9, 19, 10, 6, 5, 6, 8, 5, 9, 7, 9, 7, 4, 7, 5, 7, 3, 9, 4, 13, 9, 42, 175, 49, 10 ## Spotted in now - GitHub Trending Rust today: #5, +44 stars ## Similar by description 1. **Tencent/WeVisDoc** — 1.3 · Cooling · Language models · Python · +1 star today, ≈ 4 by evening WeVisDoc is an end-to-end document parser built on Qwen3-VL that converts page images into structured Markdown with LaTeX formulas and HTML tables. 2B and 4B versions are available, served via vLLM. Full card: https://gitnova.dev/en/r/Tencent/WeVisDoc.md 2. **docling-project/docling** — 3.3 · Peaking · Language models · Python · +32 stars today, ≈ 66 by evening Docling is a document parsing library for many formats (PDF, DOCX, PPTX, XLSX, HTML, EPUB, audio, video) with advanced PDF understanding: page layout, tables, formulas, OCR. It prepares structured data for generative AI and RAG pipelines. Full card: https://gitnova.dev/en/r/docling-project/docling.md 3. **firecrawl/anydoc** — 2.5 · Steady · Language models · Rust · +12 stars today, ≈ 24 by evening Rust library that converts office documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, PDF) into clean Markdown, with Node.js, Python and browser bindings. Built to make documents LLM-ready. Full card: https://gitnova.dev/en/r/firecrawl/anydoc.md 4. **opendatalab/MinerU** — 2.5 · Steady · Language models · Python · +15 stars today, ≈ 32 by evening MinerU is a Python tool that parses PDFs, images, and Office documents into LLM-ready Markdown/JSON. It supports OCR, layout analysis, tables and formulas, plus a local document library with search and citation locators. Full card: https://gitnova.dev/en/r/opendatalab/MinerU.md 5. **PaddlePaddle/PaddleOCR** — 2.5 · Steady · Language models · Python · +25 stars today, ≈ 43 by evening OCR toolkit and document AI engine that turns PDFs and images into structured, LLM-ready data (JSON/Markdown), with support for 100+ languages, tables, formulas and seals. Full card: https://gitnova.dev/en/r/PaddlePaddle/PaddleOCR.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-25 14:45 UTC, updated every 30 minutes.