Seismograph

What’s gaining stars on GitHub right now

run-llama/liteparse

A fast, helpful, and open-source document parser

Data & analyticsRust#document-ocr#document-processing#ocr#ocr-recognition#pdf
3.7 Breakout Magnitude out of 10 — how fast interest is growing, not a quality score. Star growth looks organic. Data as of September 25, 2026, 13:57 UTC.
Open on Seismograph Open on GitHub

About the project

A local Rust document parser that extracts text with bounding boxes from PDF, DOCX, XLSX, PPTX and images, with OCR via Tesseract or HTTP servers. Runs without cloud or LLMs, with bindings for Python, Node.js, WASM and a CLI.

Useful for

README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.

Why it’s trending

Stars per day

0100200June 28, 2026September 25, 2026

Bars are daily stars, the line is the usual pace. Red marks spike days.

Numbers

Total stars
12,625
Today
10 · ≈ 23 by evening
Forks
863
Issues and pull requests
466
Watchers
40
Language
Rust
License
Apache-2.0
Latest release
node-v2.14.7 · September 22, 2026
Created
February 9, 2026
Last push
September 22, 2026

Star trust

Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.

These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.

Spotted in

Share

README badge

Paste it into your README — the badge shows the current magnitude and links to this page.

Seismograph: 3.7
[![Seismograph](https://gitnova.dev/badge/run-llama/liteparse.svg?lang=en)](https://gitnova.dev/en/r/run-llama/liteparse)

Similar by description

  1. 1.3
    Tencent/WeVisDoc

    WeVisDoc is an end-to-end document parser built on Qwen3-VL that converts page images into structured Markdown with LaTeX formulas and HTML tables. 2B and 4B versions are available, served via vLLM.

    CoolingLanguage modelsPython+1 star today, ≈ 5 by evening

  2. 3.4
    docling-project/docling

    Docling is a document parsing library for many formats (PDF, DOCX, PPTX, XLSX, HTML, EPUB, audio, video) with advanced PDF understanding: page layout, tables, formulas, OCR. It prepares structured data for generative…

    PeakingLanguage modelsPython+30 stars today, ≈ 68 by evening

  3. 2.5
    firecrawl/anydoc

    Rust library that converts office documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, PDF) into clean Markdown, with Node.js, Python and browser bindings. Built to make documents LLM-ready.

    SteadyLanguage modelsRust+10 stars today, ≈ 23 by evening

  4. 2.5
    opendatalab/MinerU

    MinerU is a Python tool that parses PDFs, images, and Office documents into LLM-ready Markdown/JSON. It supports OCR, layout analysis, tables and formulas, plus a local document library with search and citation locators.

    SteadyLanguage modelsPython+14 stars today, ≈ 33 by evening

  5. 2.5
    PaddlePaddle/PaddleOCR

    OCR toolkit and document AI engine that turns PDFs and images into structured, LLM-ready data (JSON/Markdown), with support for 100+ languages, tables, formulas and seals.

    SteadyLanguage modelsPython+21 stars today, ≈ 40 by evening

  6. 1.4
    tesseract-ocr/tesseract

    Open-source OCR engine libtesseract and CLI tool tesseract for recognizing text in images. Supports 100+ languages, PNG/JPEG/TIFF input, and output in text, hOCR, PDF, TSV, ALTO.

    SteadyLanguage modelsC+++3 stars today, ≈ 6 by evening