Seismograph

What’s gaining stars on GitHub right now

beatrizalmeidaf/papero-pdf-text-extractor

Fast, open-source PDF text extraction API. Files never stored.

Data & analyticsPython#pdf#pdf-extraction#pdf-processing#pdf-to-text#pdf-tools
5.1 Early signal Magnitude out of 10 — how fast interest is growing, not a quality score. Star growth looks organic. Data as of October 1, 2026, 20:19 UTC.
Open on Seismograph Open on GitHub

About the project

Open-source API and Python library for extracting PDF structure: reading order, tables, formulas, figures and block positions. Runs on CPU without ML models, files are never stored.

Useful for

README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.

Why it’s trending

Stars per day

050100July 4, 2026October 1, 2026

Bars are daily stars, the line is the usual pace. Red marks spike days.

Numbers

Total stars
76
Today
83 · ≈ 96 by evening
Forks
2
Issues and pull requests
0
Watchers
1
Language
Python
License
MIT
Created
April 12, 2025
Last push
October 1, 2026

Star trust

Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.

These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.

Hacker News discussions

Spotted in

Share

README badge

Paste it into your README — the badge shows the current magnitude and links to this page.

Seismograph: 5.1
[![Seismograph](https://gitnova.dev/badge/beatrizalmeidaf/papero-pdf-text-extractor.svg?lang=en)](https://gitnova.dev/en/r/beatrizalmeidaf/papero-pdf-text-extractor)

Similar by description

  1. 2.1
    PaddlePaddle/PaddleOCR

    OCR toolkit and document AI engine that turns PDFs and images into structured, LLM-ready data (JSON/Markdown), with support for 100+ languages, tables, formulas and seals.

    SteadyLanguage modelsPython+16 stars today, ≈ 20 by evening

  2. 2.3
    docling-project/docling

    Docling is a document parsing library for many formats (PDF, DOCX, PPTX, XLSX, HTML, EPUB, audio, video) with advanced PDF understanding: page layout, tables, formulas, OCR. It prepares structured data for generative…

    CoolingLanguage modelsPython+32 stars today, ≈ 38 by evening

  3. 6.6
    Edwardxlai/easyread

    A local app for reading English-language research papers: it imports PDFs, translates them page by page in the background, and shows the translation with formulas and tables, side-by-side original text, annotations,…

    Early signalProductivity & self-hostedJavaScript+444 stars today, ≈ 512 by evening

  4. 2.3
    firecrawl/anydoc

    Rust library that converts office documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, PDF) into clean Markdown, with Node.js, Python and browser bindings. Built to make documents LLM-ready.

    SteadyLanguage modelsRust+33 stars today, ≈ 39 by evening