beatrizalmeidaf/papero-pdf-text-extractor
Fast, open-source PDF text extraction API. Files never stored.
About the project
Open-source API and Python library for extracting PDF structure: reading order, tables, formulas, figures and block positions. Runs on CPU without ML models, files are never stored.
Useful for
- Extract tables from a research paper into CSV or Excel for analysis
- Convert a PDF to Markdown preserving reading order for RAG ingestion
- Run a local REST API via Docker for batch document processing
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 83 stars so far today, about 96 expected by the end of the day. The usual pace is 0 per day, so that's 96× as much.
- Before this, the repository barely got any stars — about 0 per day.
- Hacker News: “Lightweight PDF parser with layout, tables, formulas and bounding boxes” — 44 points, 4 h ago.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 76
- Today
- 83 · ≈ 96 by evening
- Forks
- 2
- Issues and pull requests
- 0
- Watchers
- 1
- Language
- Python
- License
- MIT
- Created
- April 12, 2025
- Last push
- October 1, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Hacker News discussions
- Lightweight PDF parser with layout, tables, formulas and bounding boxes44 points, October 1, 2026
Spotted in
- October 1, 2026Spotted on Hacker News
Similar by description
-
2.1
PaddlePaddle/PaddleOCR
OCR toolkit and document AI engine that turns PDFs and images into structured, LLM-ready data (JSON/Markdown), with support for 100+ languages, tables, formulas and seals.
-
2.3
docling-project/docling
Docling is a document parsing library for many formats (PDF, DOCX, PPTX, XLSX, HTML, EPUB, audio, video) with advanced PDF understanding: page layout, tables, formulas, OCR. It prepares structured data for generative…
-
6.6
Edwardxlai/easyread
A local app for reading English-language research papers: it imports PDFs, translates them page by page in the background, and shows the translation with formulas and tables, side-by-side original text, annotations,…
-
2.3
firecrawl/anydoc
Rust library that converts office documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, PDF) into clean Markdown, with Node.js, Python and browser bindings. Built to make documents LLM-ready.