# beatrizalmeidaf/papero-pdf-text-extractor > Open-source API and Python library for extracting PDF structure: reading order, tables, formulas, figures and block positions. Runs on CPU without ML models, files are never stored. - Magnitude: 5.4 out of 10 — Early signal - Stars: 102 total · +107 stars today, ≈ 116 by evening - Star trust: star growth looks organic - Category: Data & analytics · Language: Python · License: MIT · Created: 2025-04-12 · Last push: 2026-10-01 - GitHub: https://github.com/beatrizalmeidaf/papero-pdf-text-extractor · Homepage: https://beatrizalmeidaf.github.io/papero-pdf-text-extractor/ · Page: https://gitnova.dev/en/r/beatrizalmeidaf/papero-pdf-text-extractor ## Useful for - Extract tables from a research paper into CSV or Excel for analysis - Convert a PDF to Markdown preserving reading order for RAG ingestion - Run a local REST API via Docker for batch document processing ## Why it’s here - 107 stars so far today, about 116 expected by the end of the day. The usual pace is 0 per day, so that's 116× as much. - Before this, the repository barely got any stars — about 0 per day. - Hacker News: “Lightweight PDF parser with layout, tables, formulas and bounding boxes” — 59 points, 5 h ago. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 2 - Issues and pull requests: 0 - Watchers: 1 - Average over the last week: 17 per day - Usual pace: 0 per day - Stars in the last hour (measured): 15 ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-02 … 2026-10-01: 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 107 ## Hacker News - Lightweight PDF parser with layout, tables, formulas and bounding boxes — 59 points, 7 comments: https://news.ycombinator.com/item?id=49923638 ## Spotted in now - Spotted on Hacker News ## Similar by description 1. **PaddlePaddle/PaddleOCR** — 2.1 · Steady · Language models · Python · +17 stars today, ≈ 19 by evening OCR toolkit and document AI engine that turns PDFs and images into structured, LLM-ready data (JSON/Markdown), with support for 100+ languages, tables, formulas and seals. Full card: https://gitnova.dev/en/r/PaddlePaddle/PaddleOCR.md 2. **docling-project/docling** — 2.4 · Cooling · Language models · Python · +36 stars today, ≈ 39 by evening Docling is a document parsing library for many formats (PDF, DOCX, PPTX, XLSX, HTML, EPUB, audio, video) with advanced PDF understanding: page layout, tables, formulas, OCR. It prepares structured data for generative AI and RAG pipelines. Full card: https://gitnova.dev/en/r/docling-project/docling.md 3. **Edwardxlai/easyread** — 6.4 · Early signal · Productivity & self-hosted · JavaScript · +454 stars today, ≈ 493 by evening A local app for reading English-language research papers: it imports PDFs, translates them page by page in the background, and shows the translation with formulas and tables, side-by-side original text, annotations, and AI Q&A. All data… Full card: https://gitnova.dev/en/r/Edwardxlai/easyread.md 4. **firecrawl/anydoc** — 2.4 · Steady · Language models · Rust · +44 stars today, ≈ 48 by evening Rust library that converts office documents (Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, PDF) into clean Markdown, with Node.js, Python and browser bindings. Built to make documents LLM-ready. Full card: https://gitnova.dev/en/r/firecrawl/anydoc.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-01 21:55 UTC, updated every 30 minutes.