# allenai/olmocr > Toolkit that converts PDF, PNG and JPEG documents into clean Markdown, preserving equations, tables and handwriting, powered by a 7B VLM. Used to prepare PDF documents for LLM datasets and training. - Magnitude: 2.0 out of 10 — Steady - Stars: 19,723 total · +3 stars measured 2026-10-07, 11:04–12:47 UTC, ≈ 9 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: Apache-2.0 · Created: 2024-09-17 · Last push: 2026-03-25 - GitHub: https://github.com/allenai/olmocr · Page: https://gitnova.dev/en/r/allenai/olmocr ## Useful for - Convert scientific papers to Markdown for an LLM training dataset - Extract text from scans with tables and equations in reading order - Run olmOCR-Bench to compare OCR system quality ## Why it’s here - Star-counter measurements on 2026-10-07 (UTC), 11:04–12:47: 19720 → 19723 stars (+3). This is the change over that interval. - Estimated end-of-day forecast: about +9 stars, using observed gains and the previous day. - Over the last two days the pace is 3.4× that of the previous week and a half. - GitHub Trending Python today: #3, +22 stars. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 1,645 - Issues and pull requests: 476 - Watchers: 109 - Average over the last week: 8 per day - Usual pace: 9 per day - Stars in the last hour (measured): -1 - Latest release: v0.4.27 (2026-03-12) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-08 … 2026-10-07: 10, 6, 6, 5, 0, 0, 83, 53, 21, 6, 3, 1, 2, 12, 8, 7, 4, 5, 4, 12, 4, 5, 5, 5, 5, 3, 2, 10, 24, 3 ## Spotted in now - GitHub Trending Python today: #3, +22 stars ## Similar by description 1. **datalab-to/chandra** — 0.9 · Steady · Language models · Python · +0 stars measured 2026-10-07, 00:27–12:53 UTC, ≈ 1 by evening Chandra OCR 2 is an OCR model that converts images and PDFs into structured HTML, Markdown, or JSON while preserving layout, tables, math, and handwriting. Full card: https://gitnova.dev/en/r/datalab-to/chandra.md 2. **beatrizalmeidaf/papero-pdf-text-extractor** — 0.8 · Steady · Data & analytics · Python · +0 stars measured 2026-10-07, 00:25–12:54 UTC, ≈ 0 by evening Open-source API and Python library for extracting PDF structure: reading order, tables, formulas, figures and block positions. Runs on CPU without ML models, files are never stored. Full card: https://gitnova.dev/en/r/beatrizalmeidaf/papero-pdf-text-extractor.md 3. **microsoft/markitdown** — 3.1 · Steady · Language models · Python · +6 stars measured 2026-10-07, 11:06–12:42 UTC, ≈ 49 by evening A Python utility from Microsoft for converting PDFs, Office documents, images, audio, HTML, and other formats to Markdown, aimed at LLM and text-analysis pipelines. It preserves document structure (headings, lists, tables, links) and is… Full card: https://gitnova.dev/en/r/microsoft/markitdown.md 4. **Edwardxlai/easyread** — 1.8 · Cooling · Productivity & self-hosted · Python · +7 stars measured 2026-10-07, 00:22–12:47 UTC, ≈ 13 by evening A local app for reading English-language research papers: it imports PDFs, translates them page by page in the background, and shows the translation with formulas and tables, side-by-side original text, annotations, and AI Q&A. All data… Full card: https://gitnova.dev/en/r/Edwardxlai/easyread.md 5. **LaoFeng-mouse/flyingmouse-format** — 1.7 · Steady · Apps · JavaScript · +9 stars measured 2026-10-07, 00:25–12:48 UTC, ≈ 16 by evening Offline Electron file converter for Windows and macOS covering images, documents, spreadsheets, PDF, audio and video with OCR, batch mode and a CLI. Bundles FFmpeg, LibreOffice, Poppler and Tesseract, no network needed. Full card: https://gitnova.dev/en/r/LaoFeng-mouse/flyingmouse-format.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-07 12:58 UTC, updated every 30 minutes.