# datalab-to/chandra > Chandra OCR 2 is an OCR model that converts images and PDFs into structured HTML, Markdown, or JSON while preserving layout, tables, math, and handwriting. - Magnitude: 2.2 out of 10 — Steady - Stars: 12,395 total · +8 stars today, ≈ 14 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: Apache-2.0 · Created: 2025-10-08 · Last push: 2026-06-26 - GitHub: https://github.com/datalab-to/chandra · Homepage: https://www.datalab.to · Page: https://gitnova.dev/en/r/datalab-to/chandra ## Useful for - Convert scanned PDFs to Markdown while preserving tables and layout - Extract data from forms and handwritten documents into JSON - Run local OCR via CLI or the Streamlit app ## Why it’s here - 8 stars so far today, about 14 expected by the end of the day. The usual pace is 6 per day, so that's 2.4× as much. - Over the last two days the pace is 3.3× that of the previous week and a half. - The spike has held for 2 days in a row — not a one-off blip. - GitHub Trending Python today: #5, +20 stars. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 1,250 - Issues and pull requests: 117 - Watchers: 89 - Average over the last week: 11 per day - Usual pace: 6 per day - Stars in the last hour (measured): 3 - Latest release: v0.2.0 (2026-03-18) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-04 … 2026-10-03: 4, 3, 4, 6, 7, 8, 5, 6, 4, 2, 5, 3, 7, 5, 15, 4, 11, 4, 5, 4, 9, 0, 7, 2, 7, 10, 2, 22, 23, 8 ## Spotted in now - GitHub Trending Python today: #5, +20 stars ## Similar by description 1. **PaddlePaddle/PaddleOCR** — 2.1 · Steady · Language models · Python · +19 stars today, ≈ 30 by evening OCR toolkit and document AI engine that turns PDFs and images into structured, LLM-ready data (JSON/Markdown), with support for 100+ languages, tables, formulas and seals. Full card: https://gitnova.dev/en/r/PaddlePaddle/PaddleOCR.md 2. **SeerRay-Lab/Xiaomi-OCR-0** — 0.7 · Steady · Language models · Python · +0 stars today, ≈ 1 by evening A unified 0.8B vision-language model for document parsing and OCR-centric understanding: outputs structured Markdown, extracts fields as JSON, and answers questions about document images. Full card: https://gitnova.dev/en/r/SeerRay-Lab/Xiaomi-OCR-0.md 3. **run-llama/liteparse** — 1.2 · Cooling · Data & analytics · Rust · +6 stars today, ≈ 10 by evening A local Rust document parser that extracts text with bounding boxes from PDF, DOCX, XLSX, PPTX and images, with OCR via Tesseract or HTTP servers. Runs without cloud or LLMs, with bindings for Python, Node.js, WASM and a CLI. Full card: https://gitnova.dev/en/r/run-llama/liteparse.md 4. **beatrizalmeidaf/papero-pdf-text-extractor** — 4.4 · Early signal · Data & analytics · Python · +3 stars today, ≈ 9 by evening Open-source API and Python library for extracting PDF structure: reading order, tables, formulas, figures and block positions. Runs on CPU without ML models, files are never stored. Full card: https://gitnova.dev/en/r/beatrizalmeidaf/papero-pdf-text-extractor.md 5. **chrisryugj/kordoc** — 1.2 · Cooling · Developer tools · TypeScript · +1 star today, ≈ 3 by evening Converts Korean government and office documents (HWP, HWPX, HWPML, PDF, XLS/XLSX, DOCX, images) to Markdown with structure-preserving tables, plus CLI and MCP server for form filling, document diffing, and HWPX generation. Full card: https://gitnova.dev/en/r/chrisryugj/kordoc.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-03 15:15 UTC, updated every 30 minutes.