# huggingface/tokenizers > Hugging Face tokenization library written in Rust with Python and Node.js bindings: fast BPE, Unigram, WordPiece and WordLevel implementations for training and inference of NLP models. - Magnitude: 2.5 out of 10 — Steady - Stars: 11,071 total · +4 stars today, ≈ 10 by evening - Star trust: star growth looks organic - Category: Language models · Language: Rust · License: Apache-2.0 · Created: 2019-11-01 · Last push: 2026-09-21 - GitHub: https://github.com/huggingface/tokenizers · Homepage: https://huggingface.co/docs/tokenizers · Page: https://gitnova.dev/en/r/huggingface/tokenizers ## Useful for - Load the Llama-3.1-8B tokenizer via from_pretrained and encode text into ids - Train a custom BPE tokenizer on a text corpus for a domain-specific model - Embed fast tokenization inference into a Rust service without Python dependencies ## Why it’s here - 4 stars so far today, about 10 expected by the end of the day. - Before this, the repository barely got any stars — about 3 per day. - The spike has held for 2 days in a row — not a one-off blip. - GitHub Trending Rust today: #13, +18 stars. - Recent forks include notable developers: @albertvillanova (631 followers), @johnmai-dev (209 followers). ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 1,206 - Issues and pull requests: 2,412 - Watchers: 122 - Average over the last week: 6 per day - Usual pace: 3 per day - Stars in the last hour (measured): 3 - Latest release: v1.0.0-rc.2 (2026-09-21) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-08-24 … 2026-09-22: 3, 1, 2, 3, 0, 3, 3, 1, 3, 2, 7, 1, 1, 4, 5, 1, 1, 2, 4, 4, 2, 4, 1, 3, 7, 0, 3, 1, 17, 4 ## Spotted in now - GitHub Trending Rust today: #13, +18 stars ## Similar by description 1. **deepseek-ai/deepseek-recipe** — 1.2 · Cooling · Language models · Rust · +0 stars today, ≈ 1 by evening A collection of Rust libraries and Python bindings that convert API requests in different formats (Messages, Chat Completions, Responses) into the Conversation format, encode them into prompts for DeepSeek V4/V4.1 models, and convert… Full card: https://gitnova.dev/en/r/deepseek-ai/deepseek-recipe.md 2. **marin-community/marin** — 1.2 · Cooling · Language models · Python · +4 stars today, ≈ 8 by evening Open platform and community for research and development of foundation models: data curation, tokenization, pretraining, posttraining and evaluation of LLMs. Aimed at researchers and engineers training language and multimodal models. Full card: https://gitnova.dev/en/r/marin-community/marin.md 3. **huggingface/transformers** — 2.6 · Cooling · Language models · Python · +20 stars today, ≈ 44 by evening Model-definition framework for state-of-the-art pretrained ML models (text, vision, audio, multimodal) for inference and training, compatible with most training and inference engines. Full card: https://gitnova.dev/en/r/huggingface/transformers.md 4. **ggml-org/llama.cpp** — 3.3 · Steady · Language models · C++ · +38 stars today, ≈ 89 by evening LLM and VLM inference implemented in C/C++ with no dependencies, running on CPUs and GPUs across many hardware backends. Enables local model execution with minimal setup and quantization. Full card: https://gitnova.dev/en/r/ggml-org/llama.cpp.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-22 12:33 UTC, updated every 30 minutes.