Seismograph

What’s gaining stars on GitHub right now

huggingface/tokenizers

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

Language modelsRust#nlp#natural-language-processing#natural-language-understanding#language-model#transformers
2.5 Steady Magnitude out of 10 — how fast interest is growing, not a quality score. Star growth looks organic. Data as of September 22, 2026, 12:33 UTC.
Open on Seismograph Open on GitHub

About the project

Hugging Face tokenization library written in Rust with Python and Node.js bindings: fast BPE, Unigram, WordPiece and WordLevel implementations for training and inference of NLP models.

Useful for

README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.

Why it’s trending

Stars per day

01020June 25, 2026September 22, 2026

Bars are daily stars, the line is the usual pace. Red marks spike days.

Numbers

Total stars
11,071
Today
4 · ≈ 10 by evening
Forks
1,206
Issues and pull requests
2,412
Watchers
122
Language
Rust
License
Apache-2.0
Latest release
v1.0.0-rc.2 · September 21, 2026
Created
November 1, 2019
Last push
September 21, 2026

Star trust

Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.

These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.

Spotted in

Share

README badge

Paste it into your README — the badge shows the current magnitude and links to this page.

Seismograph: 2.5
[![Seismograph](https://gitnova.dev/badge/huggingface/tokenizers.svg?lang=en)](https://gitnova.dev/en/r/huggingface/tokenizers)

Similar by description

  1. 1.2
    deepseek-ai/deepseek-recipe

    A collection of Rust libraries and Python bindings that convert API requests in different formats (Messages, Chat Completions, Responses) into the Conversation format, encode them into prompts for DeepSeek V4/V4.1…

    CoolingLanguage modelsRust+0 stars today, ≈ 1 by evening

  2. 1.2
    marin-community/marin

    Open platform and community for research and development of foundation models: data curation, tokenization, pretraining, posttraining and evaluation of LLMs. Aimed at researchers and engineers training language and…

    CoolingLanguage modelsPython+4 stars today, ≈ 8 by evening

  3. 2.6
    huggingface/transformers

    Model-definition framework for state-of-the-art pretrained ML models (text, vision, audio, multimodal) for inference and training, compatible with most training and inference engines.

    CoolingLanguage modelsPython+20 stars today, ≈ 44 by evening

  4. 3.3
    ggml-org/llama.cpp

    LLM and VLM inference implemented in C/C++ with no dependencies, running on CPUs and GPUs across many hardware backends. Enables local model execution with minimal setup and quantization.

    SteadyLanguage modelsC+++38 stars today, ≈ 89 by evening