Seismograph

What’s gaining stars on GitHub right now

ashhart/TensorFold

Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint

Language modelsPython#ai#ai-tools#llm#llm-inference#llm-tools
5.6 Early signal Magnitude out of 10 — how fast interest is growing, not a quality score. Star growth looks organic. Data as of September 28, 2026, 12:10 UTC.
Open on Seismograph Open on GitHub

About the project

LLM inference server for Apple Silicon (MLX) and NVIDIA GPUs with an OpenAI-compatible API and exact speculative decoding. Supports Nemotron, Qwen3.8, GLM-5.3 and Gemma 4 families with per-family kernels and draft models.

Useful for

README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.

Why it’s trending

Stars per day

0150300July 1, 2026September 28, 2026

Bars are daily stars, the line is the usual pace. Red marks spike days.

Numbers

Total stars
534
Today
10 · ≈ 65 by evening
Forks
52
Issues and pull requests
66
Watchers
9
Language
Python
License
MIT
Latest release
v0.3.6.1 · September 28, 2026
Created
June 19, 2026
Last push
September 28, 2026

Star trust

Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.

These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.

Spotted in

Share

README badge

Paste it into your README — the badge shows the current magnitude and links to this page.

Seismograph: 5.6
[![Seismograph](https://gitnova.dev/badge/ashhart/TensorFold.svg?lang=en)](https://gitnova.dev/en/r/ashhart/TensorFold)

Similar by description

  1. 0.5
    MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks

    A local EXL3 2.9 bpw checkpoint of DeepSeek-V4.1-Flash (196 GiB, 39 shards) served via an OpenAI-compatible vLLM API on a 2x NVIDIA GB10 (DGX Spark) kit with tensor-parallel 2. Includes DSpark speculative decoding and…

    SteadyLanguage modelsPython+0 stars today, ≈ 0 by evening

  2. 1.6
    NVIDIA/TensorRT-LLM

    NVIDIA library for optimizing inference of large language models and visual generative models on GPUs, with a Python API, specialized kernels, and an efficient C++/Python runtime.

    SteadyLanguage modelsPython+5 stars today, ≈ 9 by evening

  3. 4.3
    NVIDIA/Model-Optimizer

    NVIDIA library for compressing and accelerating models: quantization, pruning, distillation, NAS and speculative decoding with export to TensorRT-LLM, vLLM, SGLang. For ML engineers preparing models for deployment.

    PeakingLanguage modelsPython+15 stars today, ≈ 58 by evening

  4. 2.5
    incoai/splash

    A local LLM inference engine for Apple silicon, specialized per model: serves OpenAI- and Anthropic-compatible APIs to coding agents on a single Mac.

    CoolingLanguage modelsC+++10 stars today, ≈ 28 by evening