cactus-compute/needle
Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.
About the project
A 2-bit, 8-29 MB foundation model for phones, wearables, robots and microcontrollers that handles tool calls, structured extraction and text embeddings on-device.
Useful for
- Route a voice command to the right app function on a phone
- Extract typed fields from invoices or bookings into a JSON record
- Fine-tune an 8-layer subnetwork on a product's own tools
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 30 stars so far today, about 100 expected by the end of the day.
- Over the last two days the pace is 2.4× that of the previous week and a half.
- GitHub Trending today: #14, +207 stars.
- GitHub Trending Python today: #3, +207 stars.
- Hacker News: “Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash” — 200 points, 1 day ago.
- About 9 forks a day — people are taking the code.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 11,414
- Stars in a day
- 100
- Forks
- 727
- Issues and pull requests
- 130
- Watchers
- 64
- Language
- Python
- License
- Apache-2.0
- Created
- February 24, 2026
- Last push
- September 18, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Hacker News discussions
- Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash200 points, September 18, 2026
Spotted in
- September 19, 2026GitHub Trending today: #14, +207 stars; GitHub Trending Python today: #3, +207 stars
Similar projects
-
7.7
TheoLeeCJ/SemIf
An open reproduction of the Jev-style semantic decision interface: reads typed option probabilities directly from a 4B model's logits without generating text. Runs on a single RTX 3090.
-
7.4
TianyuCodings/NanoJev
A nano replica of Jev built on Qwen3-0.6B that outputs probability distributions over candidates in a single forward pass with no token decoding, plus a training pipeline and game demos (maze, Snake).
-
6.4
Continuum-AI-Corp/OrcaBonsai-27B-Uncensored
A tool for runtime refusal ablation in the compressed Ternary Bonsai 2 27B LLM: it projects the residual stream to remove the refusal direction without modifying or re-quantizing weights. Runs on Apple Silicon via MLX.
-
6.1
jaredpalmer/kev
kev is a LoRA adapter with a small readout head on top of Qwen2.5-0.5B that answers many typed questions about a document in a single forward pass, returning calibrated probabilities instead of text.
-
6.0
githubnext/localjev
A local TypeScript/Bun HTTP service implementing a Jev-compatible POST /v1/systemone API on top of an OpenAI-compatible Chat Completions endpoint (DiffusionGemma via oMLX). It acts as a bridge for typed decisions…
-
6.0
higgsfield-ai/higgsfield
Fault-tolerant GPU orchestration and ML framework for distributed training of models with billions to trillions of parameters, including LLMs. Manages node access, experiment queueing, and GitHub Actions integration.