General-Instinct/InstinctFlash
High-Performance Serving Runtime for Robotics Models
About the project
A high-performance serving runtime for robotics models (VLA, WAM, policies) with FP8 acceleration and optimized kernels on Jetson Thor and RTX 4090/5090.
Useful for
- Serve a fine-tuned robotics checkpoint for inference via `instinctflash serve`
- Expose robot action predictions over the openpi WebSocket protocol to existing clients
- Speed up pi0.5 or LingBot-VA inference with FP8 on Jetson Thor
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 10 stars so far today, about 12 expected by the end of the day. The usual pace is 0 per day, so that's 12× as much.
- Hacker News: “Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor” — 10 points, 3 h ago.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 21
- Today
- 10 · ≈ 12 by evening
- Forks
- 5
- Issues and pull requests
- 11
- Watchers
- 0
- Language
- C++
- License
- AGPL-3.0
- Latest release
- thor-2026-09-15 · September 15, 2026
- Created
- May 28, 2026
- Last push
- September 18, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Hacker News discussions
- Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor10 points, September 22, 2026
Spotted in
- September 22, 2026Spotted on Hacker News
Similar by description
-
0.0
joykraft/vla_pi0
Project for the Episode1 single-arm robot: teleoperation data collection via LeRobot, LoRA fine-tuning of the π₀ (pi0) VLA model on OpenPI/JAX, and deployment of the policy to a real robot over WebSocket.
-
1.5
NVIDIA/TensorRT-LLM
NVIDIA library for optimizing inference of large language models and visual generative models on GPUs, with a Python API, specialized kernels, and an efficient C++/Python runtime.
-
1.3
microsoft/onnxruntime
Cross-platform accelerator for ML inference and training. Runs models from PyTorch, TensorFlow, scikit-learn and others via the ONNX format with graph optimizations and hardware acceleration.
-
0.0
Binaire-0101/FRZi-inference
FRZi-inference is an inference engine for running machine learning models. It is designed to execute pre-trained models and obtain predictions.