Tencent/WeMM-Embedding
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
About the project
A family of universal multimodal embedding models from Tencent's WeChat Vision team (2B, 4B, 9B) producing unified vectors for text, images, videos, visual documents, and interleaved multimodal inputs. Supports Matryoshka dimensions and serving via vLLM/SGLang.
Useful for
- Build multimodal retrieval over a mixed corpus of images, videos, and documents with one index
- Truncate embeddings to 256 dimensions via Matryoshka to save memory with minimal quality loss
- Serve embeddings for text and video online using vLLM or SGLang
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 52 stars today.
- The repository is 19 days old and already has 1,562 stars.
- Top new repositories this month: #38.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 1,562
- Stars in a day
- 53
- Forks
- 109
- Issues and pull requests
- 8
- Watchers
- 6
- Language
- Python
- Created
- August 25, 2026
- Last push
- September 3, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 13, 2026Top new repositories this month: #37
Similar projects
-
6.1
JustVugg/colibri
A pure-C, zero-dependency inference engine for running large MoE models (744B–2.8T parameters) on consumer hardware by treating VRAM, RAM, and storage as a single multitier hierarchy and streaming experts from disk.
-
5.7
asgeirtj/system_prompts_leaks
A collection of extracted system prompts from Anthropic, OpenAI, Google, xAI and others — the hidden instructions chatbots receive before a user's first message.
-
5.1
kennethwolters/litelm
A lightweight litellm alternative: routes LLM calls across 19 providers and translates message formats in ~2,900 lines with two dependencies, without proxy, caching, or cost tracking.
-
5.0
MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
A local EXL3 2.9 bpw checkpoint of DeepSeek-V4.1-Flash (196 GiB, 39 shards) served via an OpenAI-compatible vLLM API on a 2x NVIDIA GB10 (DGX Spark) kit with tensor-parallel 2. Includes DSpark speculative decoding and…
-
4.8
unclecode/crawl4ai
Open-source web crawler and scraper that turns pages into clean, LLM-ready Markdown for RAG, agents and data pipelines. Runs via Python API, CLI and Docker with no API keys.
-
4.5
Edge0-AI/Edge0
An open-source streaming MoE inference framework: expert weights are offloaded from SSD on demand while a trained prerouter predicts routing ahead of time. Runs on Apple Silicon via MLX and ships with two ready-to-run…