Weidows/WeMM-Embedding-2B-GGUF overview
WeMM Embedding 2B — Quantization Quality STS B Evaluation set: STS B test 1,379 sentence pairs, human similarity 0 5 . Baseline: BF16 GGUF run in llama.cpp sam…
Runs locally from ~640.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| WeMM-Embedding-2B-BF16.gguf | GGUF | BF16 | 4.46 GB | Download |
| WeMM-Embedding-2B-IQ3_M.gguf | GGUF | IQ3_M | 1.19 GB | Download |
| WeMM-Embedding-2B-IQ4_XS.gguf | GGUF | IQ4_XS | 1.37 GB | Download |
| WeMM-Embedding-2B-Q4_K_M.gguf | GGUF | Q4_K_M | 1.45 GB | Download |
| WeMM-Embedding-2B-Q5_K_M.gguf | GGUF | Q5_K_M | 1.64 GB | Download |
| WeMM-Embedding-2B-Q6_K.gguf | GGUF | Q6_K | 1.84 GB | Download |
| WeMM-Embedding-2B-Q8_0.gguf | GGUF | Q8_0 | 2.38 GB | Download |
| mmproj-WeMM-Embedding-2B-BF16.gguf | GGUF | BF16 | 640.3 MB | Download |
Model Details
| Model ID | Weidows/WeMM-Embedding-2B-GGUF |
|---|---|
| Author | Weidows |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | tencent/WeMM-Embedding-2B |
| Last modified | 2026-08-30T15:44:10.000Z |
Model README
---
base_model: tencent/WeMM-Embedding-2B
library_name: gguf
license: apache-2.0
tags:
- embedding
- multimodal
- gguf
- qwen3_5
- image-text-to-text
- video-text-to-text
---
WeMM-Embedding-2B — Quantization Quality (STS-B)
Evaluation set: STS-B test (1,379 sentence pairs, human similarity 0-5). Baseline: BF16 GGUF run in llama.cpp (same engine as all quants), so the measured difference reflects quantization error only.
| Model | Bits/Weight | Size (MB) | STS-B Spearman ρ | Δρ vs BF16 | Emb Cosine vs BF16 | Pair Cosine Pearson vs BF16 |
|---|---|---|---|---|---|---|
| BF16 | 16.00 | 4790.8 | 0.8360 | — | — | — |
| Q8_0 | 8.00 | 2551.3 | 0.8357 | +0.03% | 0.9997 | 1.0000 |
| Q6_K | 5.80 | 1972.8 | 0.8358 | +0.03% | 0.9987 | 0.9999 |
| Q5_K_M | 5.17 | 1760.0 | 0.8361 | -0.01% | 0.9956 | 0.9995 |
| Q4_K_M | 4.85 | 1559.8 | 0.8310 | +0.60% | 0.9854 | 0.9984 |
| IQ4_XS | 4.33 | 1471.4 | 0.8360 | +0.00% | 0.9855 | 0.9985 |
| IQ3_M | 3.76 | 1277.3 | 0.8286 | +0.89% | 0.9263 | 0.9895 |
Metrics
- STS-B Spearman ρ: rank correlation between model cosine similarities and human similarity scores. Higher is better.
- Δρ vs BF16: relative drop of ρ against the BF16 baseline. Negative means the quant scored slightly above baseline (within noise).
- Emb Cosine vs BF16: mean cosine similarity between each sentence's embedding and its BF16 counterpart (space fidelity). 1.0 = identical.
- Pair Cosine Pearson vs BF16: Pearson correlation of per-pair cosine similarities vs BF16 (ranking fidelity). 1.0 = identical ordering.
Conclusion
- Q8_0, Q6_K, Q5_K_M and IQ4_XS show negligible quality loss (|Δρ| < 0.05%, Emb Cosine > 0.985) and are safe drop-in replacements.
- Q4_K_M (4.85 bpw) shows a small but visible drop (Δρ ≈ +0.60%, Emb Cosine 0.985) — notably worse than the equally-sized IQ4_XS, so prefer IQ4_XS or Q5_K_M over Q4_K_M when size is comparable.
- IQ3_M (3.76 bpw) is the only variant with a clearly measurable drop (Δρ ≈ +0.89%, Emb Cosine 0.93); use only when storage is critical.
Usage (llama.cpp GGUF)
All files here are GGUF and run with llama.cpp. Replace the model file with the quant you downloaded. Use -ngl 999 to offload layers to GPU (omit or -ngl 0 for CPU-only).
Text embedding — command line
llama-embedding \
-m WeMM-Embedding-2B-Q5_K_M.gguf \
-p "Represent the meaning of this sentence." \
--pooling last
Text embedding — HTTP server
llama-server \
-m WeMM-Embedding-2B-Q5_K_M.gguf \
--embedding \
-ngl 999 --host 0.0.0.0 --port 8080
Then request embeddings via the OpenAI-compatible endpoint:
curl http://localhost:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input": "Represent the meaning of this sentence.", "model": "WeMM-Embedding-2B-Q5_K_M"}'
Multimodal (image / video) — HTTP server
The visual projector (mmproj-WeMM-Embedding-2B-BF16.gguf) is required for image and video inputs:
llama-server \
-m WeMM-Embedding-2B-Q5_K_M.gguf \
--mmproj mmproj-WeMM-Embedding-2B-BF16.gguf \
--embedding \
-ngl 999 --host 0.0.0.0 --port 8080
Send image/video inside the chat content the same way as the base model (interleave image/video before text).
Notes
- Output is a 2048-dim L2-normalized vector; matryoshka truncation (e.g.
--embd-normalize+ slicing) follows the base model'smatryoshka_dimensions[64, 128, 256, 512, 1024, 2048]. - Q8_0 / Q4_K_M / BF16 are mirrored from
DreamBlooms/WeMM-Embedding-2B-GGUF; Q6_K / Q5_K_M / IQ4_XS / IQ3_M were produced for this repo withllama-quantizefrom the same BF16 master.
Run Weidows/WeMM-Embedding-2B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models