constructai/VibeThinker-3B-GGUF overview
constructai/VibeThinker 3B GGUF This is a quantized version of the original VibeThinker 3B , converted to the GGUF format for efficient CPU/GPU inference with …
Runs locally from ~754.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| VibeThinker-3B-GGUF-F16.gguf | GGUF | F16 | 5.75 GB | Download |
| VibeThinker-3B-GGUF-Q2_K.gguf | GGUF | Q2_K | 1.19 GB | Download |
| VibeThinker-3B-GGUF-Q3_K_M.gguf | GGUF | Q3_K_M | 1.48 GB | Download |
| VibeThinker-3B-GGUF-Q3_K_S.gguf | GGUF | Q3_K_S | 1.35 GB | Download |
| VibeThinker-3B-GGUF-Q4_K_M.gguf | GGUF | Q4_K_M | 1.80 GB | Download |
| VibeThinker-3B-GGUF-Q4_K_S.gguf | GGUF | Q4_K_S | 1.71 GB | Download |
| VibeThinker-3B-GGUF-Q5_K_M.gguf | GGUF | Q5_K_M | 2.07 GB | Download |
| VibeThinker-3B-GGUF-Q5_K_S.gguf | GGUF | Q5_K_S | 2.02 GB | Download |
| VibeThinker-3B-GGUF-Q6_K.gguf | GGUF | Q6_K | 2.36 GB | Download |
| VibeThinker-3B-GGUF-Q8_0.gguf | GGUF | Q8_0 | 3.06 GB | Download |
| VibeThinker-3B-GGUF-UD-IQ1_M.gguf | GGUF | IQ1_M | 810.6 MB | Download |
| VibeThinker-3B-GGUF-UD-IQ1_S.gguf | GGUF | IQ1_S | 754.4 MB | Download |
| VibeThinker-3B-GGUF-UD-IQ2_M.gguf | GGUF | IQ2_M | 1.06 GB | Download |
| VibeThinker-3B-GGUF-UD-IQ2_XXS.gguf | GGUF | IQ2_XXS | 904.3 MB | Download |
| VibeThinker-3B-GGUF-UD-IQ3_S.gguf | GGUF | IQ3_S | 1.36 GB | Download |
| VibeThinker-3B-GGUF-UD-IQ3_XXS.gguf | GGUF | IQ3_XXS | 1.19 GB | Download |
| VibeThinker-3B-GGUF-UD-IQ4_NL.gguf | GGUF | IQ4_NL | 1.70 GB | Download |
| VibeThinker-3B-GGUF-UD-IQ4_XS.gguf | GGUF | IQ4_XS | 1.62 GB | Download |
| VibeThinker-3B-GGUF-UD-Q2_K_XL.gguf | GGUF | Q2_K_XL | 1.19 GB | Download |
| VibeThinker-3B-GGUF-UD-Q3_K_M.gguf | GGUF | Q3_K_M | 1.48 GB | Download |
| VibeThinker-3B-GGUF-UD-Q3_K_XL.gguf | GGUF | Q3_K_XL | 1.59 GB | Download |
| VibeThinker-3B-GGUF-UD-Q4_K_XL.gguf | GGUF | Q4_K_XL | 1.80 GB | Download |
| VibeThinker-3B-GGUF-UD-Q5_K_M.gguf | GGUF | Q5_K_M | 2.07 GB | Download |
| VibeThinker-3B-GGUF-UD-Q5_K_S.gguf | GGUF | Q5_K_S | 2.02 GB | Download |
| VibeThinker-3B-GGUF-UD-Q5_K_XL.gguf | GGUF | Q5_K_XL | 2.07 GB | Download |
| VibeThinker-3B-GGUF-UD-Q6_K.gguf | GGUF | Q6_K | 2.36 GB | Download |
| VibeThinker-3B-GGUF-UD-Q6_K_XL.gguf | GGUF | Q6_K_XL | 2.36 GB | Download |
| VibeThinker-3B-GGUF-UD-Q8_K_XL.gguf | GGUF | Q8_K_XL | 3.06 GB | Download |
Model Details
| Model ID | constructai/VibeThinker-3B-GGUF |
|---|---|
| Author | constructai |
| Pipeline | text-generation |
| License | mit |
| Base model | WeiboAI/VibeThinker-3B |
| Last modified | 2026-06-19T19:30:30.000Z |
Model README
---
license: mit
base_model:
- WeiboAI/VibeThinker-3B
pipeline_tag: text-generation
tags:
- math
- code
- reasoning
- gpqa
- instruction-following
- gguf
---
constructai/VibeThinker-3B-GGUF
This is a quantized version of the original VibeThinker-3B , converted to the GGUF format for efficient CPU/GPU inference with llama.cpp, Ollama, or any GGUF‑compatible runner.
---
Original Model
- Author(s): WeiboAI
- Source: VibeThinker-3B
- Original License: MIT
---
Available Quantizations
Choose the quantization that fits your needs:
| Quantization | File Size |
|--------------|-----------|
| UD-IQ1_S | 791 MB |
| UD-IQ1_M | 850 MB |
| UD-IQ2_XXS | 948 MB |
| Q2_K | 1.27 GB |
| UD-IQ2_M | 1.14 GB |
| UD-Q2_K_XL | 1.27 GB |
| UD-IQ3_XXS | 1.28 GB |
| Q3_K_S | 1.45 GB |
| UD-IQ3_S | 1.46 GB |
| Q3_K_M | 1.59 GB |
| UD-Q3_K_M | 1.59 GB |
| UD-Q3_K_XL | 1.71 GB |
| UD-IQ4_XS | 1.74 GB |
| Q4_K_S | 1.83 GB |
| UD-IQ4_NL | 1.83 GB |
| Q4_K_M | 1.93 GB |
| UD-Q4_K_XL | 1.93 GB |
| Q5_K_S | 2.17 GB |
| UD-Q5_K_S | 2.17 GB |
| Q5_K_M | 2.22 GB |
| UD-Q5_K_M | 2.22 GB |
| UD-Q5_K_XL | 2.22 GB |
| Q6_K | 2.54 GB |
| UD-Q6_K | 2.54 GB |
| UD-Q6_K_XL | 2.54 GB |
| Q8_0 | 3.29 GB |
| UD-Q8_K_XL | 3.29 GB |
| F16 | 6.18 GB |
For a 3B‑parameter model, even the larger files are quite manageable. Here’s what I recommend: F16 (6.18 GB) or Q8_0 (3.29 GB).
The other quants are also usable!
---
Usage
With ollama
ollama run hf.co/constructai/VibeThinker-3B-GGUF:F16
---
With llama.cpp
llama-server -hf constructai/VibeThinker-3B-GGUF:VibeThinker-3B-GGUF-F16.gguf
or
llama-cli -hf constructai/VibeThinker-3B-GGUF:VibeThinker-3B-GGUF-F16.gguf
---
With LM Studio
lms get constructai/VibeThinker-3B-GGUF@F16
---
Run constructai/VibeThinker-3B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models