GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

constructai/VibeThinker-3B-GGUF overview

constructai/VibeThinker 3B GGUF This is a quantized version of the original VibeThinker 3B , converted to the GGUF format for efficient CPU/GPU inference with …

ggufmathcodereasoninggpqainstruction-followingtext-generationbase_model:WeiboAI/VibeThinker-3Bbase_model:quantized:WeiboAI/VibeThinker-3Blicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~754.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

28 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
VibeThinker-3B-GGUF-F16.ggufGGUFF165.75 GBDownload
VibeThinker-3B-GGUF-Q2_K.ggufGGUFQ2_K1.19 GBDownload
VibeThinker-3B-GGUF-Q3_K_M.ggufGGUFQ3_K_M1.48 GBDownload
VibeThinker-3B-GGUF-Q3_K_S.ggufGGUFQ3_K_S1.35 GBDownload
VibeThinker-3B-GGUF-Q4_K_M.ggufGGUFQ4_K_M1.80 GBDownload
VibeThinker-3B-GGUF-Q4_K_S.ggufGGUFQ4_K_S1.71 GBDownload
VibeThinker-3B-GGUF-Q5_K_M.ggufGGUFQ5_K_M2.07 GBDownload
VibeThinker-3B-GGUF-Q5_K_S.ggufGGUFQ5_K_S2.02 GBDownload
VibeThinker-3B-GGUF-Q6_K.ggufGGUFQ6_K2.36 GBDownload
VibeThinker-3B-GGUF-Q8_0.ggufGGUFQ8_03.06 GBDownload
VibeThinker-3B-GGUF-UD-IQ1_M.ggufGGUFIQ1_M810.6 MBDownload
VibeThinker-3B-GGUF-UD-IQ1_S.ggufGGUFIQ1_S754.4 MBDownload
VibeThinker-3B-GGUF-UD-IQ2_M.ggufGGUFIQ2_M1.06 GBDownload
VibeThinker-3B-GGUF-UD-IQ2_XXS.ggufGGUFIQ2_XXS904.3 MBDownload
VibeThinker-3B-GGUF-UD-IQ3_S.ggufGGUFIQ3_S1.36 GBDownload
VibeThinker-3B-GGUF-UD-IQ3_XXS.ggufGGUFIQ3_XXS1.19 GBDownload
VibeThinker-3B-GGUF-UD-IQ4_NL.ggufGGUFIQ4_NL1.70 GBDownload
VibeThinker-3B-GGUF-UD-IQ4_XS.ggufGGUFIQ4_XS1.62 GBDownload
VibeThinker-3B-GGUF-UD-Q2_K_XL.ggufGGUFQ2_K_XL1.19 GBDownload
VibeThinker-3B-GGUF-UD-Q3_K_M.ggufGGUFQ3_K_M1.48 GBDownload
VibeThinker-3B-GGUF-UD-Q3_K_XL.ggufGGUFQ3_K_XL1.59 GBDownload
VibeThinker-3B-GGUF-UD-Q4_K_XL.ggufGGUFQ4_K_XL1.80 GBDownload
VibeThinker-3B-GGUF-UD-Q5_K_M.ggufGGUFQ5_K_M2.07 GBDownload
VibeThinker-3B-GGUF-UD-Q5_K_S.ggufGGUFQ5_K_S2.02 GBDownload
VibeThinker-3B-GGUF-UD-Q5_K_XL.ggufGGUFQ5_K_XL2.07 GBDownload
VibeThinker-3B-GGUF-UD-Q6_K.ggufGGUFQ6_K2.36 GBDownload
VibeThinker-3B-GGUF-UD-Q6_K_XL.ggufGGUFQ6_K_XL2.36 GBDownload
VibeThinker-3B-GGUF-UD-Q8_K_XL.ggufGGUFQ8_K_XL3.06 GBDownload

Model Details

Model IDconstructai/VibeThinker-3B-GGUF
Authorconstructai
Pipelinetext-generation
Licensemit
Base modelWeiboAI/VibeThinker-3B
Last modified2026-06-19T19:30:30.000Z

Model README

---

license: mit

base_model:

  • WeiboAI/VibeThinker-3B

pipeline_tag: text-generation

tags:

  • math
  • code
  • reasoning
  • gpqa
  • instruction-following
  • gguf

---

constructai/VibeThinker-3B-GGUF

This is a quantized version of the original VibeThinker-3B , converted to the GGUF format for efficient CPU/GPU inference with llama.cpp, Ollama, or any GGUF‑compatible runner.

---

Original Model

---

Available Quantizations

Choose the quantization that fits your needs:

| Quantization | File Size |

|--------------|-----------|

| UD-IQ1_S | 791 MB |

| UD-IQ1_M | 850 MB |

| UD-IQ2_XXS | 948 MB |

| Q2_K | 1.27 GB |

| UD-IQ2_M | 1.14 GB |

| UD-Q2_K_XL | 1.27 GB |

| UD-IQ3_XXS | 1.28 GB |

| Q3_K_S | 1.45 GB |

| UD-IQ3_S | 1.46 GB |

| Q3_K_M | 1.59 GB |

| UD-Q3_K_M | 1.59 GB |

| UD-Q3_K_XL | 1.71 GB |

| UD-IQ4_XS | 1.74 GB |

| Q4_K_S | 1.83 GB |

| UD-IQ4_NL | 1.83 GB |

| Q4_K_M | 1.93 GB |

| UD-Q4_K_XL | 1.93 GB |

| Q5_K_S | 2.17 GB |

| UD-Q5_K_S | 2.17 GB |

| Q5_K_M | 2.22 GB |

| UD-Q5_K_M | 2.22 GB |

| UD-Q5_K_XL | 2.22 GB |

| Q6_K | 2.54 GB |

| UD-Q6_K | 2.54 GB |

| UD-Q6_K_XL | 2.54 GB |

| Q8_0 | 3.29 GB |

| UD-Q8_K_XL | 3.29 GB |

| F16 | 6.18 GB |

For a 3B‑parameter model, even the larger files are quite manageable. Here’s what I recommend: F16 (6.18 GB) or Q8_0 (3.29 GB).

The other quants are also usable!

---

Usage

With ollama

ollama run hf.co/constructai/VibeThinker-3B-GGUF:F16

---

With llama.cpp

llama-server -hf constructai/VibeThinker-3B-GGUF:VibeThinker-3B-GGUF-F16.gguf

or

llama-cli -hf constructai/VibeThinker-3B-GGUF:VibeThinker-3B-GGUF-F16.gguf

---

With LM Studio

lms get constructai/VibeThinker-3B-GGUF@F16

---

Run constructai/VibeThinker-3B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models