GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Yigit-Karaman/Cozum-4B-GGUF overview

Cozum 4B GGUF This repository contains a comprehensive collection of quantized GGUF weights for Yigit Karaman/Cozum 4B https://huggingface.co/Yigit Karaman/Coz…

ggufquantizationtext-generationqwenbase_model:Yigit-Karaman/Cozum-4Bbase_model:quantized:Yigit-Karaman/Cozum-4Bendpoints_compatibleregion:usconversational

Runs locally from ~644.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
46
Likes
1
Pipeline
text-generation

Repository Files & Downloads

8 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Cozum-4B_BF16-mmproj.ggufGGUFBF16644.3 MBDownload
Cozum-4B_BF16.ggufGGUFBF167.85 GBDownload
Cozum-4B_Q2_K_L.ggufGGUFQ2_K_L1.93 GBDownload
Cozum-4B_Q3_K_M.ggufGGUFQ3_K_M2.11 GBDownload
Cozum-4B_Q4_K_M.ggufGGUFQ4_K_M2.52 GBDownload
Cozum-4B_Q5_K_M.ggufGGUFQ5_K_M2.86 GBDownload
Cozum-4B_Q6_K.ggufGGUFQ6_K3.23 GBDownload
Cozum-4B_Q8_0.ggufGGUFQ8_04.17 GBDownload

Model Details

Model IDYigit-Karaman/Cozum-4B-GGUF
AuthorYigit-Karaman
Pipelinetext-generation
License
Base modelYigit-Karaman/Cozum-4B
Last modified2026-07-22T14:42:21.000Z

Model README

---

base_model: Yigit-Karaman/Cozum-4B

tags:

  • gguf
  • quantization
  • text-generation
  • qwen

---

Cozum-4B-GGUF

This repository contains a comprehensive collection of quantized GGUF weights for Yigit-Karaman/Cozum-4B, optimized for efficient local execution and offline inference across various hardware setups.

---

Available Files & Quantizations

We provide a "ladder" of quantization options so you can choose the best fit for your machine's hardware capabilities. Compatibility notes are tailored for an 8GB VRAM profile (e.g., RTX 4060 Laptop):

| File Name | Size | Recommended Use |

| :--- | :--- | :--- |

| Cozum-4B_BF16.gguf | 8.42 GB | Unquantized 16-bit: Maximum precision, but exceeds standard 8GB VRAM limits. Will require system RAM offloading and run slower. |

| Cozum-4B_Q8_0.gguf | 4.48 GB | Maximum Quality (8-bit): Near-unquantized performance. Ideal for dedicated GPUs with 8GB+ VRAM, fitting perfectly with room for context. |

| Cozum-4B_Q6_K.gguf | 3.46 GB | High Quality (6-bit): Very minimal degradation. Leaves excellent VRAM headroom for large system prompts or extended conversations. |

| Cozum-4B_Q5_K_M.gguf | 3.07 GB | The Sweet Spot (5-bit): Excellent balance of quality and size. Highly recommended for general use and coding. |

| Cozum-4B_Q4_K_M.gguf | 2.71 GB | Mainstream Standard (4-bit): The best balance of speed and acceptable quality for mid-range hardware. |

| Cozum-4B_Q3_K_M.gguf | 2.26 GB | Low-End Hardware (3-bit): Maximum compression for severe memory constraints. |

| Cozum-4B_Q2_K_L.gguf | 2.07 GB | Extreme Compression (2-bit): Only use if absolutely necessary; expects noticeable degradation in complex logic and generation quality. |

| Cozum-4B_BF16-mmproj.gguf | 676 MB | Multimodal Projection: Required file if utilizing the model's vision/image processing capabilities. |

---

Model Overview

  • Base Model: Yigit-Karaman/Cozum-4B
  • Architecture: Qwen (4B Parameters)
  • Format: GGUF (Various Quantization Levels)
  • Primary Use Case: Local text generation, code assistance, and lightweight AI deployment.

---

Prompt Format (ChatML)

This model natively uses the ChatML format. Ensure your local runner's prompt template matches the structure below to prevent repetitive or unexpected generation output:

<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant

Run Yigit-Karaman/Cozum-4B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models