Yigit-Karaman/Cozum-4B-GGUF overview
Cozum 4B GGUF This repository contains a comprehensive collection of quantized GGUF weights for Yigit Karaman/Cozum 4B https://huggingface.co/Yigit Karaman/Coz…
Runs locally from ~644.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Cozum-4B_BF16-mmproj.gguf | GGUF | BF16 | 644.3 MB | Download |
| Cozum-4B_BF16.gguf | GGUF | BF16 | 7.85 GB | Download |
| Cozum-4B_Q2_K_L.gguf | GGUF | Q2_K_L | 1.93 GB | Download |
| Cozum-4B_Q3_K_M.gguf | GGUF | Q3_K_M | 2.11 GB | Download |
| Cozum-4B_Q4_K_M.gguf | GGUF | Q4_K_M | 2.52 GB | Download |
| Cozum-4B_Q5_K_M.gguf | GGUF | Q5_K_M | 2.86 GB | Download |
| Cozum-4B_Q6_K.gguf | GGUF | Q6_K | 3.23 GB | Download |
| Cozum-4B_Q8_0.gguf | GGUF | Q8_0 | 4.17 GB | Download |
Model Details
| Model ID | Yigit-Karaman/Cozum-4B-GGUF |
|---|---|
| Author | Yigit-Karaman |
| Pipeline | text-generation |
| License | — |
| Base model | Yigit-Karaman/Cozum-4B |
| Last modified | 2026-07-22T14:42:21.000Z |
Model README
---
base_model: Yigit-Karaman/Cozum-4B
tags:
- gguf
- quantization
- text-generation
- qwen
---
Cozum-4B-GGUF
This repository contains a comprehensive collection of quantized GGUF weights for Yigit-Karaman/Cozum-4B, optimized for efficient local execution and offline inference across various hardware setups.
---
Available Files & Quantizations
We provide a "ladder" of quantization options so you can choose the best fit for your machine's hardware capabilities. Compatibility notes are tailored for an 8GB VRAM profile (e.g., RTX 4060 Laptop):
| File Name | Size | Recommended Use |
| :--- | :--- | :--- |
| Cozum-4B_BF16.gguf | 8.42 GB | Unquantized 16-bit: Maximum precision, but exceeds standard 8GB VRAM limits. Will require system RAM offloading and run slower. |
| Cozum-4B_Q8_0.gguf | 4.48 GB | Maximum Quality (8-bit): Near-unquantized performance. Ideal for dedicated GPUs with 8GB+ VRAM, fitting perfectly with room for context. |
| Cozum-4B_Q6_K.gguf | 3.46 GB | High Quality (6-bit): Very minimal degradation. Leaves excellent VRAM headroom for large system prompts or extended conversations. |
| Cozum-4B_Q5_K_M.gguf | 3.07 GB | The Sweet Spot (5-bit): Excellent balance of quality and size. Highly recommended for general use and coding. |
| Cozum-4B_Q4_K_M.gguf | 2.71 GB | Mainstream Standard (4-bit): The best balance of speed and acceptable quality for mid-range hardware. |
| Cozum-4B_Q3_K_M.gguf | 2.26 GB | Low-End Hardware (3-bit): Maximum compression for severe memory constraints. |
| Cozum-4B_Q2_K_L.gguf | 2.07 GB | Extreme Compression (2-bit): Only use if absolutely necessary; expects noticeable degradation in complex logic and generation quality. |
| Cozum-4B_BF16-mmproj.gguf | 676 MB | Multimodal Projection: Required file if utilizing the model's vision/image processing capabilities. |
---
Model Overview
- Base Model: Yigit-Karaman/Cozum-4B
- Architecture: Qwen (4B Parameters)
- Format: GGUF (Various Quantization Levels)
- Primary Use Case: Local text generation, code assistance, and lightweight AI deployment.
---
Prompt Format (ChatML)
This model natively uses the ChatML format. Ensure your local runner's prompt template matches the structure below to prevent repetitive or unexpected generation output:
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistantRun Yigit-Karaman/Cozum-4B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models