PocketWeights/Qwen2.5-Coder-7B-Instruct-GGUF-6GB overview
⚡ PocketWeights: Qwen 2.5 Coder 7B Instruct 6GB VRAM Edition Heavy models, made light. We specialize in high quality GGUF quantizations optimized for edge devi…
Runs locally from ~4.36 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen2.5-Coder-7B-Instruct-GGUF-6GB-Q4_K_M.gguf | GGUF | Q4_K_M | 4.36 GB | Download |
Model Details
| Model ID | PocketWeights/Qwen2.5-Coder-7B-Instruct-GGUF-6GB |
|---|---|
| Author | PocketWeights |
| Pipeline | — |
| License | apache-2.0 |
| Base model | Qwen/Qwen2.5-Coder-7B-Instruct |
| Last modified | 2026-08-03T15:00:08.000Z |
Model README
---
base_model: Qwen/Qwen2.5-Coder-7B-Instruct
library_name: gguf
license: apache-2.0
tags:
- gguf
- qwen2.5
- code
- quantized
---
⚡ PocketWeights: Qwen 2.5 Coder 7B Instruct (6GB VRAM Edition)
Heavy models, made light. We specialize in high-quality GGUF quantizations optimized for edge devices and gaming laptops.
What is this model?
This is a heavily optimized GGUF of Qwen2.5-Coder-7B-Instruct. It preserves 99% of the coding logic but shrinks the massive 15GB model down to fit easily onto a laptop, making it the perfect offline coding assistant.
📦 Pick Your Hardware
We bypass the confusing wall of files to provide specific sizes optimized for your hardware:
Q4_K_M(Fits 6GB VRAM): The standard. Perfect for RTX 3060/4050.
🚀 Quick Start Guide
Option 1: LM Studio (Easiest)
- Download & open LM Studio.
- In the search bar, type:
PocketWeights/Qwen2.5-Coder-7B-Instruct-GGUF-6GB. - Select your hardware size, click Download, and hit Play!
Option 2: Ollama
If you have Ollama installed, simply run this command in your terminal:
ollama run hf.co/PocketWeights/Qwen2.5-Coder-7B-Instruct-GGUF-6GB:Q4_K_MRun PocketWeights/Qwen2.5-Coder-7B-Instruct-GGUF-6GB with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models