Model Intelligence Sheet
ynanxiu/qwen25-1.5b-coffee-lora-gguf overview
Qwen2.5 1.5B Coffee GGUF Q4 K M quantized GGUF for llama.cpp. 940 MB. Usage bash llama cli m qwen25 15b coffee v2 q4km.gguf p "<|im start| user Hello<|im end| …
Runs locally from ~940.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
1 GGUF files detected
Direct downloads for local inference
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| qwen25_15b_coffee_v2_q4km.gguf | GGUF | Q4KM | 940.4 MB | Download |
Model Details
| Model ID | ynanxiu/qwen25-1.5b-coffee-lora-gguf |
|---|---|
| Author | ynanxiu |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | — |
| Last modified | 2026-06-18T14:01:48.000Z |
Model README
---
language: zh
license: apache-2.0
tags:
- qwen2.5
- gguf
- coffee
- llama.cpp
- q4_k_m
pipeline_tag: text-generation
---
Qwen2.5-1.5B Coffee GGUF
Q4_K_M quantized GGUF for llama.cpp. 940 MB.
Usage
llama-cli -m qwen25_15b_coffee_v2_q4km.gguf -p "<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
"
Performance
- Size: 940 MB
- Quality: 34/40 (tool 19/20, knowledge 15/20)
- 2h4g memory: ~2.5 GB
- RTX 4060 GPU: ~1s per query
Run ynanxiu/qwen25-1.5b-coffee-lora-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models