GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ynanxiu/qwen25-1.5b-coffee-lora-gguf overview

Qwen2.5 1.5B Coffee GGUF Q4 K M quantized GGUF for llama.cpp. 940 MB. Usage bash llama cli m qwen25 15b coffee v2 q4km.gguf p "<|im start| user Hello<|im end| …

ggufqwen2.5coffeellama.cppq4_k_mtext-generationzhlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~940.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
qwen25_15b_coffee_v2_q4km.ggufGGUFQ4KM940.4 MBDownload

Model Details

Model IDynanxiu/qwen25-1.5b-coffee-lora-gguf
Authorynanxiu
Pipelinetext-generation
Licenseapache-2.0
Base model
Last modified2026-06-18T14:01:48.000Z

Model README

---

language: zh

license: apache-2.0

tags:

  • qwen2.5
  • gguf
  • coffee
  • llama.cpp
  • q4_k_m

pipeline_tag: text-generation

---

Qwen2.5-1.5B Coffee GGUF

Q4_K_M quantized GGUF for llama.cpp. 940 MB.

Usage

llama-cli -m qwen25_15b_coffee_v2_q4km.gguf -p "<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
"

Performance

  • Size: 940 MB
  • Quality: 34/40 (tool 19/20, knowledge 15/20)
  • 2h4g memory: ~2.5 GB
  • RTX 4060 GPU: ~1s per query

Run ynanxiu/qwen25-1.5b-coffee-lora-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models