GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Ma7ee7/Qwen3-1.7B-Depth-Aggressive-Q4_K_M-GGUF overview

Qwen3 1.7B Depth Aggressive — Q4 K M GGUF A structured depth pruned and recovery trained variant of Qwen/Qwen3 1.7B https://huggingface.co/Qwen/Qwen3 1.7B , co…

transformersggufqwen3llama-cpppruningstructured-pruningq4_k_mtext-generationbase_model:Qwen/Qwen3-1.7Bbase_model:quantized:Qwen/Qwen3-1.7Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~940.8 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3-1.7B-Depth-Aggressive-Q4_K_M.ggufGGUFQ4_K_M940.8 MBDownload

Model Details

Model IDMa7ee7/Qwen3-1.7B-Depth-Aggressive-Q4_K_M-GGUF
AuthorMa7ee7
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3-1.7B
Last modified2026-07-16T23:34:18.000Z

Model README

---

base_model: Qwen/Qwen3-1.7B

library_name: transformers

pipeline_tag: text-generation

license: apache-2.0

tags:

  • qwen3
  • gguf
  • llama-cpp
  • pruning
  • structured-pruning
  • q4_k_m

---

Qwen3-1.7B Depth-Aggressive — Q4_K_M GGUF

A structured depth-pruned and recovery-trained variant of

Qwen/Qwen3-1.7B, converted to GGUF

and quantized as Q4_K_M for llama.cpp-compatible runtimes.

Model changes

  • Transformer layers: 28 → 24
  • Removed source layers: 9, 14, 19, 7
  • Parameters: approximately 1.721B → 1.519B
  • Parameters retained: 88.30%
  • Recovery: 100 chat-recovery optimizer steps
  • Recovery objective: assistant-only cross-entropy plus teacher distillation
  • GGUF quantization: Q4_K_M

File

| File | Quantization | Size | SHA-256 |

|---|---:|---:|---|

| Qwen3-1.7B-Depth-Aggressive-Q4_K_M.gguf | Q4_K_M | 940.82 MiB | eb836a316bad3d3672d0e182533f8cb306f5783ae68d9fdb8955c5d2fb1184ae |

llama.cpp

llama-cli -m Qwen3-1.7B-Depth-Aggressive-Q4_K_M.gguf -cnv -ngl 99

Notes

The pruning experiment showed the aggressive depth-pruned 1.7B checkpoint at

24 layers and about 1.519B parameters before GGUF quantization. This repository

contains the quantized GGUF build, not the full-precision Transformers weights.

License

The base model is distributed under the Apache 2.0 license. Review the original

Qwen model repository for its complete terms and documentation.

Run Ma7ee7/Qwen3-1.7B-Depth-Aggressive-Q4_K_M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models