GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-GGUF overview

<img src="https://serve.thelabsource.com/u/6OiIHw.png" alt="adi qwen3.5 4b glm5.2 general" width="800" adi qwen3.5 4b glm5.2 general Part of the ADI Advanced D…

ggufdistillationqwen3.5adiadvanced-data-intelligencetext-generationtool-callingenbase_model:Qwen/Qwen3.5-4Bbase_model:quantized:Qwen/Qwen3.5-4Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~2.52 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
836
Likes
2
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
adi-qwen3.5-4b-glm5.2-general-q4_k_m.ggufGGUFQ4_K_M2.52 GBDownload

Model Details

Model IDAdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-GGUF
AuthorAdvancedDataIntelligence
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3.5-4B
Last modified2026-07-01T01:49:58.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen3.5-4B

tags:

- gguf

- distillation

- qwen3.5

- adi

- advanced-data-intelligence

- text-generation

- tool-calling

language:

- en

pipeline_tag: text-generation

library_name: gguf

---

<img src="https://serve.thelabsource.com/u/6OiIHw.png" alt="adi-qwen3.5-4b-glm5.2-general" width="800">

adi-qwen3.5-4b-glm5.2-general

Part of the ADI (Advanced Data Intelligence) model line — ADI Qwen3 series.

A small, fully local model that reasons and answers like a frontier teacher.

Built by distilling glm-5.2 general-knowledge responses into a Qwen3.5-4B

student with a bf16 LoRA fine-tune, then merged, converted, and quantized to GGUF.

The student base retains native tool calling and a long context window.

Capabilities

| Size | Context | Input | Output | Tools |

|---|---|---|---|---|

| 2.7 GB | 262K | 🅣 Text | Text | ✅ |

| | |

|---|---|

| Base model | Qwen/Qwen3.5-4B |

| Teacher | glm-5.2 (responses distilled, thinking disabled) |

| Method | bf16 LoRA SFT (rank 16) → merge → GGUF |

| Quantization | Q4_K_M (~2.7 GB) |

| License | Apache-2.0 (inherited from Qwen3.5-4B) |

| Context | 262K (inherited from base) |

| Tool calling | Supported (inherited from base) |

Run it

Pull directly into Ollama:

ollama run hf.co/AdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-GGUF:Q4_K_M

Or download the .gguf and point any llama.cpp-based runtime at it.

Try it live

A hosted demo is available as a Hugging Face Space — chat with the model directly in your browser, no install required.

<a href="https://huggingface.co/spaces/AdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-demo">

<img src="https://serve.thelabsource.com/u/4Kb3iS.gif" alt="adi-qwen3.5-4b-glm5.2-general live demo" width="800">

</a>

▶ Launch the demo

Chat with the model directly in your browser — no install required.

What this model is

This is a knowledge distillation: a strong teacher (glm-5.2) generated

high-quality answers across ~2,000 diverse general-knowledge prompts, and the

Qwen3.5-4B student was fine-tuned to imitate them. The result reasons and

responds noticeably more like its teacher on general topics, while staying small

enough to run on a single consumer GPU.

What distillation does — and doesn't do. It transfers the teacher's

reasoning style and answer quality, not net-new facts. A 4B model won't become

an encyclopedia. For raw factual recall, retrieval-augmented generation (RAG) is

the right tool, not fine-tuning. What you get here is a 4B that *structures and

explains* like a much larger model on topics it already partly knows.

Training

| Metric | Value |

|---|---|

| Training pairs | 2,068 |

| Teacher tokens generated | ~1.36M |

| Epochs | 3 |

| Steps | 777 |

| Final train loss | 0.9346 |

| LoRA rank / alpha | 16 / 16 |

| Trainable params | 21.2M (0.47% of 4.56B) |

| Precision | bf16 (not 4-bit — see note) |

| Hardware | single RTX 5060 Ti (16 GB) |

| Training time | 2h 53m |

The seed prompts were drawn from the human-written

Databricks Dolly-15k

dataset (filtered to remove items requiring an attached context passage, then

deduplicated). The teacher was queried with thinking disabled so the student

learns clean final answers rather than chain-of-thought it is too small to

reproduce well.

Notes for re-builders

  • Qwen3.5 trains in bf16 LoRA, not 4-bit QLoRA. Its gated-delta / Mamba-hybrid

layers quantize poorly during training; 4-bit costs accuracy. bf16 LoRA uses

~10 GB on a 4B — comfortable on a 16 GB card.

  • Version pins: Qwen3.5 requires transformers >= 5.2.0 to be recognized, while

the Unsloth training stack caps at <= 5.5.0. The working version is

transformers == 5.5.0 with numpy < 2.3.

  • GGUF conversion was done with llama.cpp's convert_hf_to_gguf.py, which already

understands the Qwen3.5 SSM/MTP architecture.

Intended use

General-purpose local assistant: explanations, reasoning, Q&A, and tool-calling

workflows where a small, private, offline-capable model is preferred over a

hosted API. Not intended as a source of authoritative facts without retrieval.

License

Apache-2.0, inherited from the Qwen3.5-4B

base model. You are free to use, modify, and redistribute under the terms of that

license. Distilled training data was generated using glm-5.2; users should review

the teacher model's terms for their own use case.

---

Built at theLAB — Learning. Algorithms. Breakthroughs.

Run AdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models