DKTechin/local-llm-gguf overview
local llm gguf — weight mirror Files loaded by the on device assistant of the KakaoWork desktop app, re uploaded byte identical from their upstream repositorie…
Runs locally from ~1.19 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
Model README
---
license: apache-2.0
tags:
- gguf
- mirror
- llama-cpp
---
local-llm-gguf — weight mirror
Files loaded by the on-device assistant of the KakaoWork desktop app,
re-uploaded byte-identical from their upstream repositories. The app needs a
location whose bytes do not move under it; this repository is that location.
Nothing here is our own work.
Every upstream repository below is Apache-2.0 and ungated. Please prefer the
upstream repositories — they hold the other quantizations and the model cards.
Files
| File | Size | Upstream |
| --- | --- | --- |
| gemma-4-E2B_q4_0-it.gguf | 3.35 GB | google/gemma-4-E2B-it-qat-q4_0-gguf |
| gemma-4-E4B_q4_0-it.gguf | 5.15 GB | google/gemma-4-E4B-it-qat-q4_0-gguf |
| Qwen3.5-2B-Q4_K_M.gguf | 1.28 GB | unsloth/Qwen3.5-2B-GGUF |
| Qwen3.5-4B-Q4_K_M.gguf | 2.74 GB | unsloth/Qwen3.5-4B-GGUF |
kanana-2-3b-instruct-Q4_K_M.gguf is not here — it is our own conversion and
lives in DKTechin/kanana.
Credits
The models are by their creators — Google (Gemma 4) and Alibaba/Qwen
(Qwen3.5) — and the GGUF builds are by the repositories listed above. Their
licenses and terms of use apply to these copies unchanged.
If you maintain one of the upstream repositories and would rather this mirror
not exist, open a discussion and we will remove it.
Run DKTechin/local-llm-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models