Adiuk/Qwen2.5-32B-Instruct-AdiTurbo-GGUF overview
Qwen2.5 32B Instruct AdiTurbo GGUF Low bit GGUF quantizations of Qwen/Qwen2.5 32B Instruct https://huggingface.co/Qwen/Qwen2.5 32B Instruct , produced with the…
Runs locally from ~8.41 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
Model README
---
license: apache-2.0
base_model: Qwen/Qwen2.5-32B-Instruct
tags: [gguf, quantized, aditurbo, llama.cpp]
---
Qwen2.5-32B-Instruct-AdiTurbo-GGUF
Low-bit GGUF quantizations of Qwen/Qwen2.5-32B-Instruct,
produced with the AdiTurbo Engine (private llama.cpp fork, build 8661).
Quantized with a fresh 50-chunk wikitext-2 importance matrix. Intended for
offline/on-device inference (Eyla AIOS). See the AdiTurbo benchmark for
perplexity/throughput numbers.
Files
Qwen2.5-32B-Instruct-TQ3_0.ggufQwen2.5-32B-Instruct-TQ4_0.ggufQwen2.5-32B-Instruct-Q3_K_M.gguf
This is a derivative quantization; the original model's license (apache-2.0) applies.
Run Adiuk/Qwen2.5-32B-Instruct-AdiTurbo-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models