GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

bluehawana/Qwen3.8-27B-Q8-GGUF overview

Qwen3.8 27B — GGUF Q8 0 Single file Q8 0 8 bit, near lossless GGUF of Qwen/Qwen3.8 27B https://huggingface.co/Qwen/Qwen3.8 27B — 28.9 GB. Quantized by AtomicCh…

ggufollamallama.cppapple-siliconqwen3.8q8_0base_model:Qwen/Qwen3.8-27Bbase_model:quantized:Qwen/Qwen3.8-27Blicense:apache-2.0endpoints_compatibleregion:usimatrixconversational

Runs locally from ~26.90 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-27B-Q8_0.ggufGGUFQ8_026.90 GBDownload

Model Details

Model IDbluehawana/Qwen3.8-27B-Q8-GGUF
Authorbluehawana
Pipeline
Licenseapache-2.0
Base modelQwen/Qwen3.8-27B
Last modified2026-08-18T09:43:03.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen3.8-27B

base_model_relation: quantized

tags:

- gguf

- ollama

- llama.cpp

- apple-silicon

- qwen3.8

- q8_0

---

Qwen3.8-27B — GGUF Q8_0

Single-file Q8_0 (8-bit, near-lossless) GGUF of

Qwen/Qwen3.8-27B — 28.9 GB.

Quantized by AtomicChat;

re-hosted here alongside our Apple-Silicon serving research for one-command use.

Run it

# Ollama (from the Ollama registry — easiest)
ollama run bluehawana/qwen3.8-27b-q8

# Ollama (straight from this repo)
ollama run hf.co/bluehawana/Qwen3.8-27B-Q8-GGUF

# llama.cpp / LM Studio / Jan: download Qwen3.8-27B-Q8_0.gguf directly

Needs ≥48 GB unified memory on Apple Silicon (comfortable at 64 GB+).

Concurrent serving on a Mac

This model serves 16 concurrent requests with zero errors on an M-series

128 GB Mac — benchmarks across Ollama / oMLX / SGLang-MLX, plus the SGLang

patches that make it possible (upstream PR

sgl-project/sglang#35137):

bluehawana/qwen3.8-27b-apple-silicon-concurrency

Quick concurrent Ollama serving:

OLLAMA_HOST=127.0.0.1:11500 OLLAMA_NUM_PARALLEL=16 OLLAMA_CONTEXT_LENGTH=8192 ollama serve
# OpenAI-compatible endpoint: http://127.0.0.1:11500/v1

Credits: base model © Qwen (Apache-2.0) · Q8_0 quant by AtomicChat · benchmarks & patches by bluehawana.

Run bluehawana/Qwen3.8-27B-Q8-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models