kingjones777/Qwen3-VL-8B-Instruct-ROCmFP4-GGUF overview
Qwen3 VL 8B Instruct — ROCmFP4 / ROCmFPX GGUF AMD native FP4 / FP8 GGUF builds of Qwen/Qwen3 VL 8B Instruct for RDNA3.5 / Strix Halo gfx1151 . A vision languag…
Runs locally from ~1.08 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3-VL-8B-Instruct-Q4_0_ROCMFP4_COHERENT.gguf | GGUF | Q4_0_ROCMFP4_COHERENT | 4.60 GB | Download |
| Qwen3-VL-8B-Instruct-Q6_0_ROCMFPX_AGENT.gguf | GGUF | Q6_0_ROCMFPX_AGENT | 7.22 GB | Download |
| Qwen3-VL-8B-Instruct-Q8_0_ROCMFPX.gguf | GGUF | Q8_0_ROCMFPX | 7.91 GB | Download |
| Qwen3-VL-8B-Instruct-Q8_0_ROCMFPX_AGENT.gguf | GGUF | Q8_0_ROCMFPX_AGENT | 8.02 GB | Download |
| mmproj-BF16.gguf | GGUF | BF16 | 1.08 GB | Download |
Model Details
| Model ID | kingjones777/Qwen3-VL-8B-Instruct-ROCmFP4-GGUF |
|---|---|
| Author | kingjones777 |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | Qwen/Qwen3-VL-8B-Instruct |
| Last modified | 2026-08-18T02:02:27.000Z |
Model README
---
license: apache-2.0
base_model: Qwen/Qwen3-VL-8B-Instruct
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- gguf
- rocm
- rocmfp4
- amd
- strix-halo
- gfx1151
- llama.cpp
- qwen
- vision-language
---
Qwen3-VL-8B-Instruct — ROCmFP4 / ROCmFPX GGUF
AMD-native FP4 / FP8 GGUF builds of Qwen/Qwen3-VL-8B-Instruct for RDNA3.5 / Strix Halo
(gfx1151). A vision-language model — the bundled mmproj-BF16.gguf is the point of the build.
Variants
| file | ftype | size | decode | spread |
|---|---|---|---|---|
| Qwen3-VL-8B-Instruct-Q4_0_ROCMFP4_COHERENT.gguf | 102 | 4.60 GiB | 44.86 t/s | 1.0013 |
| Qwen3-VL-8B-Instruct-Q6_0_ROCMFPX_AGENT.gguf | 114 | 7.22 GiB | 28.50 t/s | 1.0004 |
| Qwen3-VL-8B-Instruct-Q8_0_ROCMFPX.gguf | 111 | 7.91 GiB | 26.29 t/s | 1.0015 |
| Qwen3-VL-8B-Instruct-Q8_0_ROCMFPX_AGENT.gguf | 115 | 8.02 GiB | 26.08 t/s | 1.0012 |
Measured on an idle Ryzen AI MAX+ 395 (Strix Halo, gfx1151, ROCm 7.2.4):
-ngl 999 -c 4096 -fa on -fit off -np 1, 300-token generations, 12 samples with
two warm-ups on the same prompt as the measurement. Spread = slowest/fastest.
⚠️ An earlier pass of these same files, taken while other jobs shared the GPU, read
20% low with 20%+ spread. On this hardware a co-resident job is the single largest
source of benchmark error — measure on an idle box or say what else was resident.
Vision verified 4/4 on a four-quadrant colour image (red / blue / yellow / green) with
the bundled mmproj-BF16.gguf. ⛔ Vision needs -fa off.
Verification
Every artifact was loaded on real hardware and checked for: exact stat bytes vs the
--dry-run projection (a constant header delta; a varying one means truncation), the
actual token_embd / output.weight types, three correctness answers asserted against
content + reasoning with finish_reason recorded, and a decode median.
Credits
FP4/FP8 tensor types from the ROCmFPX fork of llama.cpp. These types do not exist in
mainline llama.cpp — a ROCmFPX-capable build is required to load them.
Run kingjones777/Qwen3-VL-8B-Instruct-ROCmFP4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models