kingjones777/BTL-4-Q4_0_ROCMFP4_STRIX-GGUF overview
BTL 4 — Q4 0 ROCMFP4 STRIX GGUF First ROCmFP4 quantization of badtheorylabs/BTL 4 https://huggingface.co/badtheorylabs/BTL 4 that exists anywhere verified agai…
Runs locally from ~857.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | kingjones777/BTL-4-Q4_0_ROCMFP4_STRIX-GGUF |
|---|---|
| Author | kingjones777 |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | badtheorylabs/BTL-4 |
| Last modified | 2026-08-07T20:01:01.000Z |
Model README
---
license: apache-2.0
base_model: badtheorylabs/BTL-4
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- gguf
- rocmfp4
- strix-halo
- gfx1151
- ryzen-ai-max
- moe
- qwen3.5
- vision
- llama-cpp
language:
- en
- zh
---
BTL-4 — Q4_0_ROCMFP4_STRIX GGUF
**First ROCmFP4 quantization of badtheorylabs/BTL-4
that exists anywhere** (verified against the Hub before publish). Built for AMD Ryzen AI Max+ /
Strix Halo (gfx1151).
BTL-4 is a Qwen3.5 MoE vision model: architecture Qwen3_5MoeForConditionalGeneration /
model_type: qwen3_5_moe, 40 layers, 256 experts / 8 active, hidden 2048, shared-expert
512, vocab 248320. Upstream text_config.mtp_num_hidden_layers: 0 — there is no MTP head.
Do not enable speculative / MTP drafting against this file; a spec flag with no tensors is a
silent garbage drafter.
Same architecture family as KAT-Coder-V2.5-Dev (8-of-256 active-param shape), which is where
ROCmFP4 STRIX already beat Q4_K_M on Strix Halo. This build reproduces that pattern on BTL-4.
Files
| File | Notes |
|---|---|
| BTL-4-Q4_0_ROCMFP4_STRIX.gguf | Text MoE trunk, recipe 105 (Q4_0_ROCMFP4_STRIX) |
| mmproj-BTL-4-F16.gguf | Vision projector (F16), load with -mm / --mmproj |
Single-shard: 17.39 GiB text + 0.84 GiB mmproj. Under the HF 50 GB file cap — no split.
Measured A/B (gfx1151, 128 GB unified, ROCm)
Equal conditions for both quants:
- Binary:
charlie12345/ROCmFPXLaguna Strix export6255cc8
(export: Laguna Strix ROCmFP4 recipe on top of charlie12345/ROCmFPX@3edc3d3)
- Runtime:
-dio,HSA_OVERRIDE_GFX_VERSION=11.5.1,GGML_HIP_ENABLE_UNIFIED_MEMORY=1,
-ngl 999, --no-warmup, --ignore-eos
- 256-token generations, nonce-prefixed prompts (prefix cache defeated;
cache_n == 0
asserted every run)
- 3-run medians; Q4_K_M baseline run twice (noise control)
- Quality: greedy (
temp 0,top_k 1), thinking disabled via chat template kwargs, 10 prompts
Size
| Artifact | Bytes | BPW (real) |
|---|---:|---:|
| Upstream BF16 (HF) | 70,242,700,904 | bf16 |
| F16 GGUF intermediate | 69,376,637,024 | 16.01 |
| This ROCmFP4 STRIX | 18,664,879,904 | 4.31 (dry-run + build; advertised ~4.49) |
| Same-model Q4_K_M control | 21,166,757,664 | 4.88 |
STRIX is −11.8% smaller than the Q4_K_M control.
Decode throughput (tok/s, median of 3)
| Context | Q4_K_M A | Q4_K_M B | ROCmFP4 STRIX | vs doubled baseline |
|---|---:|---:|---:|---:|
| ~8K prompt | 54.28 | 54.02 | 61.06 | +12.8% |
| ~32K prompt | 46.59 | 46.72 | 51.83 | +11.1% |
Prompt-eval medians (tok/s): STRIX 1146.8 @8K / 857.4 @32K; Q4_K_M ~1105 / ~835.
Quality (10-prompt greedy battery)
| | Score |
|---|---:|
| ROCmFP4 STRIX | 10 / 10 |
| Q4_K_M | 10 / 10 |
Equal quality, clear speed win, smaller file → ship.
Recipe notes
- Prefer
Q4_0_ROCMFP4_STRIX(105) over_STRIX_LEAN(106): same speed class, better quality
headroom on this fork’s prior Strix A/Bs.
- Real dry-run BPW was 4.31, not the type’s advertised ~4.49. Always read dry-run.
- Converted from the official BF16 with the fork’s
convert_hf_to_gguf.py
(Qwen3_5MoeForConditionalGeneration + --mmproj). No --mtp.
Launch (Strix Halo / gfx1151)
env HSA_OVERRIDE_GFX_VERSION=11.5.1 \
GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
-m BTL-4-Q4_0_ROCMFP4_STRIX.gguf \
--mmproj mmproj-BTL-4-F16.gguf \
-ngl 999 -dio --no-warmup --jinja \
-c 32768 --parallel 1 \
--temp 0.0 --top-k 1
Do not pass MTP / speculative draft flags. Upstream has zero MTP layers.
Requires a ROCmFP4-capable llama.cpp build (ROCmFPX / Laguna Strix recipe), not stock llama.cpp
alone, for the ROCmFP4 tensor types.
License
Inherited from badtheorylabs/BTL-4
(Apache-2.0 on the base card at publish time). All credit to the base authors; this repo is a
quantization only (base_model_relation: quantized).
Run kingjones777/BTL-4-Q4_0_ROCMFP4_STRIX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models