GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

kingjones777/Ling-3.0-flash-ROCmFP4-STRIX_LEAN-GGUF overview

Ling 3.0 flash — ROCmFP4 STRIX LEAN ftype 106 This is the one you want if you just want to run Ling 3.0 flash on a Strix Halo box. I quantized inclusionAI/Ling…

ggufllama.cpprocmamdrocmfp4rocmfpxstrix-halogfx1151bailingmoe3text-generationbase_model:inclusionAI/Ling-3.0-flashbase_model:quantized:inclusionAI/Ling-3.0-flashlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~63.46 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
15
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-flash-I-ROCmFP4-STRIX_LEAN.ggufGGUFGGUF63.46 GBDownload

Model Details

Model IDkingjones777/Ling-3.0-flash-ROCmFP4-STRIX_LEAN-GGUF
Authorkingjones777
Pipelinetext-generation
Licensemit
Base modelinclusionAI/Ling-3.0-flash
Last modified2026-08-23T12:50:34.000Z

Model README

---

license: mit

base_model: inclusionAI/Ling-3.0-flash

base_model_relation: quantized

library_name: gguf

pipeline_tag: text-generation

tags: [gguf, llama.cpp, rocm, amd, rocmfp4, rocmfpx, strix-halo, gfx1151, bailingmoe3]

---

Ling-3.0-flash — ROCmFP4 STRIX_LEAN (ftype 106)

This is the one you want if you just want to run Ling-3.0-flash on a Strix Halo box.

I quantized inclusionAI/Ling-3.0-flash (instruct) to ftype 106 Q4_0_ROCMFP4_STRIX_LEAN for AMD gfx1151 — Ryzen AI MAX+ 395, 128 GB unified memory. It's my standard serving build: 4-bit body, lean Q5_K token embeddings, and a protected Q6_K LM head.

> ⚠️ You need a ROCmFPX-capable llama.cpp build. These tensor types (Q4_0_ROCMFP4_, Q_0_ROCMFPX*) are not in mainline, so this will not load in stock llama.cpp, Ollama, or LM Studio.

What kind of build this is

I publish three Ling-3.0-flash builds and they are not interchangeable:

| build | what it's for |

|---|---|

| this one — 106 STRIX_LEAN | default. Fastest sensible quality-per-GB. Serve from this. |

| Research 106 | my research quant recipe — bare ftype path, unprotected head. Comparison work, not serving. |

| Research Q6 AGENT (114) | my research quant recipe at 8-bit. Highest fidelity, biggest, slowest. |

All three come from the same aligned checkpoint. The word "Research" in the other two refers to how I quantized them, not to a different or de-aligned model.

Why the head matters here

tie_word_embeddings = false on this model, so output.weight is a real standalone tensor and --output-tensor-type does actual work. My earlier 4-bit instruct builds left output.weight at 4-bit — that is a real quality defect, not a rounding detail. This build forces the LM head to Q6_K, and I verify it by exact tensor name (output.weight, never a substring — it matches inside attn_output.weight):

| tensor (exact name) | type |

|---|---|

| output.weight | Q6_K (14) |

| token_embd.weight | Q5_K (13) |

⚠️ ftype 106 does not protect the head on its own. If you build this yourself and skip the flag, you get a 4-bit head and a worse model that looks identical from the outside.

Receipts

| | |

|---|---|

| bytes | 68,136,565,408 |

| sha256 | cbf2521b08a6dca0bddf386ae418436b87a0569fc23b380ee2abc83cf093abd4 |

| general.file_type | 106 |

| arch / tensors / ctx | bailingmoe3 / 938 / 262144 |

| blocks | 43, nextn_predict_layers = 1 (MTP head at blk.42) |

| dry-run → built | 64973.85 MiB (4.28 BPW) → +6.242 MiB |

| decode | 30.7 t/s (prompt 107.9 t/s) |

Source: inclusionAI/Ling-3.0-flash @ 42766a814ab117e75e2e61465d5e131b72d931a3, converted to a 255,091,083,232-byte BF16 GGUF (938 tensors), then quantized.

llama-quantize --output-tensor-type q6_K \
  Ling-3.0-flash-I-BF16.gguf \
  Ling-3.0-flash-I-ROCmFP4-STRIX_LEAN.gguf \
  Q4_0_ROCMFP4_STRIX_LEAN 16

Converter note

Upstream emits the KDA gate tensors as blk.N.ssm_f_a / ssm_g_a (vestigial kimi-linear names); the runtime I use wants blk.N.ssm_f / ssm_g. Same shapes, same semantics — a naming dialect. I patched the converter and verified the result: 35× ssm_f + 35× ssm_g, zero _a leftovers (42 layers − 7 full-attention = 35 KDA layers).

Measured on my own hardware. Every number above is read back from the built file.

Run kingjones777/Ling-3.0-flash-ROCmFP4-STRIX_LEAN-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models