GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

kingjones777/Ling-3.0-flash-Research-ROCmFP4-STRIX_LEAN-GGUF overview

Ling 3.0 flash — Research ROCmFP4 STRIX LEAN ftype 106 This is a research build. It is not my serving build — grab Ling 3.0 flash ROCmFP4 STRIX LEAN GGUF https…

ggufllama.cpprocmamdrocmfp4rocmfpxstrix-halogfx1151bailingmoe3researchtext-generationbase_model:inclusionAI/Ling-3.0-flashbase_model:quantized:inclusionAI/Ling-3.0-flashlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~63.35 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
12
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-flash-I-Research-ROCmFP4-STRIX_LEAN.ggufGGUFGGUF63.35 GBDownload

Model Details

Model IDkingjones777/Ling-3.0-flash-Research-ROCmFP4-STRIX_LEAN-GGUF
Authorkingjones777
Pipelinetext-generation
Licensemit
Base modelinclusionAI/Ling-3.0-flash
Last modified2026-08-23T12:50:36.000Z

Model README

---

license: mit

base_model: inclusionAI/Ling-3.0-flash

base_model_relation: quantized

library_name: gguf

pipeline_tag: text-generation

tags: [gguf, llama.cpp, rocm, amd, rocmfp4, rocmfpx, strix-halo, gfx1151, bailingmoe3, research]

---

Ling-3.0-flash — Research ROCmFP4 STRIX_LEAN (ftype 106)

This is a research build. It is not my serving build — grab Ling-3.0-flash-ROCmFP4-STRIX_LEAN-GGUF instead if you just want to run the model.

⚠️ Read this before you download — "Research" here means the QUANT RECIPE

I build a research family alongside my serving builds so I can measure what individual quantization choices actually cost. "Research" refers to how I quantized this file. It is the same aligned inclusionAI/Ling-3.0-flash checkpoint as my serving build — same weights, same behavior, same alignment. It is not uncensored, not ablated, not a different model. If you came here expecting a modified model, this isn't one.

What's different is one deliberate choice: I ran the bare ftype 106 path — no --output-tensor-type override — so the LM head stays at 4-bit.

| | this research build | my serving build |

|---|---|---|

| output.weight | Q4_0_ROCMFP4_FAST (101) — 4-bit | Q6_K (14) — protected |

| token_embd.weight | Q5_K (13) | Q5_K (13) |

| decode | 32.0 t/s | 30.7 t/s |

| bytes | 68,020,249,248 | 68,136,565,408 |

So this is the faster, cheaper, lower-fidelity half of a matched pair. It exists so the cost of head protection is a measured number instead of an assumption: on this model, protecting the head costs about 1.3 t/s and 116 MB. I think that's worth paying, which is why my serving build pays it — but now you can see the trade instead of taking my word for it.

> ⚠️ A 4-bit LM head is a real quality risk on this architecture. I've seen unprotected heads take a model from 4/5 to 1/5 on my own evals. Use this build to study that effect, not to serve users.

> ⚠️ Requires a ROCmFPX-capable llama.cpp build — these tensor types aren't in mainline, so stock llama.cpp / Ollama / LM Studio won't load it.

Receipts

| | |

|---|---|

| bytes | 68,020,249,248 |

| sha256 | f7eff54932de653b7ac2b1ee04d0f5d0d9c26a94b484bbcea07a110fef6588e8 |

| general.file_type | 106 |

| arch / tensors / ctx | bailingmoe3 / 938 / 262144 |

| blocks | 43, nextn_predict_layers = 1 (MTP head at blk.42) |

| dry-run → built | 64862.93 MiB (4.27 BPW) → +6.235 MiB |

| decode | 32.0 t/s (prompt 103.4 t/s) |

Heads read back from the built file by exact tensor name — never a substring, since output.weight matches inside attn_output.weight:

| tensor (exact name) | type |

|---|---|

| output.weight | Q4_0_ROCMFP4_FAST (101) |

| token_embd.weight | Q5_K (13) |

Source: inclusionAI/Ling-3.0-flash @ 42766a814ab117e75e2e61465d5e131b72d931a3 → 255,091,083,232-byte BF16 GGUF (938 tensors) → quantized.

llama-quantize \
  Ling-3.0-flash-I-BF16.gguf \
  Ling-3.0-flash-I-Research-ROCmFP4-STRIX_LEAN.gguf \
  Q4_0_ROCMFP4_STRIX_LEAN 16

Note the absence of --output-tensor-type — that omission is the experiment.

Measured on my own hardware (gfx1151, Strix Halo). Every number is read back from the built file.

Run kingjones777/Ling-3.0-flash-Research-ROCmFP4-STRIX_LEAN-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models