kingjones777/Ling-3.0-flash-Research-ROCmFP4-STRIX_LEAN-GGUF overview
Ling 3.0 flash — Research ROCmFP4 STRIX LEAN ftype 106 This is a research build. It is not my serving build — grab Ling 3.0 flash ROCmFP4 STRIX LEAN GGUF https…
Runs locally from ~63.35 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Ling-3.0-flash-I-Research-ROCmFP4-STRIX_LEAN.gguf | GGUF | GGUF | 63.35 GB | Download |
Model Details
| Model ID | kingjones777/Ling-3.0-flash-Research-ROCmFP4-STRIX_LEAN-GGUF |
|---|---|
| Author | kingjones777 |
| Pipeline | text-generation |
| License | mit |
| Base model | inclusionAI/Ling-3.0-flash |
| Last modified | 2026-08-23T12:50:36.000Z |
Model README
---
license: mit
base_model: inclusionAI/Ling-3.0-flash
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
tags: [gguf, llama.cpp, rocm, amd, rocmfp4, rocmfpx, strix-halo, gfx1151, bailingmoe3, research]
---
Ling-3.0-flash — Research ROCmFP4 STRIX_LEAN (ftype 106)
This is a research build. It is not my serving build — grab Ling-3.0-flash-ROCmFP4-STRIX_LEAN-GGUF instead if you just want to run the model.
⚠️ Read this before you download — "Research" here means the QUANT RECIPE
I build a research family alongside my serving builds so I can measure what individual quantization choices actually cost. "Research" refers to how I quantized this file. It is the same aligned inclusionAI/Ling-3.0-flash checkpoint as my serving build — same weights, same behavior, same alignment. It is not uncensored, not ablated, not a different model. If you came here expecting a modified model, this isn't one.
What's different is one deliberate choice: I ran the bare ftype 106 path — no --output-tensor-type override — so the LM head stays at 4-bit.
| | this research build | my serving build |
|---|---|---|
| output.weight | Q4_0_ROCMFP4_FAST (101) — 4-bit | Q6_K (14) — protected |
| token_embd.weight | Q5_K (13) | Q5_K (13) |
| decode | 32.0 t/s | 30.7 t/s |
| bytes | 68,020,249,248 | 68,136,565,408 |
So this is the faster, cheaper, lower-fidelity half of a matched pair. It exists so the cost of head protection is a measured number instead of an assumption: on this model, protecting the head costs about 1.3 t/s and 116 MB. I think that's worth paying, which is why my serving build pays it — but now you can see the trade instead of taking my word for it.
> ⚠️ A 4-bit LM head is a real quality risk on this architecture. I've seen unprotected heads take a model from 4/5 to 1/5 on my own evals. Use this build to study that effect, not to serve users.
> ⚠️ Requires a ROCmFPX-capable llama.cpp build — these tensor types aren't in mainline, so stock llama.cpp / Ollama / LM Studio won't load it.
Receipts
| | |
|---|---|
| bytes | 68,020,249,248 |
| sha256 | f7eff54932de653b7ac2b1ee04d0f5d0d9c26a94b484bbcea07a110fef6588e8 |
| general.file_type | 106 |
| arch / tensors / ctx | bailingmoe3 / 938 / 262144 |
| blocks | 43, nextn_predict_layers = 1 (MTP head at blk.42) |
| dry-run → built | 64862.93 MiB (4.27 BPW) → +6.235 MiB |
| decode | 32.0 t/s (prompt 103.4 t/s) |
Heads read back from the built file by exact tensor name — never a substring, since output.weight matches inside attn_output.weight:
| tensor (exact name) | type |
|---|---|
| output.weight | Q4_0_ROCMFP4_FAST (101) |
| token_embd.weight | Q5_K (13) |
Source: inclusionAI/Ling-3.0-flash @ 42766a814ab117e75e2e61465d5e131b72d931a3 → 255,091,083,232-byte BF16 GGUF (938 tensors) → quantized.
llama-quantize \
Ling-3.0-flash-I-BF16.gguf \
Ling-3.0-flash-I-Research-ROCmFP4-STRIX_LEAN.gguf \
Q4_0_ROCMFP4_STRIX_LEAN 16
Note the absence of --output-tensor-type — that omission is the experiment.
Measured on my own hardware (gfx1151, Strix Halo). Every number is read back from the built file.
Run kingjones777/Ling-3.0-flash-Research-ROCmFP4-STRIX_LEAN-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models