GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE β†’
Model Intelligence Sheet

kingjones777/Ling-3.0-flash-base-ROCmFP4-COHERENT-GGUF overview

πŸ”§ Runtime: build the ROCmFPX fork below Stock llama.cpp will not load this file. You need both the bailing hybrid architecture and the ROCmFP4 tensor types in…

ggufllama.cpprocmamdrocmfp4rocmfpxstrix-haloamd-strix-halogfx1151ryzen-ai-maxryzen-ai-max-395radeon-8060smtpspeculative-decodingbase-modelpretrainedlingbailing-hybridquantizedtext-generationbase_model:inclusionAI/Ling-3.0-flash-basebase_model:quantized:inclusionAI/Ling-3.0-flash-baselicense:mitendpoints_compatible

Runs locally from ~67.17 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
164
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-flash-base-Q4_0_ROCMFP4_COHERENT.ggufGGUFQ4_0_ROCMFP4_COHERENT67.17 GBDownload

Model Details

Model IDkingjones777/Ling-3.0-flash-base-ROCmFP4-COHERENT-GGUF
Authorkingjones777
Pipelinetext-generation
Licensemit
Base modelinclusionAI/Ling-3.0-flash-base
Last modified2026-08-28T04:08:00.000Z

Model README

---

license: mit

base_model: inclusionAI/Ling-3.0-flash-base

base_model_relation: quantized

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- llama.cpp

- rocm

- amd

- rocmfp4

- rocmfpx

- strix-halo

- amd-strix-halo

- gfx1151

- ryzen-ai-max

- ryzen-ai-max-395

- radeon-8060s

- mtp

- speculative-decoding

- base-model

- pretrained

- ling

- bailing-hybrid

- quantized

---

> ### πŸ”§ Runtime: build the ROCmFPX fork below

> Stock llama.cpp will not load this file. You need both the bailing-hybrid architecture

> and the ROCmFP4 tensor types in one tree. Upstream

> charlie12345/ROCmFPX has the ROCmFP4 types but

> not bailing-hybrid. Our fork has both:

>

> kingjones30/ROCmFPX β€” a fork of charlie12345/ROCmFPX, branch main.

>

> ```bash

> git clone https://github.com/kingjones30/ROCmFPX.git

> cd ROCmFPX

> cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release

> cmake --build build --target llama-server llama-quantize -j$(nproc)

> ```

>

> Verified 2026-08-27 on gfx1151: clean clone β†’ 0 build errors β†’ llama-server loads a

> bailing-hybrid ROCmFP4 GGUF from this family and generates coherent text.

Ling-3.0-flash-base β€” ROCmFP4 for AMD Strix Halo (gfx1151)

<!-- LEAD-CLAIM:START -->

> βœ… the first ROCmFP4 build of Ling-3.0-flash-base, published with a measured MTP curve

>

> *Checked 2026-08-22 against every public GGUF of this checkpoint. The two other GGUF builds

> β€” mradermacher/Ling-3.0-flash-base-GGUF (Q2_K through Q5_K_S) and

> avar6/Ling-3.0-flash-base-gguf (Q5_K_M, uploaded 2026-08-22) β€” both ship standard

> k-quants. ROCmFP4 is a runtime tensor format that exists only in the

> ROCmFPX fork of llama.cpp. Repository-content

> comparison only β€” no third-party build was run or benchmarked here.*

<!-- LEAD-CLAIM:END -->

A 4-bit ROCmFP4 quantisation of Ling-3.0-flash-base for AMD Ryzen AI Max+ 395 / Radeon 8060S /

gfx1151, with the multi-token-prediction (MTP) draft head preserved.

⚠️ This is a base checkpoint, not an instruct model

Ling-3.0-flash-base is a pretrained / base checkpoint released by inclusionAI for continued

pretraining, domain adaptation and fine-tuning. It is not instruction-tuned. It ships a

chat_template.jinja, but that is a tokenizer asset β€” it does not make the weights conversational.

Prompt it as a text continuation model. For chat, use inclusionAI/Ling-3.0-flash instead.

The file

| | |

|---|---|

| ftype | 102 β€” Q4_0_ROCMFP4_COHERENT |

| size | 72,123,713,664 bytes (67.16 GiB) |

| parameters | 127.49 B (512 experts Γ— 3.9 B, 8 active) |

| architecture | bailing-hybrid β€” hybrid KDA linear attention + MLA |

| tensors | 938 Β· block_count 43 (42 layers + 1 MTP layer) |

| context | 262,144 |

Head protection, verified in the finished file (not merely requested at quantise time):

output.weight        Q6_K     2560 x 157184    402.4M params
token_embd.weight    Q6_K     2560 x 157184    402.4M params
histogram: ROCmFP4(type 100) x545, F32 x390, Q6_K x2, Q8_0 x1

tie_word_embeddings is false on this model, so --output-tensor-type does real work here β€”

the COHERENT tier on its own leaves output.weight at 4-bit. Both heads were forced to Q6_K and

audited on exact tensor names after the build.

Architecture notes

Ling-3.0-flash interleaves two attention types. head_count_kv is a per-layer array β€”

[0,0,0,0,0,1,0,0,0,0,0,1,…] β€” where 0 marks a KDA linear-attention layer and 1 a full MLA

layer: 1 MLA layer in every 6. MLA uses a compressed KV path (kv_lora_rank 512) with a plain

wide query projection (q_lora_rank: null). The blk.42 MTP layer is retained in full, including

nextn.eh_proj, nextn.enorm, nextn.hnorm and nextn.shared_head_norm, with the unfused

attn_k_b / attn_v_b form that the MTP path requires.

Measured throughput

AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), ROCm 7.2.4, 128 GB unified memory.

llama-cli, -dio -ngl 999 -st -c 2048 -n 512 --temp 0 --seed 1234, 5 repetitions per config.

| config | flags | generation (median) | range |

|---|---|---:|---|

| no drafter | --spec-type none | 36.6 t/s | 36.6 – 36.6 |

| MTP n-max 3 | --spec-type draft-mtp --spec-draft-ngl 999 --spec-draft-n-max 3 | 42.3 t/s | 42.3 – 42.5 |

MTP is worth +15.6% on this model, with disjoint ranges and a 0.0% baseline spread across

5 repetitions.

The two sibling checkpoints measure the same baseline to the decimal on identical hardware and

flags, which is the cross-check for this figure:

| checkpoint | no drafter | MTP n-max 3 | effect |

|---|---:|---:|---:|

| Ling-3.0-flash-base (this file) | 36.6 t/s | 42.3 t/s | +15.6% |

| Ling-3.0-flash-base-30T | 36.6 t/s | 41.4 t/s | +13.1% |

| Ling-3.0-flash-base-midtrain | 36.6 t/s | 43.0 t/s | +17.5% |

MTP is reliably positive across the whole Ling-3.0-flash base family. It is not reliable on

Ling-3.0-tiny, where the same measurement gives +7.5% / +5.1% / βˆ’4.4% across the three

checkpoints β€” the draft head is trained with the model, so its value belongs to the specific

(size, checkpoint) pair rather than to the architecture. Measure before enabling it.

n-max 3 is the recommended setting. Deeper speculation costs prefill throughput: prompt

processing measures ~90–98 t/s at n3 versus ~122–130 t/s with no drafter.

⚠️ MTP is not output-deterministic here

At --temp 0 with a fixed seed, the no-drafter arm produced **byte-identical text across all 5

repetitions**. Both MTP arms did not β€” n-max 3 produced 3 distinct outputs across 5 runs, n-max 5

produced 4, and every MTP output differed from the no-drafter output. If bit-reproducible generation

matters more to you than +15.6%, run without the draft head.

Requirements

This file uses the ROCmFP4 tensor format and the bailing-hybrid architecture. It requires a build

of ROCmFPX that carries both. Stock llama.cpp will

not load it. Verify with strings libllama.so | grep bailing-hybrid β€” the architecture table lives

in the shared library, not in the thin CLI binary.

llama-cli -m Ling-3.0-flash-base-Q4_0_ROCMFP4_COHERENT.gguf \
  -dio -ngl 999 -c 2048 -n 512 \
  --spec-type draft-mtp --spec-draft-ngl 999 --spec-draft-n-max 3 \
  -p "The history of mathematics begins in ancient times. One of the earliest known"

-dio (direct I/O) is recommended. At -ngl 999 the HIP backend copies offloaded tensors out of

file-backed pages into device allocations, so without direct I/O the source pages and the device

buffer are resident simultaneously β€” roughly twice the model size, which is tight on a 128 GB box.

Sample output

Continuation from "The history of mathematics begins in ancient times. One of the earliest known":

> mathematical texts is the Rhind Mathematical Papyrus, which dates back to around 1650 BCE in

> ancient Egypt. This papyrus, named after the Scottish antiquarian Alexander Henry Rhind who

> purchased it in 1858, contains a collection of mathematical problems and solutions that provide

> insight into the mathematical knowledge of the time. The problems cover a range of topics,

> including arithmetic, geometry, and algebra, and demonstrate the Egyptians' practical approach

> to mathematics.

Not measured

Perplexity is not published for this build. No perplexity figure is quoted because none was

completed on this hardware β€” a 127 B model at this size exceeds a practical evaluation budget on a

single Strix Halo box. Quality evidence here is limited to the coherence check above and the

verified tensor-level audit.

Sibling quantisations

Also built but not published: Q4_0_ROCMFP4_LEAN (72.07 GB β€” within 0.07% of COHERENT, since the

tiers differ only in token-embedding type, which is negligible on a 127 B model),

Q6_0_ROCMFPX (105.22 GB) and Q8_0_ROCMFPX (131.66 GB β€” exceeds the addressable memory of a

128 GB Strix Halo box).

Provenance

Converted from inclusionAI/Ling-3.0-flash-base at revision

0a96b0a3c68f93e5565d6f2d56c98557bded74fb to BF16 GGUF (255.1 GB, 938 tensors), then quantised to

ftype 102 with --output-tensor-type q6_K. Licence MIT, inherited from the base model.

Run kingjones777/Ling-3.0-flash-base-ROCmFP4-COHERENT-GGUF with guIDE

Download guIDE β€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE β†’ Β· Browse 524k+ models Β· Compare models

Source: Hugging Face Β· Compare models