GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE β†’
Model Intelligence Sheet

kingjones777/Ling-3.0-flash-base-30T-ROCmFP4-COHERENT-GGUF overview

πŸ”§ Runtime: build the ROCmFPX fork below Stock llama.cpp will not load this file. You need both the bailing hybrid architecture and the ROCmFP4 tensor types in…

ggufllama.cpprocmamdrocmfp4rocmfpxstrix-haloamd-strix-halogfx1151ryzen-ai-maxryzen-ai-max-395radeon-8060smtpspeculative-decodingbase-modelpretrainedlingbailing-hybridquantizedtext-generationbase_model:inclusionAI/Ling-3.0-flash-base-30Tbase_model:quantized:inclusionAI/Ling-3.0-flash-base-30Tlicense:mitendpoints_compatible

Runs locally from ~67.17 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
92
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-flash-base-30T-Q4_0_ROCMFP4_COHERENT.ggufGGUFQ4_0_ROCMFP4_COHERENT67.17 GBDownload

Model Details

Model IDkingjones777/Ling-3.0-flash-base-30T-ROCmFP4-COHERENT-GGUF
Authorkingjones777
Pipelinetext-generation
Licensemit
Base modelinclusionAI/Ling-3.0-flash-base-30T
Last modified2026-08-28T04:08:14.000Z

Model README

---

license: mit

base_model: inclusionAI/Ling-3.0-flash-base-30T

base_model_relation: quantized

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- llama.cpp

- rocm

- amd

- rocmfp4

- rocmfpx

- strix-halo

- amd-strix-halo

- gfx1151

- ryzen-ai-max

- ryzen-ai-max-395

- radeon-8060s

- mtp

- speculative-decoding

- base-model

- pretrained

- ling

- bailing-hybrid

- quantized

---

> ### πŸ”§ Runtime: build the ROCmFPX fork below

> Stock llama.cpp will not load this file. You need both the bailing-hybrid architecture

> and the ROCmFP4 tensor types in one tree. Upstream

> charlie12345/ROCmFPX has the ROCmFP4 types but

> not bailing-hybrid. Our fork has both:

>

> kingjones30/ROCmFPX β€” a fork of charlie12345/ROCmFPX, branch main.

>

> ```bash

> git clone https://github.com/kingjones30/ROCmFPX.git

> cd ROCmFPX

> cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release

> cmake --build build --target llama-server llama-quantize -j$(nproc)

> ```

>

> Verified 2026-08-27 on gfx1151: clean clone β†’ 0 build errors β†’ llama-server loads a

> bailing-hybrid ROCmFP4 GGUF from this family and generates coherent text.

Ling-3.0-flash-base-30T β€” ROCmFP4 for AMD Strix Halo (gfx1151)

> βœ… the first ROCmFP4 build of Ling-3.0-flash-base-30T, published with a measured MTP curve

>

> *Checked 2026-08-22 against every public GGUF of this checkpoint. The only other GGUF build

> (avar6/Ling-3.0-flash-base-30T-gguf) ships a single standard k-quant, Q5_K_M. ROCmFP4 is a

> runtime tensor format that exists only in the

> ROCmFPX fork of llama.cpp. Repository-content

> comparison only β€” no third-party build was run or benchmarked here.*

A 4-bit ROCmFP4 quantisation of Ling-3.0-flash-base-30T for AMD Ryzen AI Max+ 395 /

Radeon 8060S / gfx1151, with the multi-token-prediction (MTP) draft head preserved.

⚠️ This is a base checkpoint, not an instruct model

Ling-3.0-flash-base-30T is a pretrained / base checkpoint released by inclusionAI for

continued pretraining, domain adaptation and fine-tuning. It is not instruction-tuned. It ships

a chat_template.jinja, but that is a tokenizer asset β€” it does not make the weights

conversational. Prompt it as a text continuation model. For chat, use

inclusionAI/Ling-3.0-flash instead.

30T is the 30-trillion-token pretraining checkpoint β€” the earliest of the three public stages, taken before long-context extension.

⚠️ Short context: 8,192 tokens

This checkpoint declares context_length = 8,192 and rope_theta = 10000, not the 262,144 / 6,000,000 of the other two Flash base

checkpoints β€” it predates the long-context extension stage. Do not assume 262k here; the

GGUF metadata carries the real value and llama.cpp will honour it.

The file

| | |

|---|---|

| ftype | 102 β€” Q4_0_ROCMFP4_COHERENT |

| size | 72,123,709,248 bytes (67.17 GiB) |

| parameters | 127.49 B (512 experts Γ— 3.9 B, 8 active) |

| architecture | bailing-hybrid β€” hybrid KDA linear attention + MLA |

| tensors | 938 Β· block_count 43 (42 layers + 1 MTP layer) |

| context | 8,192 |

| rope_theta | 10000 |

Head protection, verified in the finished file (not merely requested at quantise time, and

re-audited after the metadata rename that produced the final bytes above):

output.weight        Q6_K
token_embd.weight    Q6_K
histogram: ROCmFP4(type 100) x545, F32 x390, Q6_K x2, Q8_0 x1

tie_word_embeddings is false on this model, so --output-tensor-type does real work here β€”

the COHERENT tier on its own leaves output.weight at 4-bit. Both heads were forced to Q6_K and

audited on exact tensor names.

Architecture notes

Ling-3.0-flash interleaves two attention types. head_count_kv is a per-layer array where 0

marks a KDA linear-attention layer and 1 a full MLA layer: 1 MLA layer in every 6. MLA uses a

compressed KV path (kv_lora_rank 512) with a plain wide query projection (q_lora_rank: null).

The blk.42 MTP layer is retained in full, including nextn.eh_proj, nextn.enorm,

nextn.hnorm and nextn.shared_head_norm, with the unfused attn_k_b / attn_v_b form that

the MTP path requires.

Measured throughput

AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), ROCm 7.2.4, 128 GB unified memory.

llama-cli, -dio -ngl 999 -st -c 2048 -n 512 --temp 0 --seed 1234, **3 repetitions per

config**, measured on an otherwise idle box.

| config | flags | generation (median) | runs |

|---|---|---:|---|

| no drafter | --spec-type none | 36.6 t/s | 36.6 / 36.5 / 36.6 |

| MTP n-max 3 | --spec-type draft-mtp --spec-draft-ngl 999 --spec-draft-n-max 3 | 41.4 t/s | 41.4 / 41.3 / 41.4 |

MTP is worth +13.1% on this checkpoint, with disjoint ranges.

All three Ling-3.0-flash base checkpoints measure the same no-drafter baseline to the decimal on

identical hardware and flags, which is the cross-check for this figure:

| checkpoint | no drafter | MTP n-max 3 | effect |

|---|---:|---:|---:|

| Ling-3.0-flash-base | 36.6 t/s | 42.3 t/s | +15.6% |

| Ling-3.0-flash-base-30T (this file) | 36.6 t/s | 41.4 t/s | +13.1% |

| Ling-3.0-flash-base-midtrain | 36.6 t/s | 43.0 t/s | +17.5% |

MTP is reliably positive across the whole Ling-3.0-flash base family. It is not reliable on

Ling-3.0-tiny, where the same measurement gives +7.5% / +5.1% / βˆ’4.4% across the three

checkpoints β€” the draft head is trained with the model, so its value belongs to the specific

(size, checkpoint) pair rather than to the architecture. Measure before enabling it.

Requirements

This file uses the ROCmFP4 tensor format and the bailing-hybrid architecture. It requires a build

of ROCmFPX that carries both. Stock llama.cpp will

not load it. Verify with strings libllama.so | grep bailing-hybrid β€” the architecture table lives

in the shared library, not in the thin CLI binary.

llama-cli -m Ling-3.0-flash-base-30T-Q4_0_ROCMFP4_COHERENT.gguf \
  -dio -ngl 999 -c 2048 -n 512 \
  --spec-type draft-mtp --spec-draft-ngl 999 --spec-draft-n-max 3 \
  -p "The history of mathematics begins in ancient times. One of the earliest known"

-dio (direct I/O) is recommended. At -ngl 999 the HIP backend copies offloaded tensors out of

file-backed pages into device allocations, so without direct I/O the source pages and the device

buffer are resident simultaneously β€” roughly twice the model size, which is tight on a 128 GB box.

Sample output

Continuation from "The history of mathematics begins in ancient times. One of the earliest known":

> mathematical texts is the Rhind Mathematical Papyrus, which dates back to around 1650 BCE. This ancient Egyptian document contains a collection of mathematical problems and solutions, covering topics such as arithmetic, geometry, and algebra. The Rhind Papyrus provides valuable insights into the mathematical knowledge and techniques of the time, showcasing the Egyptians' understanding of

Not measured

  • Perplexity is not published for this build. A 127 B model at this size exceeds a practical

evaluation budget on a single Strix Halo box. Quality evidence here is limited to the coherence

check above and the verified tensor-level audit.

  • Output determinism under MTP was not tested on this checkpoint. On the sibling

Ling-3.0-flash-base, MTP was found not to be output-deterministic at --temp 0 with a fixed

seed. Assume the same here unless you verify it.

  • n-max 5 was not swept on this checkpoint. n-max 3 is the published setting.

Provenance

Converted from inclusionAI/Ling-3.0-flash-base-30T at revision

4c7a67683c9e7ea9fad199b58e7b54cd835b63c2 to BF16 GGUF (938 tensors), then quantised to ftype 102 with

--output-tensor-type q6_K, then general.name set to Ling-3.0-flash-base-30T and the heads

re-audited on the finished file. Licence MIT, inherited from the base model.

Run kingjones777/Ling-3.0-flash-base-30T-ROCmFP4-COHERENT-GGUF with guIDE

Download guIDE β€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE β†’ Β· Browse 524k+ models Β· Compare models

Source: Hugging Face Β· Compare models