GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

bloomer010/Ling-3.0-flash-VL-GGUF overview

Ling 3.0 flash VL GGUF GGUF conversions of inclusionAI/Ling 3.0 flash VL https://huggingface.co/inclusionAI/Ling 3.0 flash VL Built upon Ling 3.0 flash, it bri…

llama.cppggufbailingmoe3mixture-of-expertsvisionvideoconversationalimage-text-to-textbase_model:inclusionAI/Ling-3.0-flash-VLbase_model:quantized:inclusionAI/Ling-3.0-flash-VLlicense:mitendpoints_compatibleregion:us

Runs locally from ~513.8 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
6,863
Likes
5
Pipeline
image-text-to-text

Repository Files & Downloads

22 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-flash-DSpark-BF16.ggufGGUFBF161.80 GBDownload
Ling-3.0-flash-DSpark-Q2_K.ggufGGUFQ2_K513.8 MBDownload
Ling-3.0-flash-DSpark-Q4_K_M.ggufGGUFQ4_K_M633.8 MBDownload
Ling-3.0-flash-DSpark-Q6_K.ggufGGUFQ6_K758.3 MBDownload
Ling-3.0-flash-VL-BF16.ggufGGUFBF16231.85 GBDownload
Ling-3.0-flash-VL-IQ2_M.ggufGGUFIQ2_M37.83 GBDownload
Ling-3.0-flash-VL-IQ2_XS.ggufGGUFIQ2_XS34.02 GBDownload
Ling-3.0-flash-VL-MXFP4_MOE.ggufGGUFGGUF63.49 GBDownload
Ling-3.0-flash-VL-Q3_K_M.ggufGGUFQ3_K_M55.22 GBDownload
Ling-3.0-flash-VL-Q4_K_M.ggufGGUFQ4_K_M70.11 GBDownload
Ling-3.0-flash-VL-Q4_K_S.ggufGGUFQ4_K_S65.82 GBDownload
Ling-3.0-flash-VL-Q5_K_M.ggufGGUFQ5_K_M82.28 GBDownload
Ling-3.0-flash-VL-Q6_K.ggufGGUFQ6_K95.22 GBDownload
Ling-3.0-flash-VL-Q8_0.ggufGGUFQ8_0123.27 GBDownload
Ling-3.0-flash-VL-UD-Q2_K_XL.ggufGGUFQ2_K_XL39.03 GBDownload
Ling-3.0-flash-VL-UD-Q3_K_XL.ggufGGUFQ3_K_XL56.46 GBDownload
Ling-3.0-flash-VL-UD-Q4_K_XL.ggufGGUFQ4_K_XL76.31 GBDownload
Ling-3.0-flash-VL-UD-Q5_K_XL.ggufGGUFQ5_K_XL85.97 GBDownload
Ling-3.0-flash-VL-UD-Q6_K_XL.ggufGGUFQ6_K_XL96.02 GBDownload
Ling-3.0-flash-VL-UD-Q6_K_XXL.ggufGGUFQ6_K_XXL105.05 GBDownload
Ling-3.0-flash-VL-UD-Q8_K_XL.ggufGGUFQ8_K_XL160.61 GBDownload
mmproj-model-f16.ggufGGUFF16834.1 MBDownload

Model Details

Model IDbloomer010/Ling-3.0-flash-VL-GGUF
Authorbloomer010
Pipelineimage-text-to-text
Licensemit
Base modelinclusionAI/Ling-3.0-flash-VL
Last modified2026-09-26T04:11:58.000Z

Model README

---

license: mit

base_model:

  • inclusionAI/Ling-3.0-flash-VL

pipeline_tag: image-text-to-text

library_name: llama.cpp

tags:

  • gguf
  • bailingmoe3
  • mixture-of-experts
  • vision
  • video
  • conversational

---

Ling-3.0-flash-VL GGUF

GGUF conversions of inclusionAI/Ling-3.0-flash-VL

> Built upon Ling-3.0-flash, it brings visual information into the complete process of understanding, reasoning, acting, and verification—advancing beyond image and video perception to solving real-world tasks through vision. With 124B total parameters, only 5.5B activated parameters per token, support for image and video inputs, and a context window of up to 256K tokens, Ling-3.0-flash-VL delivers powerful multimodal reasoning and agentic capabilities with exceptional efficiency.

Every text quant requires the bundled 875 MB mmproj-model-f16.gguf file for vision input.<br>

Text-only chat works without it.

🦙🦙 llama.cpp 🦙🦙

🎉 Now supported in stock llama.cpp 🎉 <br>

Merged 2026-09-24 in #29151 (commit f830688e9).<br>

Any build b11190 or newer loads these files as-is.

📝 Note: the GGUFs in this repo were re-published on 2026-09-22 under the final

bailingmoe3 architecture name. Files downloaded before that date, or builds

of the obsolete temporary ling3-vl branch are incompatible. You will need to re-download

to use llama.cpp on build b11190 or newer.

To run with llama-server:

llama-server \
  -m Ling-3.0-flash-VL-Q4_K_M.gguf \
  --mmproj mmproj-model-f16.gguf \
  --jinja

Quant Sizing

Generally... <BR>

Larger files = More precision.<BR>

Smaller files = More compression = More slop and misbehavin'.

Weights and context share your memory, so be sure to leave headroom.

| your memory | file | size |

| --- | --- | ---: |

| 256 GB+ | BF16 | 249 GB |

| 178 GB+ | UD-Q8_K_XL | 172.5 GB |

| 136 GB+ | Q8_0 | 132 GB |

| 116 GB+ | UD-Q6_K_XXL | 112.8 GB |

| 128 GB | UD-Q6_K_XL | 103 GB |

| 104 GB+ | Q6_K | 102 GB |

| 94 GB+ | UD-Q5_K_XL | 92.3 GB |

| 90 GB+ | Q5_K_M | 88.3 GB |

| 84 GB+ | UD-Q4_K_XL | 81.9 GB |

| 76 GB+ | Q4_K_M | 75.3 GB |

| 72 GB+ | Q4_K_S | 70.7 GB |

| 62 GB+ | UD-Q3_K_XL | 60.6 GB |

| 60 GB+ | Q3_K_M | 59.3 GB |

| 44 GB+ | UD-Q2_K_XL | 41.9 GB |

| 42 GB+ | IQ2_M | 40.6 GB |

| 38 GB+ | IQ2_XS | 36.5 GB |

| (vision, required for images/video) | mmproj-model-f16.gguf | 0.87 GB |

With less VRAM than the file size, keep the experts on CPU and the rest on GPU, e.g.:

llama-server \
  -m Ling-3.0-flash-VL-Q4_K_M.gguf \
  --mmproj mmproj-model-f16.gguf \
  -ngl 99 -ot "ffn_.*_exps\.weight=CPU" -c 32768 \
  --jinja

Usage

Recommended sampling from the source model card: temperature 0.6, top_p 0.95, top_k 20.

Thinking mode is on by default; disable per request with

"chat_template_kwargs": {"enable_thinking": false}.

Images

./build/bin/llama-server \
  -m Ling-3.0-flash-VL-Q4_K_M.gguf \
  --mmproj mmproj-model-f16.gguf \
  -c 131072 \
  -ngl auto \
  --flash-attn auto \
  --temp 0.6 --top-p 0.95 --top-k 20 \
  --jinja

Then attach an image in the web UI, or via the API:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Describe this image."},
        {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
      ]
    }]
  }'

Video

Video input uses the same chat API with video_url content parts. Frames are sampled and encoded

by the same vision tower.

Unlike the text-only Ling-3.0-flash GGUFs, these files contain no MTP/NextN block: the VL release

does not ship one. Speculative drafting via --spec-type draft-mtp is not available for VL.

Long context (256K)

The GGUFs declare a native 131,072-token context. The advertised 256K window is reached with YaRN at factor 2, mirroring the upstream Ling-3.0-flash-VL recipe (yarn, factor 2.0, original context 131072):

llama-server \
  -m Ling-3.0-flash-VL-Q4_K_M.gguf \
  --mmproj mmproj-model-f16.gguf \
  -c 262144 \
  --rope-scaling yarn --rope-scale 2 --yarn-orig-ctx 131072 \
  --jinja

The attention KV cache scales with -c, doubling KV memory versus the native 131,072. Long-context quality at 256K was not validated here, the flags simply mirror the upstream recommendation.

Speculative decoding (DSpark)

The Ling-3.0-flash DSpark draft heads are compatible with the VL model: same vocabulary, and

acceptance on VL is at least as good as on the text-only model the draft was trained for.

Measured with llama-server (VL Q6_K target, --spec-type draft-dspark --spec-draft-n-max 8,

32K context, 24 requests):

| config | decode speed |

| --- | ---: |

| Q6_K | 26.8 tok/s |

| Q6_K + DSpark Q4_K_M | 43.6 tok/s (1.63x) |

Draft acceptance on VL Q6_K: 0.32 (Q4_K_M draft), 0.30 (Q2_K draft). The same Q4_K_M draft

measures 0.26 against text-only Ling-3.0-flash.

llama-server \
  -m Ling-3.0-flash-VL-Q6_K.gguf \
  -md Ling-3.0-flash-DSpark-Q4_K_M.gguf \
  --spec-type draft-dspark --spec-draft-n-max 8 \
  -ngl 99 -ngld 99 \
  --mmproj mmproj-model-f16.gguf \
  --jinja

The DSpark draft's attention does not use flash attention, so its compute buffer

grows linearly with context length. It also carries its own KV cache; keep it at

f16, quantizing it with -ctkd q4_0 -ctvd q4_0 was measured to cut decode speed

by a further ~40%.

Benchmark your own stack before adopting the draft: the 1.63x above is from a

bare benchmark harness (single role, 32K context). In a full serving stack

(7 mixed GPUs, vision encoder loaded, q4_0 target KV cache, 48K context) the same

draft measured 24.1 tok/s versus 38.1 tok/s with no draft at all.

Additional MoE Information

MoE placement can be adjusted for available VRAM with -ncmoe N.

Native context is 128K; 256K is available via YaRN (see Long context above).

Conversion and Quantization

Taken directly from the released inclusionAI/Ling-3.0-flash-VL BF16 safetensors.

Conversion-specific tensor transformations match the text-only Ling-3.0-flash conversions:

  • A_log stored as exp(A_log)
  • MLA kv_b_proj split into separate K and V tensors, with the K tensor transposed
  • KDA convolution weights reshaped for llama.cpp
  • Per-expert tensors stacked into GGUF expert tensors
  • KDA and MLA g_proj tensors mapped separately

Vision tower and projector tensors live in the separate mmproj GGUF: Conv3D patch embedding,

learned position embeddings, 27 attention blocks, a norm-only merger, and the two-layer projector.

Norms, routing tensors, expert routing bias, KDA state scalars, dt_bias, and convolution weights

remain F32.

Notes

The text GGUF contains 42 blocks:

  • 35 KDA layers
  • 7 gated MLA layers at zero-based indices 5, 11, 17, 23, 29, 35, and 41

(No MTP/NextN block, unlike the text-only flash GGUFs.)

The first two layers use dense FFNs. The remaining layers use 512 routed experts with top-8

selection plus one shared expert. Routing uses sigmoid scoring, expert bias, eight expert groups,

and four selected groups.

Position encoding is M-RoPE with sections [8, 12, 12], shared between text and vision positions.

Validation Completed

  • BF16 architecture load and tensor round-trip (test-llama-archs, MoE fixture)
  • mmproj GGUF round-trip: 334 tensors, ling3vl_merger projector
  • End-to-end image and video inference on llama-server (Q4_K_M + mmproj)

Build

# Ling 3.0 VL support is in stock master (b11190+):
git clone https://github.com/ggml-org/llama.cpp.git

cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j --target llama-cli llama-server

---

Run bloomer010/Ling-3.0-flash-VL-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models