AtomicChat/Ling-3.0-flash-VL-GGUF overview
How to Run Ling 3.0 Flash VL Locally <p style="margin top: 0; margin bottom: 0;" <em Built from InclusionAI's original weights using Atomic Chat's existing Lin…
Runs locally from ~1.62 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Ling-3.0-flash-VL-AD-Q4_K_M-00001-of-00006.gguf | GGUF | Q4_K_M | 14.55 GB | Download |
| Ling-3.0-flash-VL-AD-Q4_K_M-00002-of-00006.gguf | GGUF | Q4_K_M | 13.65 GB | Download |
| Ling-3.0-flash-VL-AD-Q4_K_M-00003-of-00006.gguf | GGUF | Q4_K_M | 13.40 GB | Download |
| Ling-3.0-flash-VL-AD-Q4_K_M-00004-of-00006.gguf | GGUF | Q4_K_M | 13.74 GB | Download |
| Ling-3.0-flash-VL-AD-Q4_K_M-00005-of-00006.gguf | GGUF | Q4_K_M | 14.39 GB | Download |
| Ling-3.0-flash-VL-AD-Q4_K_M-00006-of-00006.gguf | GGUF | Q4_K_M | 4.12 GB | Download |
| Ling-3.0-flash-VL-AD-Q5_K_M-00001-of-00006.gguf | GGUF | Q5_K_M | 16.22 GB | Download |
| Ling-3.0-flash-VL-AD-Q5_K_M-00002-of-00006.gguf | GGUF | Q5_K_M | 15.53 GB | Download |
| Ling-3.0-flash-VL-AD-Q5_K_M-00003-of-00006.gguf | GGUF | Q5_K_M | 15.16 GB | Download |
| Ling-3.0-flash-VL-AD-Q5_K_M-00004-of-00006.gguf | GGUF | Q5_K_M | 15.50 GB | Download |
| Ling-3.0-flash-VL-AD-Q5_K_M-00005-of-00006.gguf | GGUF | Q5_K_M | 16.28 GB | Download |
| Ling-3.0-flash-VL-AD-Q5_K_M-00006-of-00006.gguf | GGUF | Q5_K_M | 4.61 GB | Download |
| Ling-3.0-flash-VL-AD-Q6_K-00001-of-00006.gguf | GGUF | Q6_K | 19.57 GB | Download |
| Ling-3.0-flash-VL-AD-Q6_K-00002-of-00006.gguf | GGUF | Q6_K | 18.39 GB | Download |
| Ling-3.0-flash-VL-AD-Q6_K-00003-of-00006.gguf | GGUF | Q6_K | 18.02 GB | Download |
| Ling-3.0-flash-VL-AD-Q6_K-00004-of-00006.gguf | GGUF | Q6_K | 18.36 GB | Download |
| Ling-3.0-flash-VL-AD-Q6_K-00005-of-00006.gguf | GGUF | Q6_K | 19.45 GB | Download |
| Ling-3.0-flash-VL-AD-Q6_K-00006-of-00006.gguf | GGUF | Q6_K | 5.98 GB | Download |
| Ling-3.0-flash-VL-AD-Q8_0-00001-of-00006.gguf | GGUF | Q8_0 | 23.21 GB | Download |
| Ling-3.0-flash-VL-AD-Q8_0-00002-of-00006.gguf | GGUF | Q8_0 | 23.61 GB | Download |
| Ling-3.0-flash-VL-AD-Q8_0-00003-of-00006.gguf | GGUF | Q8_0 | 23.25 GB | Download |
| Ling-3.0-flash-VL-AD-Q8_0-00004-of-00006.gguf | GGUF | Q8_0 | 23.58 GB | Download |
| Ling-3.0-flash-VL-AD-Q8_0-00005-of-00006.gguf | GGUF | Q8_0 | 24.00 GB | Download |
| Ling-3.0-flash-VL-AD-Q8_0-00006-of-00006.gguf | GGUF | Q8_0 | 5.98 GB | Download |
| mmproj-Ling-3.0-flash-VL-F32.gguf | GGUF | F32 | 1.62 GB | Download |
Model Details
| Model ID | AtomicChat/Ling-3.0-flash-VL-GGUF |
|---|---|
| Author | AtomicChat |
| Pipeline | image-text-to-text |
| License | mit |
| Base model | inclusionAI/Ling-3.0-flash-VL |
| Last modified | 2026-09-08T16:56:14.000Z |
Model README
---
license: mit
base_model: inclusionAI/Ling-3.0-flash-VL
library_name: llama.cpp
pipeline_tag: image-text-to-text
tags:
- gguf
- experimental
---
How to Run Ling 3.0 Flash VL Locally
<p style="margin-top: 0; margin-bottom: 0;">
<em>Built from InclusionAI's original weights using Atomic Chat's existing Ling importance matrix. The <a href="https://huggingface.co/datasets/AtomicChat/calib-corpora">calibration corpora</a> behind our builds are public.</em>
</p>
<div style="display: flex; gap: 8px; align-items: center; margin-top: 10px; margin-bottom: 10px;">
<a href="https://atomic.chat/?utm_source=huggingface&utm_medium=referral&utm_campaign=hf_ling_3_0_flash_vl&utm_content=btn_atomic"><img src="https://huggingface.co/AtomicChat/Ling-3.0-flash-VL-GGUF/resolve/main/btn_atomic.png" width="162" alt="Atomic Chat"></a>
<a href="https://discord.gg/8wGSsvmg4V"><img src="https://huggingface.co/AtomicChat/Ling-3.0-flash-VL-GGUF/resolve/main/btn_discord.png" width="119" alt="Discord"></a>
<a href="https://github.com/AtomicBot-ai/Atomic-Chat"><img src="https://huggingface.co/AtomicChat/Ling-3.0-flash-VL-GGUF/resolve/main/btn_github.png" width="115" alt="GitHub"></a>
</div>
<ul style="margin: 0 0 12px 0;">
<li>Ling 3.0 Flash VL is InclusionAI's vision-language model for text and image inputs.</li>
<li>Choose AD-Q4_K_M, AD-Q5_K_M, AD-Q6_K, or AD-Q8_0. Image inputs use the shared F32 vision projector.</li>
<li>These builds are an experimental preview. Use the accompanying runtime; full quality validation is still in progress.</li>
</ul>
<hr style="margin: 0 0 16px 0;">
Prepared from the original inclusionAI checkpoint, revision 869591498e8dbb41d4d96e3e2a5b428a2f70eb1e.
These are the Atomic AD layouts, made on CPU using the existing calibration iMatrix. The full nine-variant text comparison and CPU speed measurements are complete; see FINAL_REPORT.md. This repository is an experimental working preview, not a completed benchmark release.
Uploads arrive progressively. Check UPLOAD_STATUS.json before downloading a variant. Each language variant requires all six GGUF shards in the same directory; select shard 00001 when loading.
| Variant | Complete language files, decimal GB |
|---|---:|
| AD-Q4_K_M | 79.30 |
| AD-Q5_K_M | 89.44 |
| AD-Q6_K | 107.14 |
| AD-Q8_0 | 132.73 |
Vision requires the shared F32 mmproj, an additional 1.74 GB. File size is not a RAM requirement estimate.
Runtime requirement
Use the accompanying private runtime patch and added source files, based on AtomicBot-ai/atomic-llama-cpp-turboquant@cd560939087c95b93a1f30a95603d6b079436952. Stock llama.cpp and the released Atomic Chat app have not been validated for these artifacts. See RUNTIME.md.
Checks completed
All four variants: 917 tensor types/shapes verified, six-shard integrity checked, complete SHA-256 manifests, and all 382 protected F32 tensor payloads unchanged from BF16. Text and a spatial image passed on all four. Q4/Q6/Q8 additionally passed the OCR and object-count smoke cases. These are functional smoke checks, not a comprehensive vision benchmark.
Full text quality comparisons use the historical held-out 92 × 4096 protocol and a fresh BF16 reference from this checkpoint. Full text results are available in FINAL_REPORT.md. Video and maximum context are not validated.
The existing iMatrix comes from AtomicChat/Ling-3.0-flash-GGUF@253738fe190c15f329001f263f355fc1562bbe7c, SHA-256 7d3c0ebe9eb235cc08e0b7c91886f5c422772ea95eebbf0f53b0974a1c040991. Its 573 entries match the new language tensor dimensions. Four routed experts have no observations in that matrix; uniform importance was used for those entries.
manifest.json and SHA256SUMS describe the complete intended set. Actual upload completion is recorded separately in UPLOAD_STATUS.json.
Completed measurements
Full report · KLD chart · Metrics CSV.
The pilot results are separate from the full 92-block comparison. Completion is not a comprehensive vision or agentic quality certification.
Run AtomicChat/Ling-3.0-flash-VL-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models