Q1ngMang/Ling-3.0-tiny-sub3bit-PPLp10-GGUF overview
Ling 3.0 tiny sub3bit PPLp10 GGUF 我的天那,钠是结晶的。<br 这是我见过最小巧最有用的MoE模型。<br 所以我想随意的量化它,并为PPL跑分设计。(其实是我想内置到 TranslatorMinecraft https://github.com/lingxingmiao/Trans…
Runs locally from ~2.75 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | Q1ngMang/Ling-3.0-tiny-sub3bit-PPLp10-GGUF |
|---|---|
| Author | Q1ngMang |
| Pipeline | text-generation |
| License | mit |
| Base model | inclusionAI/Ling-3.0-tiny |
| Last modified | 2026-08-30T15:01:25.000Z |
Model README
---
license: mit
base_model:
- inclusionAI/Ling-3.0-tiny
pipeline_tag: text-generation
library_name: llama.cpp
tags:
- gguf
- bailingmoe3
- mixture-of-experts
- conversational
---
Ling-3.0-tiny-sub3bit-PPLp10-GGUF
我的天那,钠是结晶的。<br>
这是我见过最小巧最有用的MoE模型。<br>
所以我想随意的量化它,并为PPL跑分设计。(其实是我想内置到TranslatorMinecraft,因为它是MIT的。)<br>
我会在这个仓库上上传小于3bit且bartowski/Ling-3.0-tiny-calibration-v6.txt验证PPL小于+10%的模型。<br>
校准文件bartowski/Ling-3.0-tiny-imatrix.gguf<br>
| 项 | 00 | 01 | bartowski/Ling-3.0-tiny-bf16.gguf |
| - | - | - | - |
| BPW | 2.99 | 2.98 | 16 |
| PPL | 5.0691±0.03523 | 5.0480±0.03498 | 4.6431±0.03271 |
| ↑+% | 9.1749% | 8.7204% | 0% |
| 大小 | 2812.01MiB | 2807.11MiB | 15065.15MiB |
| 经验 | 输入相关的不用IQ更强,输出得用IQ。 | 继左边 | 空气与Air的混合物 |
| 观后感 | NAVI回家吧,paiN粘的要死。 | 呜呜呜我补药开学 | 网速100Mbps |
你懂 01 比 00 小 4.9MiB 是什么感觉吗?换算到 HBM4 价值整整¥1.75。
Ling-3.0-tiny-sub3bit-PPLp10-GGUF
Markdown translation model: DeepSeek V4 Flash 0731.<br>
Oh my god, sodium is crystalline.<br>
This is the smallest and most useful MoE model I have ever seen.<br>
So I decided to casually quantize it and design it for PPL benchmarking. (Actually, I want to embed it into TranslatorMinecraft, because it is MIT-licensed.)<br>
I will upload sub-3-bit models on this repository that have a PPL increase of less than +10%, verified by bartowski/Ling-3.0-tiny-calibration-v6.txt.<br>
Calibration file: bartowski/Ling-3.0-tiny-imatrix.gguf<br>
| Item | 00 | 01 | bartowski/Ling-3.0-tiny-bf16.gguf |
| - | - | - | - |
| BPW | 2.99 | 2.98 | 16 |
| PPL | 5.0691±0.03523 | 5.0480±0.03498 | 4.6431±0.03271 |
| ↑+% | 9.1749% | 8.7204% | 0% |
| Size | 2812.01 MiB | 2807.11 MiB | 15065.15 MiB |
| Experience | For input-related tasks, non-IQ is stronger; for output, use IQ. | Continuing from the left | A mixture of air and Air |
| Afterthoughts | NAVI, go home; paiN is too sticky/clingy. | Boo-hoo, I don't want school to start! | Internet speed: 100 Mbps |
Can you imagine what it means that 01 is 4.9 MiB less than 00? In HBM4 cost terms, that's a whole $0.25.
Run Q1ngMang/Ling-3.0-tiny-sub3bit-PPLp10-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models