GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

vincespeed/Ling-3.0-tiny-APEX-GGUF overview

Ling 3.0 Tiny — Apex Quant GGUF Models This repository contains 3 quantized GGUF profiles of the inclusionAI/Ling 3.0 tiny model, produced using Apex Quant tec…

ggufmoebailingmoe3text-generationbase_model:inclusionAI/Ling-3.0-tinybase_model:quantized:inclusionAI/Ling-3.0-tinylicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~3.71 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
179
Likes
1
Pipeline
text-generation

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ling-3.0-tiny-i-balanced.ggufGGUFGGUF5.54 GBDownload
Ling-3.0-tiny-i-compact.ggufGGUFGGUF3.71 GBDownload
Ling-3.0-tiny-i-quality.ggufGGUFGGUF5.37 GBDownload

Model Details

Model IDvincespeed/Ling-3.0-tiny-APEX-GGUF
Authorvincespeed
Pipelinetext-generation
Licensemit
Base modelinclusionAI/Ling-3.0-tiny
Last modified2026-08-23T18:24:10.000Z

Model README

---

license: mit

base_model: inclusionAI/Ling-3.0-tiny

base_model_relation: quantized

tags:

  • gguf
  • moe
  • bailingmoe3

pipeline_tag: text-generation

---

Ling-3.0 Tiny — Apex Quant GGUF Models

This repository contains 3 quantized GGUF profiles of the inclusionAI/Ling-3.0-tiny model, produced using Apex-Quant technology.

📦 Model Profile Summary

| Profile | Size | BPW | Use Case |

|---------|------|-----|----------|

| i-quality | 5.4 GB | 5.83 | Highest quality, production environments |

| i-balanced | 5.6 GB | 6.02 | Balanced quality and performance |

| i-compact | 3.7 GB | 4.03 | Compact deployment, low RAM |

> BPW = Bits Per Weight. Higher value = better quality.

📁 File Structure

models/
├── Ling-3.0-tiny-i-quality.gguf   # 5.4 GB — Highest quality
├── Ling-3.0-tiny-i-balanced.gguf  # 5.6 GB — Balanced
└── Ling-3.0-tiny-i-compact.gguf   # 3.7 GB — Compact

🔗 Source Model

These models were created based on the inclusionAI/Ling-3.0-tiny model from HuggingFace.

  • Model Page: https://huggingface.co/inclusionAI/Ling-3.0-tiny
  • Architecture: BailingMoeV3ForCausalLM (Mixture-of-Experts)
  • Parameter Count: 128×1.0B (128 experts, each with 1B parameters)
  • Context Length: 131,072 tokens
  • Vocabulary: 157,184 tokens
  • License: MIT

🛠️ Technology

These quantized models were produced using Apex-Quant technology.

  • Apex-Quant: MoE-aware mixed-precision quantization
  • Infrastructure: llama.cpp (llama-quantize)
  • Quantize Script: apex-quant/scripts/quantize.sh

🙏 Acknowledgments

📋 Technical Details

Architecture Information

  • Architecture: bailingmoe3
  • Block Count: 24 layers
  • Expert Count: 128 experts
  • Expert Used Count: 8 experts/token
  • Expert Group Count: 8
  • Expert Group Used Count: 4
  • Expert Gating Function: Top-K (k=8)
  • Hidden Size: 1,536
  • Feed Forward Size: 4,608
  • Attention Heads: 16
  • Attention Head Count KV: [0, 0, 0, 1, ...] (grouped query attention)
  • Rope Frequency Base: 6,000,000
  • Layer Norm Epsilon: 1e-6

Quantize Profile Details

i-quality (Q6_K/Q5_K/IQ4_XS)

  • Expert FFN: Q6_K / Q5_K / IQ4_XS (mixed)
  • Shared FFN: Q8_0
  • Attention: Q6_K
  • BPW: 5.83
  • File Size: 5.4 GB

i-balanced (Q6_K/Q5_K)

  • Expert FFN: Q6_K / Q5_K (mixed)
  • Shared FFN: Q8_0
  • Attention: Q6_K
  • BPW: 6.02
  • File Size: 5.6 GB

i-compact (Q4_K/Q3_K)

  • Expert FFN: Q4_K / Q3_K (mixed)
  • Shared FFN: Q6_K
  • Attention: Q4_K
  • BPW: 4.03
  • File Size: 3.7 GB

💻 Usage

With llama.cpp

# Run with i-quality profile
./main -m models/Ling-3.0-tiny-i-quality.gguf -n 128 -p "Hello, how are you?"

# Run with i-compact profile
./main -m models/Ling-3.0-tiny-i-compact.gguf -n 128 -p "Hello, how are you?"

With Ollama

# Create Dockerfile or Ollamafile
FROM llama.cpp
COPY models/Ling-3.0-tiny-i-quality.gguf /model.gguf

With Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="models/Ling-3.0-tiny-i-quality.gguf",
    n_ctx=4096,
    n_threads=8
)

output = llm(
    "Hello, how are you?",
    max_tokens=128
)
print(output["choices"][0]["text"])

📊 Model Comparison

| Criterion | i-quality | i-balanced | i-compact |

|-----------|-----------|------------|-----------|

| Quality | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ |

| Speed | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ |

| RAM | High | Medium | Low |

| Size | 5.4 GB | 5.6 GB | 3.7 GB |

| BPW | 5.83 | 6.02 | 4.03 |

📝 Notes

  • All models are in GGUF v3 format.
  • The BailingMoeV3 architecture uses Mixture-of-Experts (MoE) technology.
  • The model uses grouped query attention (GQA) and rope positional embeddings.
  • The i-mini profile cannot be quantized without imatrix. ~100-200 inference samples must be run on the model to generate the importance matrix.

📄 License

The original model is distributed under the MIT license. The quantized models are shared under the same license.

🔗 Related Links

  • Original Model: https://huggingface.co/inclusionAI/Ling-3.0-tiny
  • Apex-Quant: https://github.com/localai-org/apex-quant
  • llama.cpp: https://github.com/ggerganov/llama.cpp
  • GGUF Format: https://github.com/ggerganov/ggml/blob/master/docs/gguf.md

---

Note: These models are quantized for local use. Check the original model's license for commercial use.

Run vincespeed/Ling-3.0-tiny-APEX-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models