GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE โ†’
Model Intelligence Sheet

w-ahmad/LFM2.5-8B-A1B-GGUF-MoQ overview

๐Ÿš€ MoQ: Mixture of Quants image https://cdn uploads.huggingface.co/production/uploads/69ac7f5db2b3b515d77e2278/rPskixznY 6PiAd981s5S.png MoQ Mixture of Quants โ€ฆ

ggufMoQmixture-of-quantsGGUFQWENquantizationtext-generationenbase_model:LiquidAI/LFM2.5-8B-A1Bbase_model:quantized:LiquidAI/LFM2.5-8B-A1Blicense:mitendpoints_compatibleregion:usimatrixconversational

Runs locally from ~2.71 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
57,464
Likes
6
Pipeline
text-generation
Author

Repository Files & Downloads

10 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
MoQ-2.75.ggufGGUFGGUF2.71 GBDownload
MoQ-3.0.ggufGGUFGGUF2.95 GBDownload
MoQ-3.25.ggufGGUFGGUF3.20 GBDownload
MoQ-3.5.ggufGGUFGGUF3.46 GBDownload
MoQ-3.75.ggufGGUFGGUF3.69 GBDownload
MoQ-4.0.ggufGGUFGGUF3.91 GBDownload
MoQ-4.25.ggufGGUFGGUF4.16 GBDownload
MoQ-4.5.ggufGGUFGGUF4.40 GBDownload
MoQ-5.0.ggufGGUFGGUF4.93 GBDownload
bf16/LFM2.5-8B-A1B-BF16.ggufGGUFBF1615.78 GBDownload

Model Details

Model IDw-ahmad/LFM2.5-8B-A1B-GGUF-MoQ
Authorw-ahmad
Pipelinetext-generation
Licensemit
Base modelLiquidAI/LFM2.5-8B-A1B
Last modified2026-06-09T00:50:20.000Z

Model README

---

language:

  • en

library_name: gguf

tags:

  • MoQ
  • mixture-of-quants
  • GGUF
  • QWEN
  • quantization

base_model:

  • LiquidAI/LFM2.5-8B-A1B

license: mit

pipeline_tag: text-generation

---

๐Ÿš€ MoQ: Mixture of Quants

!image

>MoQ (Mixture of Quants) is a smart way to shrink AI models without losing their "brainpower." Unlike old methods that treat every part of the model the same, MoQ identifies the most important parts and keeps them high-quality, while heavily compressing the rest to save space.

The result? A model that punches significantly above its weight class.

Benjamin Marie evaluated MoQ GGUFs ("Mixture of Quants") against Unsloth Dynamic (UD) quants, focusing on low-bit versions below 4 bits on average โ€” the range where GGUF models typically struggle most.

Results: At similar bits-per-weight (Bpw), MoQ outperforms Unsloth Dynamic quants by ~10% on benchmarks, while also being roughly 2ร— more token-efficient on average.

"MoQ models are much better than UD quants on benchmarks, and they are also more token-efficient."

Comparison

Here is the comparison between MoQ and Unsloth dynamic quants for LFM 2.5 8BA1B.

MoQ perform better i guess .

!image

!image

Thanks to benjamin Marie for his evals

!image

| Folder Link | BPW | Total Size | * |

| :--- | :---: | :---: | :--- |

| ๐Ÿ“‚ Root | 2.75 | 2.91 GB

| ๐Ÿ“‚ Root | 3.0 | 3.17 GB

| ๐Ÿ“‚ Root | 3.25 | 3.43 GB

| ๐Ÿ“‚ Root | 3.5 | 3.71 GB

| ๐Ÿ“‚ Root | 3.75 | 3.96 GB

| ๐Ÿ“‚ Root | 4.0 | 4.20 GB

| ๐Ÿ“‚ Root | 4.25 | 4.46 GB

| ๐Ÿ“‚ Root | 4.5 | 4.72 GB

| ๐Ÿ“‚ Root | 5.0 | 5.30 GB

| ๐Ÿ“‚ bf16 | 16.0 | 16.95 GB

๐Ÿง  The MoQ Edge

MoQ optimizes the architecture for the Pareto frontier of memory and performance.

  • Dynamic Bitrate Allocation: No more "one-size-fits-all." MoQ assigns precision where it actually matters.
  • Cognitive Preservation: Massive VRAM savings with near-zero degradation in logic and coherence.
  • Next-Gen Efficiency: Fits "Large" model intelligence into "Small" model hardware.

##

x : https://x.com/WaleedAhmad1a10

If MoQ does not perform well, email me :

waleedahmad.1a10@gmail.com

๐Ÿ›  Usage & Deployment.

./llama-cli -m Qwen3.5-9B-MoQ-4.0.gguf -p "The future of efficient AI is..."

Run w-ahmad/LFM2.5-8B-A1B-GGUF-MoQ with guIDE

Download guIDE โ€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE โ†’ ยท Browse 524k+ models ยท Compare models

Source: Hugging Face ยท Compare models