GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Nichonauta/LFM2.5-230M-ToMoE-GGUF overview

LFM2.5 230M ToMoE GGUF GGUF quantizations of Nichonauta/LFM2.5 230M ToMoE https://huggingface.co/Nichonauta/LFM2.5 230M ToMoE — the ToMoE Mixture of Experts co…

gguflfm2mixture-of-expertstomoellama.cpptext-generationenbase_model:Nichonauta/LFM2.5-230M-ToMoEbase_model:quantized:Nichonauta/LFM2.5-230M-ToMoElicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~146.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LFM2.5-230M-ToMoE-F16.ggufGGUFF16440.5 MBDownload
LFM2.5-230M-ToMoE-Q4_K_M.ggufGGUFQ4_K_M146.3 MBDownload
LFM2.5-230M-ToMoE-Q8_0.ggufGGUFQ8_0235.2 MBDownload

Model Details

Model IDNichonauta/LFM2.5-230M-ToMoE-GGUF
AuthorNichonauta
Pipelinetext-generation
Licenseother
Base modelNichonauta/LFM2.5-230M-ToMoE
Last modified2026-08-26T01:07:59.000Z

Model README

---

license: other

license_name: lfm-open-license-v1.0

license_link: https://huggingface.co/LiquidAI/LFM2.5-230M/blob/main/LICENSE

language:

  • en

base_model: Nichonauta/LFM2.5-230M-ToMoE

tags:

  • lfm2
  • mixture-of-experts
  • tomoe
  • gguf
  • llama.cpp

pipeline_tag: text-generation

---

LFM2.5-230M-ToMoE-GGUF

GGUF quantizations of Nichonauta/LFM2.5-230M-ToMoE — the ToMoE Mixture-of-Experts conversion of LiquidAI/LFM2.5-230M.

Files

| File | Quantization | Size | BPW |

|---|---|---|---|

| LFM2.5-230M-ToMoE-F16.gguf | F16 | 438 MB | 16.0 |

| LFM2.5-230M-ToMoE-Q8_0.gguf | Q8_0 | 233 MB | 8.51 |

| LFM2.5-230M-ToMoE-Q4_K_M.gguf | Q4_K_M | 144 MB | 5.26 |

Important note about the conversion

The ToMoE channel-MoE experts and the per-token dynamic V masks have no native representation in llama.cpp's LFM2 architecture. To produce a runnable GGUF, the pruned model was rebuilt as a dense-equivalent:

  • The cut conv layers were restored to their full width using the original dense weights (the conv weights were never modified by ToMoE — only masked at runtime).
  • The MLP union cut (~99.8% of the FFN) was restored to the full FFN.
  • The attention Q/K were already full-width in the pruned model.

As a result, the GGUF behaves approximately like the dense base model rather than the pruned MoE. For the actual MoE behavior, use the safetensors version with trust_remote_code.

Usage (llama.cpp, CUDA)

llama-server.exe ^
  --model LFM2.5-230M-ToMoE-Q4_K_M.gguf ^
  --gpu-layers all ^
  --ctx-size 32768 ^
  --alias LFM2.5-230M-ToMoE

Or with llama-cli:

llama-cli.exe -m LFM2.5-230M-ToMoE-Q4_K_M.gguf -ngl 99 -p "The capital of France is" -n 25

Measured on an RTX 3060 (CUDA 13.3 build, llama.cpp b10603): ~540–550 t/s generation for Q4_K_M, ~430–450 t/s for Q8_0.

Base model

License

Derivative of LiquidAI/LFM2.5-230M — released under the LFM Open License v1.0 (see LICENSE).

Run Nichonauta/LFM2.5-230M-ToMoE-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models