Nichonauta/LFM2.5-230M-ToMoE-GGUF overview
LFM2.5 230M ToMoE GGUF GGUF quantizations of Nichonauta/LFM2.5 230M ToMoE https://huggingface.co/Nichonauta/LFM2.5 230M ToMoE — the ToMoE Mixture of Experts co…
Runs locally from ~146.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | Nichonauta/LFM2.5-230M-ToMoE-GGUF |
|---|---|
| Author | Nichonauta |
| Pipeline | text-generation |
| License | other |
| Base model | Nichonauta/LFM2.5-230M-ToMoE |
| Last modified | 2026-08-26T01:07:59.000Z |
Model README
---
license: other
license_name: lfm-open-license-v1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-230M/blob/main/LICENSE
language:
- en
base_model: Nichonauta/LFM2.5-230M-ToMoE
tags:
- lfm2
- mixture-of-experts
- tomoe
- gguf
- llama.cpp
pipeline_tag: text-generation
---
LFM2.5-230M-ToMoE-GGUF
GGUF quantizations of Nichonauta/LFM2.5-230M-ToMoE — the ToMoE Mixture-of-Experts conversion of LiquidAI/LFM2.5-230M.
Files
| File | Quantization | Size | BPW |
|---|---|---|---|
| LFM2.5-230M-ToMoE-F16.gguf | F16 | 438 MB | 16.0 |
| LFM2.5-230M-ToMoE-Q8_0.gguf | Q8_0 | 233 MB | 8.51 |
| LFM2.5-230M-ToMoE-Q4_K_M.gguf | Q4_K_M | 144 MB | 5.26 |
Important note about the conversion
The ToMoE channel-MoE experts and the per-token dynamic V masks have no native representation in llama.cpp's LFM2 architecture. To produce a runnable GGUF, the pruned model was rebuilt as a dense-equivalent:
- The cut conv layers were restored to their full width using the original dense weights (the conv weights were never modified by ToMoE — only masked at runtime).
- The MLP union cut (~99.8% of the FFN) was restored to the full FFN.
- The attention Q/K were already full-width in the pruned model.
As a result, the GGUF behaves approximately like the dense base model rather than the pruned MoE. For the actual MoE behavior, use the safetensors version with trust_remote_code.
Usage (llama.cpp, CUDA)
llama-server.exe ^
--model LFM2.5-230M-ToMoE-Q4_K_M.gguf ^
--gpu-layers all ^
--ctx-size 32768 ^
--alias LFM2.5-230M-ToMoE
Or with llama-cli:
llama-cli.exe -m LFM2.5-230M-ToMoE-Q4_K_M.gguf -ngl 99 -p "The capital of France is" -n 25
Measured on an RTX 3060 (CUDA 13.3 build, llama.cpp b10603): ~540–550 t/s generation for Q4_K_M, ~430–450 t/s for Q8_0.
Base model
- MoE source: Nichonauta/LFM2.5-230M-ToMoE
- Dense base: LiquidAI/LFM2.5-230M (LFM Open License v1.0)
License
Derivative of LiquidAI/LFM2.5-230M — released under the LFM Open License v1.0 (see LICENSE).
Run Nichonauta/LFM2.5-230M-ToMoE-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models