Koshkasa/Vortex5_Phoenix-X-26B-A4B-MXFP4_MOE-GGUF overview
What's that? MXFP4 MOE quantization of Vortex5/Phoenix X 26B A4B https://huggingface.co/Vortex5/Phoenix X 26B A4B with more aggressive compression than standar…
Runs locally from ~13.22 GB disk (16 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Phoenix-X-26B-A4B-MXFP4-MOE.gguf | GGUF | GGUF | 13.22 GB | Download |
Model Details
| Model ID | Koshkasa/Vortex5_Phoenix-X-26B-A4B-MXFP4_MOE-GGUF |
|---|---|
| Author | Koshkasa |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Vortex5/Phoenix-X-26B-A4B |
| Last modified | 2026-08-11T20:22:55.000Z |
Model README
---
license: apache-2.0
base_model:
- Vortex5/Phoenix-X-26B-A4B
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- quantized
- llama.cpp
- roleplay
- mixed precision
- mxfp4
quantized_by: Koshkasa
base_model_relation: quantized
---
What's that?
MXFP4_MOE quantization of Vortex5/Phoenix-X-26B-A4B with more aggressive compression than standard mxfp4_moe. Local attention at IQ4_XS, global attention at Q5_K, attn_output in global attention layers in Q6_K.
imatrix generated by alexokita
Disclosure
My only contribution is compute. This is not my merge. Have fun.
Model card incomplete. Tests and comparisons may be uploaded at a later date.
Run Koshkasa/Vortex5_Phoenix-X-26B-A4B-MXFP4_MOE-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models