FreedomAISVR/Qwen3.6-35B-A3B-MXFP4-MOE-Fast-GGUF overview
Qwen3.6 35B A3B Fast MXFP4 MOE GGUF GGUF MXFP4 MoE quantization of unsloth/Qwen3.6 35B A3B NVFP4 Fast https://huggingface.co/unsloth/Qwen3.6 35B A3B NVFP4 Fast…
Runs locally from ~857.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | FreedomAISVR/Qwen3.6-35B-A3B-MXFP4-MOE-Fast-GGUF |
|---|---|
| Author | FreedomAISVR |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | unsloth/Qwen3.6-35B-A3B-NVFP4-Fast |
| Last modified | 2026-07-13T15:36:12.000Z |
Model README
---
language:
- en
- zh
tags:
- qwen
- qwen3_5_moe
- qwen3_6
- mxfp4
- gguf
- moe
- vision
- multimodal
- fast
license: apache-2.0
base_model: unsloth/Qwen3.6-35B-A3B-NVFP4-Fast
pipeline_tag: image-text-to-text
library_name: gguf
---
Qwen3.6-35B-A3B-Fast-MXFP4-MOE-GGUF
GGUF MXFP4 MoE quantization of unsloth/Qwen3.6-35B-A3B-NVFP4-Fast, a 35B parameter MoE model with 3B active parameters.
What is the "Fast" Variant?
Unsloth's NVFP4 Fast variant is a speed-optimized quantization that delivers 1.79x faster throughput than other NVFP4 quants. This GGUF extends that optimization to MXFP4 MoE format:
| Variant | MMLU-Pro | GPQA | AIME 2025 |
|---------|----------|------|-----------|
| Unsloth NVFP4 Fast | 85.58 | 87.75 | 91.67 |
| Unsloth NVFP4 | 85.85 | 86.74 | 92.29 |
| NVIDIA NVFP4 | 85.60 | 87.12 | 91.88 |
| BF16 | 85.75 | 86.36 | 92.50 |
MXFP4 MoE Format
This quantization uses a hybrid approach for optimal quality:
- Expert weights: MXFP4 (E2M1 microscaling, 4-bit)
- Non-expert weights (attention, embeddings, norms): Q8_0 (8-bit)
MXFP4 is an open standard supported by AMD, NVIDIA, and Microsoft, making it compatible with a wider range of hardware.
About the Model
Qwen3.6-35B-A3B is a multimodal MoE model from Alibaba's Qwen team:
- 35B total parameters, 3B active per token (256 experts, 8 active)
- 40-layer decoder with Gated DeltaNet + full attention hybrid
- 27-layer vision encoder (SigLIP-based) for image/video understanding
- 262K native context (extensible to 1M+ via YaRN)
- Multi-Token Prediction (MTP) for faster speculative decoding
- Agentic coding with SWE-bench Verified 73.4, tool calling support
Files
| File | Size | Description |
|------|------|-------------|
| qwen36-35b-a3b-fast-mxfp4_moe.gguf | ~18.9 GB | MXFP4 MoE quantized text model |
| mmproj-qwen36-35b-a3b-f16.gguf | ~0.84 GB | Vision encoder (F16) |
Usage
llama.cpp
llama-server \
-m qwen36-35b-a3b-fast-mxfp4_moe.gguf \
--mmproj mmproj-qwen36-35b-a3b-f16.gguf \
-ngl 99 \
--host 0.0.0.0 \
--port 8080
Hardware Requirements
- Minimum: 24 GB VRAM for partial offload
- Recommended: 32+ GB VRAM for full GPU offload
Quantization
Quantized from Qwen/Qwen3.6-35B-A3B BF16 weights using llama.cpp (llama-quantize.exe --allow-requantize MXFP4_MOE).
License
Apache 2.0 - same as the base model.
Credits
- Original model: Qwen/Qwen3.6-35B-A3B
- NVFP4 Fast quantization: unsloth/Qwen3.6-35B-A3B-NVFP4-Fast
- GGUF conversion: FreedomAISVR
Run FreedomAISVR/Qwen3.6-35B-A3B-MXFP4-MOE-Fast-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models