FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF overview
Qwen3.6 35B A3B Fast NVFP4 GGUF GGUF NVFP4 quantization of unsloth/Qwen3.6 35B A3B NVFP4 Fast https://huggingface.co/unsloth/Qwen3.6 35B A3B NVFP4 Fast , a 35B…
Runs locally from ~857.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF |
|---|---|
| Author | FreedomAISVR |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | unsloth/Qwen3.6-35B-A3B-NVFP4-Fast |
| Last modified | 2026-07-13T15:36:07.000Z |
Model README
---
language:
- en
- zh
tags:
- qwen
- qwen3_5_moe
- qwen3_6
- nvfp4
- fp4
- gguf
- moe
- vision
- multimodal
- fast
license: apache-2.0
base_model: unsloth/Qwen3.6-35B-A3B-NVFP4-Fast
pipeline_tag: image-text-to-text
library_name: gguf
---
Qwen3.6-35B-A3B-Fast-NVFP4-GGUF
GGUF NVFP4 quantization of unsloth/Qwen3.6-35B-A3B-NVFP4-Fast, a 35B parameter MoE model with 3B active parameters.
What is the "Fast" Variant?
Unsloth's NVFP4 Fast variant is a speed-optimized NVFP4 quantization that delivers 1.79x faster throughput than other NVFP4 quants while maintaining competitive accuracy:
| Variant | MMLU-Pro | GPQA | AIME 2025 |
|---------|----------|------|-----------|
| Unsloth NVFP4 Fast | 85.58 | 87.75 | 91.67 |
| Unsloth NVFP4 | 85.85 | 86.74 | 92.29 |
| NVIDIA NVFP4 | 85.60 | 87.12 | 91.88 |
| BF16 | 85.75 | 86.36 | 92.50 |
The Fast variant is calibrated on a mixture of Unsloth's dataset + UltraChat dataset, optimized for throughput on vLLM's native NVFP4 backend (cute-DSL/CUTLASS/flashinfer_trtllm).
About the Model
Qwen3.6-35B-A3B is a multimodal MoE model from Alibaba's Qwen team:
- 35B total parameters, 3B active per token (256 experts, 8 active)
- 40-layer decoder with Gated DeltaNet + full attention hybrid
- 27-layer vision encoder (SigLIP-based) for image/video understanding
- 262K native context (extensible to 1M+ via YaRN)
- Multi-Token Prediction (MTP) for faster speculative decoding
- Agentic coding with SWE-bench Verified 73.4, tool calling support
Files
| File | Size | Description |
|------|------|-------------|
| qwen36-35b-a3b-fast-nvfp4.gguf | ~18.8 GB | NVFP4 quantized text model |
| mmproj-qwen36-35b-a3b-f16.gguf | ~0.84 GB | Vision encoder (F16) |
Usage
llama.cpp
llama-server \
-m qwen36-35b-a3b-fast-nvfp4.gguf \
--mmproj mmproj-qwen36-35b-a3b-f16.gguf \
-ngl 99 \
--host 0.0.0.0 \
--port 8080
vLLM (Recommended for Max Performance)
pip install vllm flashinfer-python nvidia-cutlass-dsl
vllm serve FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF --dtype auto --max-model-len 4096
Hardware Requirements
- Minimum: 24 GB VRAM for partial offload
- Recommended: 32+ GB VRAM for full GPU offload
- Optimal: Blackwell B200 for max NVFP4 throughput
Quantization
Quantized from Qwen/Qwen3.6-35B-A3B BF16 weights using llama.cpp (llama-quantize.exe NVFP4).
License
Apache 2.0 - same as the base model.
Credits
- Original model: Qwen/Qwen3.6-35B-A3B
- NVFP4 Fast quantization: unsloth/Qwen3.6-35B-A3B-NVFP4-Fast
- GGUF conversion: FreedomAISVR
Run FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models