GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF overview

Qwen3.6 35B A3B Fast NVFP4 GGUF GGUF NVFP4 quantization of unsloth/Qwen3.6 35B A3B NVFP4 Fast https://huggingface.co/unsloth/Qwen3.6 35B A3B NVFP4 Fast , a 35B…

ggufqwenqwen3_5_moeqwen3_6nvfp4fp4moevisionmultimodalfastimage-text-to-textenzhbase_model:unsloth/Qwen3.6-35B-A3B-NVFP4-Fastbase_model:quantized:unsloth/Qwen3.6-35B-A3B-NVFP4-Fastlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~857.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
image-text-to-text

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
mmproj-qwen36-35b-a3b-f16.ggufGGUFF16857.6 MBDownload
qwen36-35b-a3b-fast-nvfp4.ggufGGUFGGUF18.81 GBDownload

Model Details

Model IDFreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF
AuthorFreedomAISVR
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelunsloth/Qwen3.6-35B-A3B-NVFP4-Fast
Last modified2026-07-13T15:36:07.000Z

Model README

---

language:

  • en
  • zh

tags:

  • qwen
  • qwen3_5_moe
  • qwen3_6
  • nvfp4
  • fp4
  • gguf
  • moe
  • vision
  • multimodal
  • fast

license: apache-2.0

base_model: unsloth/Qwen3.6-35B-A3B-NVFP4-Fast

pipeline_tag: image-text-to-text

library_name: gguf

---

Qwen3.6-35B-A3B-Fast-NVFP4-GGUF

GGUF NVFP4 quantization of unsloth/Qwen3.6-35B-A3B-NVFP4-Fast, a 35B parameter MoE model with 3B active parameters.

What is the "Fast" Variant?

Unsloth's NVFP4 Fast variant is a speed-optimized NVFP4 quantization that delivers 1.79x faster throughput than other NVFP4 quants while maintaining competitive accuracy:

| Variant | MMLU-Pro | GPQA | AIME 2025 |

|---------|----------|------|-----------|

| Unsloth NVFP4 Fast | 85.58 | 87.75 | 91.67 |

| Unsloth NVFP4 | 85.85 | 86.74 | 92.29 |

| NVIDIA NVFP4 | 85.60 | 87.12 | 91.88 |

| BF16 | 85.75 | 86.36 | 92.50 |

The Fast variant is calibrated on a mixture of Unsloth's dataset + UltraChat dataset, optimized for throughput on vLLM's native NVFP4 backend (cute-DSL/CUTLASS/flashinfer_trtllm).

About the Model

Qwen3.6-35B-A3B is a multimodal MoE model from Alibaba's Qwen team:

  • 35B total parameters, 3B active per token (256 experts, 8 active)
  • 40-layer decoder with Gated DeltaNet + full attention hybrid
  • 27-layer vision encoder (SigLIP-based) for image/video understanding
  • 262K native context (extensible to 1M+ via YaRN)
  • Multi-Token Prediction (MTP) for faster speculative decoding
  • Agentic coding with SWE-bench Verified 73.4, tool calling support

Files

| File | Size | Description |

|------|------|-------------|

| qwen36-35b-a3b-fast-nvfp4.gguf | ~18.8 GB | NVFP4 quantized text model |

| mmproj-qwen36-35b-a3b-f16.gguf | ~0.84 GB | Vision encoder (F16) |

Usage

llama.cpp

llama-server \
  -m qwen36-35b-a3b-fast-nvfp4.gguf \
  --mmproj mmproj-qwen36-35b-a3b-f16.gguf \
  -ngl 99 \
  --host 0.0.0.0 \
  --port 8080

vLLM (Recommended for Max Performance)

pip install vllm flashinfer-python nvidia-cutlass-dsl
vllm serve FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF --dtype auto --max-model-len 4096

Hardware Requirements

  • Minimum: 24 GB VRAM for partial offload
  • Recommended: 32+ GB VRAM for full GPU offload
  • Optimal: Blackwell B200 for max NVFP4 throughput

Quantization

Quantized from Qwen/Qwen3.6-35B-A3B BF16 weights using llama.cpp (llama-quantize.exe NVFP4).

License

Apache 2.0 - same as the base model.

Credits

Run FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models