GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

jamiefutch/Qwen3.5-35B-A3B-MXFP4_MOE-GGUF overview

These are quantizations of the model Qwen3.5 35B A3B https://huggingface.co/unsloth/Qwen3.5 35B A3B Download the latest llama.cpp https://github.com/ggml org/l…

ggufimage-text-to-textbase_model:Qwen/Qwen3.5-35B-A3Bbase_model:quantized:Qwen/Qwen3.5-35B-A3Bendpoints_compatibleregion:usconversational

Runs locally from ~861.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
image-text-to-text

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.5-35B-A3B-MXFP4_MOE_BF16.ggufGGUFBF1620.55 GBDownload
Qwen3.5-35B-A3B-MXFP4_MOE_F16.ggufGGUFF1620.55 GBDownload
mmproj-BF16.ggufGGUFBF16861.0 MBDownload
mmproj-F32.ggufGGUFF321.66 GBDownload

Model Details

Model IDjamiefutch/Qwen3.5-35B-A3B-MXFP4_MOE-GGUF
Authorjamiefutch
Pipelineimage-text-to-text
License
Base modelQwen/Qwen3.5-35B-A3B
Last modified2026-07-06T10:19:25.000Z

Model README

---

pipeline_tag: image-text-to-text

base_model:

  • Qwen/Qwen3.5-35B-A3B

---

These are quantizations of the model Qwen3.5-35B-A3B

  • Download the latest llama.cpp to use these quantizations.
  • For the mmproj file, the F32 version is recommended for best results.

The mmproj files are the same from unsloth.

Read the guide from unsloth in order to set up the model's recommended settings:

Qwen3.5 - How to Run Locally Guide

The mainline standard is to use MXFP4 for the MoE tensors, and Q8 for the rest.

So I created 2 new variants, where the other tensors are either BF16 or FP16 instead of Q8.

The order of preference is BF16, then F16.

On some architectures BF16 will be slower, but its the highest quality, essentialy its the original tensors from the model copied over unquantized.

As of 2026-03-03 the chat template has been updated, and is using the fixed version from unsloth.

If you don't want to download the model again, you can just update the chat template.

Run jamiefutch/Qwen3.5-35B-A3B-MXFP4_MOE-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models