jamiefutch/Qwopus3.6-35B-A3B-v1-MTP-MXFP4_MOE-GGUF overview
These are MXFP4 quantizations of the model Jackrong / Qwopus3.6 35B A3B v1 https://huggingface.co/Jackrong/Qwopus3.6 35B A3B v1 This is the multi token predict…
Runs locally from ~19.71 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | jamiefutch/Qwopus3.6-35B-A3B-v1-MTP-MXFP4_MOE-GGUF |
|---|---|
| Author | jamiefutch |
| Pipeline | image-text-to-text |
| License | — |
| Base model | Jackrong/Qwopus3.6-35B-A3B-v1 |
| Last modified | 2026-07-12T09:49:08.000Z |
Model README
---
pipeline_tag: image-text-to-text
base_model:
- Jackrong/Qwopus3.6-35B-A3B-v1
tags:
- qwen
- qwen3_5_moe
- MTP
---
These are MXFP4 quantizations of the model Jackrong / Qwopus3.6-35B-A3B-v1
This is the multi-token prediction (MTP) version.
Quick Start
- Download the latest release of llama.cpp.
- Download your preferred model variant from below.
Which version should I choose?
All variants use MXFP4 for the MoE (Mixture of Experts) weights to keep the model efficient. The difference lies in how the remaining tensors are handled:
| Variant | Quality | Performance | Size | Recommendation |
| :--- | :--- | :--- | ---: | :--- |
| BF16 | ⭐⭐⭐ | Variable* | 21.39GiB | Best for maximum accuracy; original unquantized weights. |
| F16 | ⭐⭐ | Fast | 21.39GiB | Great alternative if BF16 is slow on your hardware. |
| Q8 | ⭐ | Fastest | 19.71GiB | Balanced performance and memory usage. |
\Note: On some older architectures, BF16 may be slower than F16. Check that your GPU supports native BF16 *
Read the guide from unsloth in order to set up the model's recommended settings for MTP:
On my system it works very well with the commands:
--spec-type draft-mtp
--spec-draft-p-min 0.75
--spec-draft-n-max 3Run jamiefutch/Qwopus3.6-35B-A3B-v1-MTP-MXFP4_MOE-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models