Model Intelligence Sheet
nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF overview
Pruned layers after 49 as they arent needed for use with MiniMax H3. You need only the model for T2V and model + mmproj for I2V. place both the model and the m…
Runs locally from ~1.12 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
7 GGUF files detected
Direct downloads for local inference
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3-VL-32B-Instruct-MiniMax-H3-L0-49-IQ4_XS.gguf | GGUF | IQ4_XS | 12.52 GB | Download |
| Qwen3-VL-32B-Instruct-MiniMax-H3-L0-49-Q4_K_M.gguf | GGUF | Q4_K_M | 13.91 GB | Download |
| Qwen3-VL-32B-Instruct-MiniMax-H3-L0-49-UD-IQ2_XXS.gguf | GGUF | IQ2_XXS | 6.44 GB | Download |
| Qwen3-VL-32B-Instruct-MiniMax-H3-L0-49-UD-IQ3_XXS.gguf | GGUF | IQ3_XXS | 9.10 GB | Download |
| Qwen3-VL-32B-Instruct-MiniMax-H3-L0-49-UD-Q2_K_XL.gguf | GGUF | Q2_K_XL | 8.91 GB | Download |
| Qwen3-VL-32B-Instruct-MiniMax-H3-L0-49-mmproj-BF16.gguf | GGUF | BF16 | 1.12 GB | Download |
| Qwen3-VL-32B-Instruct-MiniMax-H3-L0-49-mmproj-F32.gguf | GGUF | F32 | 2.22 GB | Download |
Model Details
Model README
---
license: apache-2.0
base_model:
- unsloth/Qwen3-VL-32B-Instruct-GGUF
---
Pruned layers after 49 as they arent needed for use with MiniMax H3.
You need only the model for T2V and model + mmproj for I2V. place both the model and the mmproj file in text_encoder folder.
Can be used in ComfyUI with my fork of City96's ComfyUI-GGUF.
https://github.com/Nif00/ComfyUI-GGUF
Run nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models