GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

tooltd/GRM-3.2-Sky-MTP-GGUF overview

tooltd/GRM 3.2 Sky MTP GGUF This model was converted to GGUF format from OrionLLM/GRM 3.2 Sky https://huggingface.co/OrionLLM/GRM 3.2 Sky using llama.cpp via t…

ggufllama-cppimage-text-to-textbase_model:OrionLLM/GRM-3.2-Skybase_model:quantized:OrionLLM/GRM-3.2-Skylicense:apache-2.0endpoints_compatibleregion:us

Runs locally from ~34.37 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
10
Likes
0
Pipeline
image-text-to-text
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
grm-3.2-sky-q8_0.ggufGGUFQ8_034.37 GBDownload

Model Details

Model IDtooltd/GRM-3.2-Sky-MTP-GGUF
Authortooltd
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelOrionLLM/GRM-3.2-Sky
Last modified2026-08-02T04:28:44.000Z

Model README

---

license: apache-2.0

base_model: OrionLLM/GRM-3.2-Sky

pipeline_tag: image-text-to-text

tags:

  • llama-cpp

---

tooltd/GRM-3.2-Sky-MTP-GGUF

This model was converted to GGUF format from OrionLLM/GRM-3.2-Sky using llama.cpp via the ggml.ai's GGUF-my-repo space.

Refer to the original model card for more details on the model.

MTP HEAD

  1. Download Qwen3.6-35B-A3B MTP-ONLY from https://huggingface.co/a4lg/Qwen3.6-35B-A3B-MTP-ONLY-GGUF
  1. Add MTP HEAD to GRM-3.2-Sky GGUF model by using my script modified. otherwise, it will cause an error.

https://huggingface.co/tooltd/GRM-3.2-Sky-MTP-GGUF/blob/main/grm_mtp_graft.py

python grm_mtp_graft.py grm-3.2-sky-q8_0.gguf Qwen3.6-35B-A3B-MTP-ONLY-Q8_0.gguf grm-3.2-sky-mtp-q8_0.gguf

Draft acceptance rate is usually above 80%. Done!

Run tooltd/GRM-3.2-Sky-MTP-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models