GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

prithivMLmods/Spatial-Interactor-Qwen3-VL-4B-GGUF overview

Spatial Interactor Qwen3 VL 4B GGUF Spatial Interactor Qwen3 VL 4B https://huggingface.co/kagakouko/Spatial Interactor Qwen3 VL 4B is a full parameter BF16 che…

transformersgguftext-generation-inferencellama-cppvision-languagevideospatial-reasoningembodied-aiqwenimage-text-to-textenbase_model:kagakouko/Spatial-Interactor-Qwen3-VL-4Bbase_model:quantized:kagakouko/Spatial-Interactor-Qwen3-VL-4Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~800.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
image-text-to-text

Repository Files & Downloads

8 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Spatial-Interactor-Qwen3-VL-4B.BF16.ggufGGUFGGUF8.22 GBDownload
Spatial-Interactor-Qwen3-VL-4B.Q3_K_L.ggufGGUFGGUF2.24 GBDownload
Spatial-Interactor-Qwen3-VL-4B.Q3_K_M.ggufGGUFGGUF2.09 GBDownload
Spatial-Interactor-Qwen3-VL-4B.Q4_K_M.ggufGGUFGGUF2.53 GBDownload
Spatial-Interactor-Qwen3-VL-4B.Q4_K_S.ggufGGUFGGUF2.42 GBDownload
Spatial-Interactor-Qwen3-VL-4B.Q5_K_M.ggufGGUFGGUF2.94 GBDownload
Spatial-Interactor-Qwen3-VL-4B.Q5_K_S.ggufGGUFGGUF2.88 GBDownload
Spatial-Interactor-Qwen3-VL-4B.mmproj-bf16.ggufGGUFBF16800.4 MBDownload

Model Details

Model IDprithivMLmods/Spatial-Interactor-Qwen3-VL-4B-GGUF
AuthorprithivMLmods
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelkagakouko/Spatial-Interactor-Qwen3-VL-4B
Last modified2026-09-17T03:23:23.000Z

Model README

---

license: apache-2.0

base_model:

  • kagakouko/Spatial-Interactor-Qwen3-VL-4B

tags:

  • text-generation-inference
  • llama-cpp
  • vision-language
  • video
  • spatial-reasoning
  • embodied-ai
  • qwen

language:

  • en

pipeline_tag: image-text-to-text

library_name: transformers

---

Spatial-Interactor-Qwen3-VL-4B-GGUF

> Spatial-Interactor-Qwen3-VL-4B is a full-parameter BF16 checkpoint from ZJU-OmniAI's Spatial-Interactor project, built on Qwen3-VL-4B-Instruct to learn spatial reasoning through interaction with the observable physical world, as detailed in the accompanying paper. It's trained in two stages: supervised fine-tuning on the LSI-108K dataset's L1-L2 split (local world-state and ego-motion transition modeling) alongside a public spatial QA mixture, followed by On-Policy Distillation (OPD) that integrates verifiable answer rewards with same-prefix privileged self-distillation to guide intermediate reasoning over long-horizon video trajectories — with the visual encoder kept frozen throughout while only the language model and multimodal projector are updated. Critically, the privileged transition trace used during training is discarded at inference time, so the deployed model takes the same image/video-plus-question inputs as its base model with no extra trace, reward model, or teacher branch; for video evaluation, the paper's main results use 32 ordered frames with chronological order preserved. It's one of four checkpoints in the Spatial-Interactor collection, loadable via the standard Transformers interface in place of the base model identifier, and released under Apache-2.0 following the base model's license.

Model Files

File Name | Quant Type | File Size | File Link |

|-----------|------------|-----------|-----------|

| Spatial-Interactor-Qwen3-VL-4B.BF16.gguf | BF16 | 8.83 GB | Download |

| Spatial-Interactor-Qwen3-VL-4B.Q3_K_L.gguf | Q3_K_L | 2.41 GB | Download |

| Spatial-Interactor-Qwen3-VL-4B.Q3_K_M.gguf | Q3_K_M | 2.24 GB | Download |

| Spatial-Interactor-Qwen3-VL-4B.Q4_K_M.gguf | Q4_K_M | 2.72 GB | Download |

| Spatial-Interactor-Qwen3-VL-4B.Q4_K_S.gguf | Q4_K_S | 2.6 GB | Download |

| Spatial-Interactor-Qwen3-VL-4B.Q5_K_M.gguf | Q5_K_M | 3.16 GB | Download |

| Spatial-Interactor-Qwen3-VL-4B.Q5_K_S.gguf | Q5_K_S | 3.09 GB | Download |

| Spatial-Interactor-Qwen3-VL-4B.mmproj-bf16.gguf | mmproj-bf16 | 839 MB | Download |

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Run prithivMLmods/Spatial-Interactor-Qwen3-VL-4B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models