prithivMLmods/oMEGA-4B-SpatialThink-0804-GGUF overview
oMEGA 4B SpatialThink 0804 GGUF oMEGA 4B SpatialThink 0804 is a vision language model built on top of Qwen/Qwen3 VL 4B Instruct and fine tuned for spatial reas…
Runs locally from ~432.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| oMEGA-4B-SpatialThink-0804.BF16.gguf | GGUF | GGUF | 7.50 GB | Download |
| oMEGA-4B-SpatialThink-0804.F16.gguf | GGUF | GGUF | 7.50 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q3_K_L.gguf | GGUF | GGUF | 2.09 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q3_K_M.gguf | GGUF | GGUF | 1.93 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q3_K_S.gguf | GGUF | GGUF | 1.76 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q4_K_M.gguf | GGUF | GGUF | 2.33 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q4_K_S.gguf | GGUF | GGUF | 2.22 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q5_K_M.gguf | GGUF | GGUF | 2.69 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q5_K_S.gguf | GGUF | GGUF | 2.63 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q6_K.gguf | GGUF | GGUF | 3.08 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q8_0.gguf | GGUF | GGUF | 3.99 GB | Download |
| oMEGA-4B-SpatialThink-0804.mmproj-bf16.gguf | GGUF | BF16 | 800.4 MB | Download |
| oMEGA-4B-SpatialThink-0804.mmproj-f16.gguf | GGUF | F16 | 800.4 MB | Download |
| oMEGA-4B-SpatialThink-0804.mmproj-q8_0.gguf | GGUF | Q8_0 | 432.9 MB | Download |
Model Details
| Model ID | prithivMLmods/oMEGA-4B-SpatialThink-0804-GGUF |
|---|---|
| Author | prithivMLmods |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | prithivMLmods/oMEGA-4B-SpatialThink-0804 |
| Last modified | 2026-08-04T09:35:44.000Z |
Model README
---
base_model:
- prithivMLmods/oMEGA-4B-SpatialThink-0804
tags:
- text-generation-inference
- spatial-reasoning
- llama-cpp
- vision-language
- multimodal
- image-captioning
- visual-question-answering
- conditional-generation
- vision
- language-model
- sft
- fine-grained-captioning
- computer-vision
datasets:
- prithivMLmods/OpenCaption-FineGrained
- remyxai/SpaceThinker
license: apache-2.0
language:
- en
pipeline_tag: image-text-to-text
library_name: transformers
---
oMEGA-4B-SpatialThink-0804-GGUF
> oMEGA-4B-SpatialThink-0804 is a vision-language model built on top of Qwen/Qwen3-VL-4B-Instruct and fine-tuned for spatial reasoning with concise notes for unfiltered vision tasks. The model is trained to produce concise yet informative reasoning for spatial understanding while maintaining strong image captioning capabilities. Training is based on remyxai's SpaceThinker and OpenCaption-FineGrained, enabling efficient spatial reasoning and detailed image understanding across diverse visual domains.
> [!NOTE]
> This model is an experimental release and may generate unexpected behaviors or reasoning artifacts in certain scenarios.
Model Files
File Name | Quant Type | File Size | File Link |
|-----------|------------|-----------|-----------|
| oMEGA-4B-SpatialThink-0804.BF16.gguf | BF16 | 8.05 GB | Download |
| oMEGA-4B-SpatialThink-0804.F16.gguf | F16 | 8.05 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q3_K_L.gguf | Q3_K_L | 2.24 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q3_K_M.gguf | Q3_K_M | 2.08 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q3_K_S.gguf | Q3_K_S | 1.89 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q4_K_M.gguf | Q4_K_M | 2.5 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q4_K_S.gguf | Q4_K_S | 2.38 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q5_K_M.gguf | Q5_K_M | 2.89 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q5_K_S.gguf | Q5_K_S | 2.82 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q6_K.gguf | Q6_K | 3.31 GB | Download |
| oMEGA-4B-SpatialThink-0804.Q8_0.gguf | Q8_0 | 4.28 GB | Download |
| oMEGA-4B-SpatialThink-0804.mmproj-bf16.gguf | mmproj-bf16 | 839 MB | Download |
| oMEGA-4B-SpatialThink-0804.mmproj-f16.gguf | mmproj-f16 | 839 MB | Download |
| oMEGA-4B-SpatialThink-0804.mmproj-q8_0.gguf | mmproj-q8_0 | 454 MB | Download |
llama.cpp
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp
Run prithivMLmods/oMEGA-4B-SpatialThink-0804-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models