prithivMLmods/OpenCaption-4B-VL-SFT-v1.0-GGUF overview
OpenCaption 4B VL SFT v1.0 GGUF OpenCaption 4B VL SFT v1.0 is a vision language captioning model built on top of Qwen/Qwen3 VL 4B Instruct . It was trained for…
Runs locally from ~432.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| OpenCaption-4B-VL-SFT-v1.0.BF16.gguf | GGUF | GGUF | 7.50 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.F16.gguf | GGUF | GGUF | 7.50 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.F32.gguf | GGUF | GGUF | 14.99 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q3_K_L.gguf | GGUF | GGUF | 2.09 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q3_K_M.gguf | GGUF | GGUF | 1.93 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q3_K_S.gguf | GGUF | GGUF | 1.76 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q4_0.gguf | GGUF | GGUF | 2.21 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q4_K_M.gguf | GGUF | GGUF | 2.33 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q4_K_S.gguf | GGUF | GGUF | 2.22 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q5_0.gguf | GGUF | GGUF | 2.63 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q5_K_M.gguf | GGUF | GGUF | 2.69 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q5_K_S.gguf | GGUF | GGUF | 2.63 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q6_K.gguf | GGUF | GGUF | 3.08 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q8_0.gguf | GGUF | GGUF | 3.99 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-bf16.gguf | GGUF | BF16 | 800.4 MB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-f16.gguf | GGUF | F16 | 800.4 MB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-f32.gguf | GGUF | F32 | 1.55 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-q8_0.gguf | GGUF | Q8_0 | 432.9 MB | Download |
Model Details
| Model ID | prithivMLmods/OpenCaption-4B-VL-SFT-v1.0-GGUF |
|---|---|
| Author | prithivMLmods |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | prithivMLmods/OpenCaption-4B-VL-SFT-v1.0 |
| Last modified | 2026-07-25T05:17:26.000Z |
Model README
---
base_model:
- prithivMLmods/OpenCaption-4B-VL-SFT-v1.0
license: apache-2.0
language:
- en
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- text-generation-inference
- llama-cpp
- vision-language
- multimodal
- image-captioning
- visual-question-answering
- conditional-generation
- vision
- language-model
- sft
- fine-grained-captioning
- computer-vision
datasets:
- prithivMLmods/OpenCaption-FineGrained
- prithivMLmods/SuperFlickr-30K-LARGE-Remastered
- prithivMLmods/OpenCaption-UHD
- prithivMLmods/OpenCaption-Unified-10K
---
OpenCaption-4B-VL-SFT-v1.0-GGUF
> OpenCaption-4B-VL-SFT-v1.0 is a vision-language captioning model built on top of Qwen/Qwen3-VL-4B-Instruct. It was trained for image captioning and high-quality dense image captioning using a fine-grained mixture of long-form image description traces. The model is designed to produce richer visual understanding than conventional captions by capturing scene composition, object relationships, spatial reasoning, attributes, actions, lighting, background context, and fine visual details.
> [!NOTE]
> This model is an experimental release and may produce unexpected outputs in some scenarios.
Model Files
File Name | Quant Type | File Size | File Link |
|-----------|------------|-----------|-----------|
| OpenCaption-4B-VL-SFT-v1.0.BF16.gguf | BF16 | 8.05 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.F16.gguf | F16 | 8.05 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.F32.gguf | F32 | 16.1 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q3_K_L.gguf | Q3_K_L | 2.24 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q3_K_M.gguf | Q3_K_M | 2.08 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q3_K_S.gguf | Q3_K_S | 1.89 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q4_0.gguf | Q4_0 | 2.37 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q4_K_M.gguf | Q4_K_M | 2.5 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q4_K_S.gguf | Q4_K_S | 2.38 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q5_0.gguf | Q5_0 | 2.82 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q5_K_M.gguf | Q5_K_M | 2.89 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q5_K_S.gguf | Q5_K_S | 2.82 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q6_K.gguf | Q6_K | 3.31 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.Q8_0.gguf | Q8_0 | 4.28 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-bf16.gguf | mmproj-bf16 | 839 MB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-f16.gguf | mmproj-f16 | 839 MB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-f32.gguf | mmproj-f32 | 1.66 GB | Download |
| OpenCaption-4B-VL-SFT-v1.0.mmproj-q8_0.gguf | mmproj-q8_0 | 454 MB | Download |
llama.cpp
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp
Run prithivMLmods/OpenCaption-4B-VL-SFT-v1.0-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models