prithivMLmods/OpenCaption-2B-VL-SFT-v1.0-GGUF overview
OpenCaption 2B VL SFT v1.0 GGUF OpenCaption 2B VL SFT v1.0 is a vision language captioning model built on top of Qwen/Qwen3 VL 2B Instruct . It was trained for…
Runs locally from ~424.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| OpenCaption-2B-VL-SFT-v1.0.BF16.gguf | GGUF | GGUF | 3.21 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.F16.gguf | GGUF | GGUF | 3.21 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.F32.gguf | GGUF | GGUF | 6.42 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q3_K_L.gguf | GGUF | GGUF | 957.0 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q3_K_M.gguf | GGUF | GGUF | 896.0 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q3_K_S.gguf | GGUF | GGUF | 827.1 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q4_0.gguf | GGUF | GGUF | 1005.6 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q4_K_M.gguf | GGUF | GGUF | 1.03 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q4_K_S.gguf | GGUF | GGUF | 1011.1 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q5_0.gguf | GGUF | GGUF | 1.15 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q5_K_M.gguf | GGUF | GGUF | 1.17 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q5_K_S.gguf | GGUF | GGUF | 1.15 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q6_K.gguf | GGUF | GGUF | 1.32 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q8_0.gguf | GGUF | GGUF | 1.71 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-bf16.gguf | GGUF | BF16 | 784.4 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-f16.gguf | GGUF | F16 | 784.4 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-f32.gguf | GGUF | F32 | 1.52 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-q8_0.gguf | GGUF | Q8_0 | 424.4 MB | Download |
Model Details
| Model ID | prithivMLmods/OpenCaption-2B-VL-SFT-v1.0-GGUF |
|---|---|
| Author | prithivMLmods |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | prithivMLmods/OpenCaption-2B-VL-SFT-v1.0 |
| Last modified | 2026-07-25T05:16:46.000Z |
Model README
---
license: apache-2.0
base_model:
- prithivMLmods/OpenCaption-2B-VL-SFT-v1.0
tags:
- text-generation-inference
- llama-cpp
- vision-language
- multimodal
- image-captioning
- visual-question-answering
- conditional-generation
- vision
- language-model
- sft
- fine-grained-captioning
- computer-vision
language:
- en
pipeline_tag: image-text-to-text
library_name: transformers
datasets:
- prithivMLmods/OpenCaption-FineGrained
- prithivMLmods/SuperFlickr-30K-LARGE-Remastered
- prithivMLmods/OpenCaption-UHD
- prithivMLmods/OpenCaption-Unified-10K
---
OpenCaption-2B-VL-SFT-v1.0-GGUF
> OpenCaption-2B-VL-SFT-v1.0 is a vision-language captioning model built on top of Qwen/Qwen3-VL-2B-Instruct. It was trained for image captioning and high-quality dense image captioning using a fine-grained mixture of long-form image description traces. The model is designed to produce richer visual understanding than conventional captions by capturing scene composition, object relationships, spatial reasoning, attributes, actions, lighting, background context, and fine visual details.
> [!NOTE]
> This model is an experimental release and may produce unexpected outputs in some scenarios.
Model Files
File Name | Quant Type | File Size | File Link |
|-----------|------------|-----------|-----------|
| OpenCaption-2B-VL-SFT-v1.0.BF16.gguf | BF16 | 3.45 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.F16.gguf | F16 | 3.45 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.F32.gguf | F32 | 6.89 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q3_K_L.gguf | Q3_K_L | 1 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q3_K_M.gguf | Q3_K_M | 940 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q3_K_S.gguf | Q3_K_S | 867 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q4_0.gguf | Q4_0 | 1.05 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q4_K_M.gguf | Q4_K_M | 1.11 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q4_K_S.gguf | Q4_K_S | 1.06 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q5_0.gguf | Q5_0 | 1.23 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q5_K_M.gguf | Q5_K_M | 1.26 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q5_K_S.gguf | Q5_K_S | 1.23 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q6_K.gguf | Q6_K | 1.42 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.Q8_0.gguf | Q8_0 | 1.83 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-bf16.gguf | mmproj-bf16 | 823 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-f16.gguf | mmproj-f16 | 823 MB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-f32.gguf | mmproj-f32 | 1.63 GB | Download |
| OpenCaption-2B-VL-SFT-v1.0.mmproj-q8_0.gguf | mmproj-q8_0 | 445 MB | Download |
llama.cpp
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp
Run prithivMLmods/OpenCaption-2B-VL-SFT-v1.0-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models