GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Abiray/OvisOCR2-GGUF overview

OvisOCR2 GGUF Quantizations <p align="center" <img src="https://cdn uploads.huggingface.co/production/uploads/658a8a837959448ef5500ce5/vRCIu5QD8VuIJolkC ZHQ.pn…

ggufocrdocument-parsingmultimodalmarkdowntablesformulasllama-cppimage-text-to-textbase_model:ATH-MaaS/OvisOCR2base_model:quantized:ATH-MaaS/OvisOCR2license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~195.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
image-text-to-text
Author

Repository Files & Downloads

12 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
OvisOCR2-BF16.ggufGGUFBF161.41 GBDownload
OvisOCR2-F16.ggufGGUFF161.41 GBDownload
OvisOCR2-Q3_K_M.ggufGGUFQ3_K_M444.6 MBDownload
OvisOCR2-Q4_K_M.ggufGGUFQ4_K_M504.8 MBDownload
OvisOCR2-Q4_K_S.ggufGGUFQ4_K_S481.8 MBDownload
OvisOCR2-Q5_K_M.ggufGGUFQ5_K_M551.2 MBDownload
OvisOCR2-Q5_K_S.ggufGGUFQ5_K_S537.5 MBDownload
OvisOCR2-Q6_K.ggufGGUFQ6_K600.6 MBDownload
OvisOCR2-Q8_0.ggufGGUFQ8_0774.2 MBDownload
mmproj-BF16.ggufGGUFBF16197.7 MBDownload
mmproj-F16.ggufGGUFF16195.5 MBDownload
mmproj-F32.ggufGGUFF32383.7 MBDownload

Model Details

Model IDAbiray/OvisOCR2-GGUF
AuthorAbiray
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelATH-MaaS/OvisOCR2
Last modified2026-07-13T18:34:25.000Z

Model README

---

license: apache-2.0

library_name: gguf

pipeline_tag: image-text-to-text

base_model: ATH-MaaS/OvisOCR2

tags:

  • ocr
  • document-parsing
  • multimodal
  • markdown
  • tables
  • formulas
  • gguf
  • llama-cpp

---

OvisOCR2 - GGUF Quantizations

<p align="center">

<img src="https://cdn-uploads.huggingface.co/production/uploads/658a8a837959448ef5500ce5/vRCIu5QD8VuIJolkC_ZHQ.png" alt="Ovis" width="30%" />

</p>

This repository contains GGUF format quantizations of OvisOCR2, a compact 0.8B end-to-end model for page-level document parsing. The original model was developed by ATH-MaaS by post-training Qwen3.5-0.8B to parse full document pages directly into clean Markdown (including LaTeX formulas, HTML tables, and layout components).

OvisOCR2 establishes a new state-of-the-art for compact document understanding, scoring 96.58 on OmniDocBench v1.6 and outperforming traditional, multi-stage layout analysis pipelines.

---

Available Files

Main Text Models

| File Name | Precision / Quantization | File Size | Description |

| :--- | :--- | :--- | :--- |

| OvisOCR2-F16.gguf | 16-bit Float | 1.52 GB | Baseline unquantized model |

| OvisOCR2-BF16.gguf | 16-bit Brain Float | 1.52 GB | Native weight precision |

| OvisOCR2-Q8_0.gguf | 8-bit | 812 MB | Near-identical precision to F16 |

| OvisOCR2-Q6_K.gguf | 6-bit | 630 MB | Excellent balance of size and accuracy |

| OvisOCR2-Q5_K_M.gguf | 5-bit (Medium) | 578 MB | Recommended for low-resource deployment |

| OvisOCR2-Q5_K_S.gguf | 5-bit (Small) | 564 MB | Highly optimized 5-bit layout |

| OvisOCR2-Q4_K_M.gguf | 4-bit (Medium) | 529 MB | Standard 4-bit quantization |

| OvisOCR2-Q4_K_S.gguf | 4-bit (Small) | 505 MB | Lightweight 4-bit footprint |

| OvisOCR2-Q3_K_M.gguf | 3-bit (Medium) | 466 MB | Maximum compression ratio |

Multimodal Projectors (mmproj)

Note: Because OvisOCR2 is a vision-language model, you must download one of these image processing units alongside your choice of the text models listed above.

  • mmproj-F32.gguf (402 MB) - Unquantized full precision projector.
  • mmproj-F16.gguf (205 MB) - Recommended standard performance/size option.
  • mmproj-BF16.gguf (207 MB) - Target alternative precision layout.

---

Inference Guide (llama.cpp)

To run multimodal OCR tasks using these GGUF files, you need to use the llama-minicpmv-cli or llama-llava-cli tool (depending on your build version of llama.cpp) to handle simultaneous image and text tokens.

Basic Command Line Example

# Run parsing via llama.cpp cli tools
./llama-minicpmv-cli \
  -m OvisOCR2-Q5_K_M.gguf \
  --mmproj mmproj-F16.gguf \
  --image /path/to/your/document_page.jpg \
  -p "<|im_start|>user\nExtract all readable content from the image in natural human reading order and output the result as a single Markdown document. Format formulas as LaTeX. Format tables as HTML: <table>...</table>. Preserve the original text without translation.<|im_end|>\n<|im_start|>assistant\n" \
  -n 4096 \
  --temp 0.0

Run Abiray/OvisOCR2-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models