GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

prithivMLmods/dots.ocr-GGUF overview

dots.ocr GGUF dots.ocr is a multilingual document layout parsing model developed by rednote hilab https://huggingface.co/dots studio/dots.ocr that unifies layo…

transformersgguftext-generation-inferencellama-cppimage-to-textocrdocument-parselayouttableformulacustom_codeimage-text-to-textenzhmultilingualbase_model:dots-studio/dots.ocrbase_model:quantized:dots-studio/dots.ocrlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~821.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
image-text-to-text

Repository Files & Downloads

14 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
dots.ocr.BF16.ggufGGUFGGUF3.32 GBDownload
dots.ocr.F16.ggufGGUFGGUF3.32 GBDownload
dots.ocr.Q3_K_L.ggufGGUFGGUF935.0 MBDownload
dots.ocr.Q3_K_M.ggufGGUFGGUF881.6 MBDownload
dots.ocr.Q3_K_S.ggufGGUFGGUF821.3 MBDownload
dots.ocr.Q4_K_M.ggufGGUFGGUF1.04 GBDownload
dots.ocr.Q4_K_S.ggufGGUFGGUF1021.9 MBDownload
dots.ocr.Q5_K_M.ggufGGUFGGUF1.20 GBDownload
dots.ocr.Q5_K_S.ggufGGUFGGUF1.17 GBDownload
dots.ocr.Q6_K.ggufGGUFGGUF1.36 GBDownload
dots.ocr.Q8_0.ggufGGUFGGUF1.76 GBDownload
dots.ocr.mmproj-bf16.ggufGGUFBF162.35 GBDownload
dots.ocr.mmproj-f16.ggufGGUFF162.35 GBDownload
dots.ocr.mmproj-q8_0.ggufGGUFQ8_01.25 GBDownload

Model Details

Model IDprithivMLmods/dots.ocr-GGUF
AuthorprithivMLmods
Pipelineimage-text-to-text
Licensemit
Base modeldots-studio/dots.ocr
Last modified2026-08-28T06:22:53.000Z

Model README

---

license: mit

base_model:

  • dots-studio/dots.ocr

language:

  • en
  • zh
  • multilingual

pipeline_tag: image-text-to-text

library_name: transformers

tags:

  • text-generation-inference
  • llama-cpp
  • image-to-text
  • ocr
  • document-parse
  • layout
  • table
  • formula
  • transformers
  • custom_code

---

dots.ocr-GGUF

> dots.ocr is a multilingual document layout parsing model developed by rednote-hilab that unifies layout detection and content recognition within a single vision-language model (VLM), built upon a compact 1.7B-parameter LLM foundation (based on Qwen2.5-VL). It achieves state-of-the-art performance on OmniDocBench across text recognition, table parsing, and reading order tasks, while delivering formula recognition results comparable to much larger models like Gemini 2.5 Pro and Doubao-1.5. The model supports over 100 languages and handles diverse document types including academic papers, books, slides, financial reports, exam papers, magazines, and newspapers, outputting structured JSON with bounding boxes, layout categories (Caption, Footnote, Formula, List-item, Page-footer, Page-header, Picture, Section-header, Table, Text, Title), and extracted text formatted as LaTeX for formulas, HTML for tables, and Markdown for all other content. It can be switched between full layout parsing, detection-only, OCR-only, and grounding OCR modes simply by changing the input prompt, and supports inference via both HuggingFace Transformers and vLLM, with vLLM version 0.9.1 recommended for production deployment.

Model Files

File Name | Quant Type | File Size | File Link |

|-----------|------------|-----------|-----------|

| dots.ocr.BF16.gguf | BF16 | 3.56 GB | Download |

| dots.ocr.F16.gguf | F16 | 3.56 GB | Download |

| dots.ocr.Q3_K_L.gguf | Q3_K_L | 980 MB | Download |

| dots.ocr.Q3_K_M.gguf | Q3_K_M | 924 MB | Download |

| dots.ocr.Q3_K_S.gguf | Q3_K_S | 861 MB | Download |

| dots.ocr.Q4_K_M.gguf | Q4_K_M | 1.12 GB | Download |

| dots.ocr.Q4_K_S.gguf | Q4_K_S | 1.07 GB | Download |

| dots.ocr.Q5_K_M.gguf | Q5_K_M | 1.29 GB | Download |

| dots.ocr.Q5_K_S.gguf | Q5_K_S | 1.26 GB | Download |

| dots.ocr.Q6_K.gguf | Q6_K | 1.46 GB | Download |

| dots.ocr.Q8_0.gguf | Q8_0 | 1.89 GB | Download |

| dots.ocr.mmproj-bf16.gguf | mmproj-bf16 | 2.53 GB | Download |

| dots.ocr.mmproj-f16.gguf | mmproj-f16 | 2.53 GB | Download |

| dots.ocr.mmproj-q8_0.gguf | mmproj-q8_0 | 1.34 GB | Download |

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Run prithivMLmods/dots.ocr-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models