GraySoft
Projects Models About FAQ Contact Download guIDE →

dinhquangson/monkeyocr-pro-1.2b-vision-gguf - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

dinhquangson/monkeyocr-pro-1.2b-vision-gguf overview

A high-performance vision-language model specialized for Optical Character Recognition (OCR) and document analysis. This repository contains GGUF format models optimized for both general use and LM Studio compatibility.

ggufqwen2_vlvisionocrdocument-analysismultimodalqwen2vllm-studioimage-text-to-textlicense:apache-2.0endpoints_compatibleregion:usconversational
dinhquangson/monkeyocr-pro-1.2b-vision-gguf visual
Downloads
383
Likes
3
Pipeline
image-text-to-text
Library
Visibility
Public
Access
Open

Repository Files & Downloads

3 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
MonkeyOCR-pro-1.2B-Recognition.gguf GGUF 1.23 GB Download
MonkeyOCR-pro-1.2B-Text-LMStudio.gguf GGUF 1.23 GB Download
mmproj-MonkeyOCR-pro-1.2B-Vision-LMStudio.gguf GGUF 1.25 GB Download

Model Details Live

Model Slug
dinhquangson/monkeyocr-pro-1.2b-vision-gguf
Author
dinhquangson
Pipeline Task
image-text-to-text
Library
Created
2025-09-12
Last Modified
2025-09-12
Gated
No
Private
No
HF SHA
b86e2addc855f5f586105a3265fa2e35693ef417
License
Unknown
Language
Unknown
Base Model
Unknown

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "license": "apache-2.0",
    "tags": [
      "vision",
      "ocr",
      "document-analysis",
      "multimodal",
      "qwen2vl",
      "gguf",
      "lm-studio"
    ],
    "model_type": "vision-language",
    "architecture": "qwen2vl",
    "quantization": "q8_0",
    "languages": [
      "en",
      "zh"
    ],
    "pipeline_tag": "image-text-to-text",
    "frontmatter": {},
    "hero_image_url": "",
    "summary": "A high-performance **vision-language model** specialized for **Optical Character Recognition (OCR)** and **document analysis**. This repository contains GGUF format models optimized for both general use and **LM Studio** compatibility.",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\r\nlicense: apache-2.0\r\ntags:\r\n- vision\r\n- ocr\r\n- document-analysis\r\n- multimodal\r\n- qwen2vl\r\n- gguf\r\n- lm-studio\r\nmodel_type: vision-language\r\narchitecture: qwen2vl\r\nquantization: q8_0\r\nlanguages:\r\n- en\r\n- zh\r\npipeline_tag: image-text-to-text\r\n---\r\n\r\n# MonkeyOCR-pro-1.2B Vision GGUF\r\n\r\nA high-performance **vision-language model** specialized for **Optical Character Recognition (OCR)** and **document analysis**. This repository contains GGUF format models optimized for both general use and **LM Studio** compatibility.\r\n\r\n## 🎯 Model Capabilities\r\n\r\n- ✅ **Vision-Language Processing**: Understand and process images with text\r\n- ✅ **OCR (Optical Character Recognition)**: Extract text from images and documents  \r\n- ✅ **Document Structure Analysis**: Analyze layout, tables, and formatting\r\n- ✅ **Multi-turn Conversations**: Chat about images and documents\r\n- ✅ **Multiple Languages**: Support for English, Chinese, and more\r\n\r\n## 📁 Files Included\r\n\r\n### Standard GGUF Models\r\n- **`MonkeyOCR-pro-1.2B-Recognition.gguf`** (1.26 GB) - Complete vision-enabled model\r\n- **`MonkeyOCR-pro-1.2B.gguf`** (1.26 GB) - Alternative version\r\n\r\n### LM Studio Optimized\r\n- **`MonkeyOCR-pro-1.2B-Text-LMStudio.gguf`** (1.23 GB) - Text model for LM Studio\r\n- **`mmproj-MonkeyOCR-pro-1.2B-Vision-LMStudio.gguf`** (1.25 GB) - Vision projection for LM Studio\r\n\r\n## 🎨 LM Studio Setup (Recommended)\r\n\r\n### Step 1: Download Files\r\nDownload both LM Studio files:\r\n- `MonkeyOCR-pro-1.2B-Text-LMStudio.gguf`\r\n- `mmproj-MonkeyOCR-pro-1.2B-Vision-LMStudio.gguf`\r\n\r\n### Step 2: Install in LM Studio\r\n1. Copy both files to your LM Studio models directory\r\n2. Keep them in the same folder\r\n3. Open LM Studio and load the text model: `MonkeyOCR-pro-1.2B-Text-LMStudio.gguf`\r\n4. LM Studio will automatically detect and load the vision projection\r\n\r\n### Step 3: Enable Vision\r\n1. Start a new chat session\r\n2. Look for the camera/image icon in the chat interface  \r\n3. Upload an image and start asking questions!\r\n\r\n## 🔧 Usage Examples\r\n\r\n### OCR Tasks\r\n```\r\nUser: [uploads image] Extract all text from this document\r\nAssistant: I can see text in the image. Here's what I extracted:\r\n[Provides accurate OCR results]\r\n```\r\n\r\n### Document Analysis\r\n```\r\nUser: [uploads form] Analyze the structure of this form\r\nAssistant: This appears to be a [form type] with the following structure:\r\n- Header section with title\r\n- Main content area with fields for...\r\n- Footer with signature area\r\n```\r\n\r\n### Vision Q&A\r\n```\r\nUser: [uploads chart] What does this chart show?\r\nAssistant: This chart displays [detailed analysis of the visual content]\r\n```\r\n\r\n## ⚙️ Technical Specifications\r\n\r\n- **Architecture**: Qwen2.5-VL (qwen2vl)\r\n- **Model Size**: 1.2B parameters\r\n- **Quantization**: Q8_0 (8-bit)\r\n- **Context Length**: 8,196 tokens\r\n- **Vision Encoder**: CLIP-based projection\r\n- **Input Support**: Images, text, multimodal conversations\r\n\r\n## 🔧 Configuration Details\r\n\r\nThe model includes proper vision token configuration:\r\n- Image Token ID: 151655 (`<|image_pad|>`)\r\n- Video Token ID: 151656 (`<|video_pad|>`)\r\n- Vision Start: 151652 (`<|vision_start|>`)\r\n- Vision End: 151653 (`<|vision_end|>`)\r\n\r\n## 🚀 Performance\r\n\r\n- **Fast inference** with Q8_0 quantization\r\n- **High accuracy** OCR capabilities\r\n- **Efficient memory usage** (~1.2-1.3 GB VRAM)\r\n- **Multi-language support** for global documents\r\n\r\n## 📝 License\r\n\r\nApache 2.0 - Free for commercial and research use\r\n\r\n## 🙏 Credits\r\n\r\nBased on the MonkeyOCR model architecture with optimizations for GGUF format and LM Studio compatibility.\r\n\r\n---\r\n\r\n**For best vision performance, use the LM Studio optimized files!** 🎯\r\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "qwen2_vl",
    "vision",
    "ocr",
    "document-analysis",
    "multimodal",
    "qwen2vl",
    "lm-studio",
    "image-text-to-text",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 3,
  "downloads": 383,
  "gated": false,
  "private": false,
  "last_modified": "2025-09-12T07:44:29.000Z",
  "created_at": "2025-09-12T04:27:34.000Z",
  "pipeline_tag": "image-text-to-text",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "68c3a136299095c16a1c8835",
  "id": "dinhquangson/MonkeyOCR-pro-1.2B-Vision-GGUF",
  "modelId": "dinhquangson/MonkeyOCR-pro-1.2B-Vision-GGUF",
  "sha": "b86e2addc855f5f586105a3265fa2e35693ef417",
  "createdAt": "2025-09-12T04:27:34.000Z",
  "lastModified": "2025-09-12T07:44:29.000Z",
  "author": "dinhquangson",
  "downloads": 383,
  "likes": 3,
  "gated": false,
  "private": false,
  "pipeline_tag": "image-text-to-text",
  "library_name": "",
  "siblings_count": 8
}