dinhquangson/monkeyocr-pro-1.2b-vision-gguf Recognition GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
dinhquangson/monkeyocr-pro-1.2b-vision-gguf overview
A high-performance vision-language model specialized for Optical Character Recognition (OCR) and document analysis. This repository contains GGUF format models optimized for both general use and LM Studio compatibility.
Downloads
383
Likes
3
Pipeline
image-text-to-text
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"license": "apache-2.0",
"tags": [
"vision",
"ocr",
"document-analysis",
"multimodal",
"qwen2vl",
"gguf",
"lm-studio"
],
"model_type": "vision-language",
"architecture": "qwen2vl",
"quantization": "q8_0",
"languages": [
"en",
"zh"
],
"pipeline_tag": "image-text-to-text",
"frontmatter": {},
"hero_image_url": "",
"summary": "A high-performance **vision-language model** specialized for **Optical Character Recognition (OCR)** and **document analysis**. This repository contains GGUF format models optimized for both general use and **LM Studio** compatibility.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\r\nlicense: apache-2.0\r\ntags:\r\n- vision\r\n- ocr\r\n- document-analysis\r\n- multimodal\r\n- qwen2vl\r\n- gguf\r\n- lm-studio\r\nmodel_type: vision-language\r\narchitecture: qwen2vl\r\nquantization: q8_0\r\nlanguages:\r\n- en\r\n- zh\r\npipeline_tag: image-text-to-text\r\n---\r\n\r\n# MonkeyOCR-pro-1.2B Vision GGUF\r\n\r\nA high-performance **vision-language model** specialized for **Optical Character Recognition (OCR)** and **document analysis**. This repository contains GGUF format models optimized for both general use and **LM Studio** compatibility.\r\n\r\n## 🎯 Model Capabilities\r\n\r\n- ✅ **Vision-Language Processing**: Understand and process images with text\r\n- ✅ **OCR (Optical Character Recognition)**: Extract text from images and documents \r\n- ✅ **Document Structure Analysis**: Analyze layout, tables, and formatting\r\n- ✅ **Multi-turn Conversations**: Chat about images and documents\r\n- ✅ **Multiple Languages**: Support for English, Chinese, and more\r\n\r\n## 📁 Files Included\r\n\r\n### Standard GGUF Models\r\n- **`MonkeyOCR-pro-1.2B-Recognition.gguf`** (1.26 GB) - Complete vision-enabled model\r\n- **`MonkeyOCR-pro-1.2B.gguf`** (1.26 GB) - Alternative version\r\n\r\n### LM Studio Optimized\r\n- **`MonkeyOCR-pro-1.2B-Text-LMStudio.gguf`** (1.23 GB) - Text model for LM Studio\r\n- **`mmproj-MonkeyOCR-pro-1.2B-Vision-LMStudio.gguf`** (1.25 GB) - Vision projection for LM Studio\r\n\r\n## 🎨 LM Studio Setup (Recommended)\r\n\r\n### Step 1: Download Files\r\nDownload both LM Studio files:\r\n- `MonkeyOCR-pro-1.2B-Text-LMStudio.gguf`\r\n- `mmproj-MonkeyOCR-pro-1.2B-Vision-LMStudio.gguf`\r\n\r\n### Step 2: Install in LM Studio\r\n1. Copy both files to your LM Studio models directory\r\n2. Keep them in the same folder\r\n3. Open LM Studio and load the text model: `MonkeyOCR-pro-1.2B-Text-LMStudio.gguf`\r\n4. LM Studio will automatically detect and load the vision projection\r\n\r\n### Step 3: Enable Vision\r\n1. Start a new chat session\r\n2. Look for the camera/image icon in the chat interface \r\n3. Upload an image and start asking questions!\r\n\r\n## 🔧 Usage Examples\r\n\r\n### OCR Tasks\r\n```\r\nUser: [uploads image] Extract all text from this document\r\nAssistant: I can see text in the image. Here's what I extracted:\r\n[Provides accurate OCR results]\r\n```\r\n\r\n### Document Analysis\r\n```\r\nUser: [uploads form] Analyze the structure of this form\r\nAssistant: This appears to be a [form type] with the following structure:\r\n- Header section with title\r\n- Main content area with fields for...\r\n- Footer with signature area\r\n```\r\n\r\n### Vision Q&A\r\n```\r\nUser: [uploads chart] What does this chart show?\r\nAssistant: This chart displays [detailed analysis of the visual content]\r\n```\r\n\r\n## ⚙️ Technical Specifications\r\n\r\n- **Architecture**: Qwen2.5-VL (qwen2vl)\r\n- **Model Size**: 1.2B parameters\r\n- **Quantization**: Q8_0 (8-bit)\r\n- **Context Length**: 8,196 tokens\r\n- **Vision Encoder**: CLIP-based projection\r\n- **Input Support**: Images, text, multimodal conversations\r\n\r\n## 🔧 Configuration Details\r\n\r\nThe model includes proper vision token configuration:\r\n- Image Token ID: 151655 (`<|image_pad|>`)\r\n- Video Token ID: 151656 (`<|video_pad|>`)\r\n- Vision Start: 151652 (`<|vision_start|>`)\r\n- Vision End: 151653 (`<|vision_end|>`)\r\n\r\n## 🚀 Performance\r\n\r\n- **Fast inference** with Q8_0 quantization\r\n- **High accuracy** OCR capabilities\r\n- **Efficient memory usage** (~1.2-1.3 GB VRAM)\r\n- **Multi-language support** for global documents\r\n\r\n## 📝 License\r\n\r\nApache 2.0 - Free for commercial and research use\r\n\r\n## 🙏 Credits\r\n\r\nBased on the MonkeyOCR model architecture with optimizations for GGUF format and LM Studio compatibility.\r\n\r\n---\r\n\r\n**For best vision performance, use the LM Studio optimized files!** 🎯\r\n",
"related_quantizations": []
},
"tags": [
"gguf",
"qwen2_vl",
"vision",
"ocr",
"document-analysis",
"multimodal",
"qwen2vl",
"lm-studio",
"image-text-to-text",
"license:apache-2.0",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 3,
"downloads": 383,
"gated": false,
"private": false,
"last_modified": "2025-09-12T07:44:29.000Z",
"created_at": "2025-09-12T04:27:34.000Z",
"pipeline_tag": "image-text-to-text",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "68c3a136299095c16a1c8835",
"id": "dinhquangson/MonkeyOCR-pro-1.2B-Vision-GGUF",
"modelId": "dinhquangson/MonkeyOCR-pro-1.2B-Vision-GGUF",
"sha": "b86e2addc855f5f586105a3265fa2e35693ef417",
"createdAt": "2025-09-12T04:27:34.000Z",
"lastModified": "2025-09-12T07:44:29.000Z",
"author": "dinhquangson",
"downloads": 383,
"likes": 3,
"gated": false,
"private": false,
"pipeline_tag": "image-text-to-text",
"library_name": "",
"siblings_count": 8
}