GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

prithivMLmods/LensVLM-9B-GGUF overview

LensVLM 9B GGUF LensVLM 9B is a 9 billion parameter vision language model from Apple, built on Qwen3.5 9B, introduced in the paper "LensVLM: Selective Context …

transformersgguftext-generation-inferencellama-cppvision-language-modellong-contextvisual-text-compressionimage-text-to-textenbase_model:apple/LensVLM-9Bbase_model:quantized:apple/LensVLM-9Blicense:apple-amlrendpoints_compatibleregion:usconversational

Runs locally from ~879.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
image-text-to-text

Repository Files & Downloads

9 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LensVLM-9B.BF16.ggufGGUFGGUF16.69 GBDownload
LensVLM-9B.Q3_K_L.ggufGGUFGGUF4.59 GBDownload
LensVLM-9B.Q3_K_M.ggufGGUFGGUF4.31 GBDownload
LensVLM-9B.Q4_K_M.ggufGGUFGGUF5.24 GBDownload
LensVLM-9B.Q4_K_S.ggufGGUFGGUF4.98 GBDownload
LensVLM-9B.Q5_K_M.ggufGGUFGGUF6.02 GBDownload
LensVLM-9B.Q5_K_S.ggufGGUFGGUF5.87 GBDownload
LensVLM-9B.Q6_K.ggufGGUFGGUF6.85 GBDownload
LensVLM-9B.mmproj-bf16.ggufGGUFBF16879.0 MBDownload

Model Details

Model IDprithivMLmods/LensVLM-9B-GGUF
AuthorprithivMLmods
Pipelineimage-text-to-text
Licenseapple-amlr
Base modelapple/LensVLM-9B
Last modified2026-09-23T05:44:22.000Z

Model README

---

license: apple-amlr

license_link: https://huggingface.co/apple/LensVLM-9B/blob/main/LICENSE

library_name: transformers

base_model:

  • apple/LensVLM-9B

tags:

  • text-generation-inference
  • llama-cpp
  • vision-language-model
  • long-context
  • visual-text-compression

language:

  • en

pipeline_tag: image-text-to-text

---

LensVLM-9B-GGUF

> LensVLM-9B is a 9-billion-parameter vision-language model from Apple, built on Qwen3.5-9B, introduced in the paper "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text." Its core mechanism scans a compressed image representation of text — at configurable compression ratios of 5x, 10x, or 15x — and then selectively expands only the pages relevant to a given question back to their uncompressed form via learned tools, allowing the model to process very long documents without holding the entire uncompressed text in context. It's run via the accompanying ml-lensvlm codebase with a simple demo script accepting a text file and a question, and is released under the Apple Machine Learning Research Model License (with the accompanying source code separately licensed under the Apple Sample Code License).

Model Files

| File Name | Quant Type | File Size | File Link | Description |

|-----------|------------|-----------|-----------|-------------|

| LensVLM-9B.BF16.gguf | BF16 | 17.9 GB | Link | Full BF16 weights. Highest quality, largest file size. |

| LensVLM-9B.Q3_K_L.gguf | Q3_K_L | 4.93 GB | Link | Lower quality but usable, good for low RAM availability. |

| LensVLM-9B.Q3_K_M.gguf | Q3_K_M | 4.62 GB | Link | Low quality. |

| LensVLM-9B.Q4_K_M.gguf | Q4_K_M | 5.63 GB | Link | Good quality, default size for most use cases, recommended. |

| LensVLM-9B.Q4_K_S.gguf | Q4_K_S | 5.35 GB | Link | Slightly lower quality with more space savings, recommended. |

| LensVLM-9B.Q5_K_M.gguf | Q5_K_M | 6.47 GB | Link | High quality, recommended. |

| LensVLM-9B.Q5_K_S.gguf | Q5_K_S | 6.31 GB | Link | High quality, recommended. |

| LensVLM-9B.Q6_K.gguf | Q6_K | 7.36 GB | Link | Very high quality, near perfect, recommended. |

| LensVLM-9B.mmproj-bf16.gguf | mmproj-bf16 | 922 MB | Link | Multimodal projection file in BF16 format. Used for vision/language models. |

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Run prithivMLmods/LensVLM-9B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models