EryriLabs/LFM2.5-VL-3B-DragOn-GGUF overview
LFM2.5 VL 3B DragOn — GGUF GGUF quants of EryriLabs/LFM2.5 VL 3B DragOn https://huggingface.co/EryriLabs/LFM2.5 VL 3B DragOn : LiquidAI's LFM2.5 VL 3B fine tun…
Runs locally from ~814.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| LFM2.5-VL-3B-DragOn-F16.gguf | GGUF | F16 | 5.03 GB | Download |
| LFM2.5-VL-3B-DragOn-Q4_K_M.gguf | GGUF | Q4_K_M | 1.56 GB | Download |
| LFM2.5-VL-3B-DragOn-Q5_K_M.gguf | GGUF | Q5_K_M | 1.81 GB | Download |
| LFM2.5-VL-3B-DragOn-Q6_K.gguf | GGUF | Q6_K | 2.07 GB | Download |
| LFM2.5-VL-3B-DragOn-Q8_0.gguf | GGUF | Q8_0 | 2.68 GB | Download |
| mmproj-LFM2.5-VL-3B-DragOn-F16.gguf | GGUF | F16 | 814.4 MB | Download |
Model Details
Model README
---
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-VL-3B/blob/main/LICENSE
base_model: EryriLabs/LFM2.5-VL-3B-DragOn
datasets:
- Hcompany/DragOn
tags:
- gguf
- llama.cpp
- gui-agents
- computer-use
- drag-and-drop
- grounding
---
LFM2.5-VL-3B-DragOn — GGUF
GGUF quants of EryriLabs/LFM2.5-VL-3B-DragOn: LiquidAI's LFM2.5-VL-3B fine-tuned for drag-and-drop grounding on GUI screenshots. Screenshot + instruction in, {"start":[x,y],"end":[x,y]} out (0-1000 normalised coordinates).
The short version of the story: the base model scores 0.7% on the DragOn public eval, this fine-tune scores 70.8% acc@5 (78.5% acc@10), and it cost about $36 to train. Details, per-domain numbers and caveats are in the main repo.
Files
You need TWO files: a main model quant plus the vision projector (mmproj).
| file | size | note |
|---|---|---|
| LFM2.5-VL-3B-DragOn-Q4_K_M.gguf | ~1.5 GB | good default, runs on almost anything |
| LFM2.5-VL-3B-DragOn-Q5_K_M.gguf | ~1.8 GB | |
| LFM2.5-VL-3B-DragOn-Q6_K.gguf | ~2.0 GB | recommended if you have the room |
| LFM2.5-VL-3B-DragOn-Q8_0.gguf | ~2.7 GB | |
| LFM2.5-VL-3B-DragOn-F16.gguf | ~5.1 GB | reference |
| mmproj-LFM2.5-VL-3B-DragOn-F16.gguf | ~0.8 GB | vision projector, always required |
Note that coordinates are a precision task, so if you see degraded accuracy at Q4, step up a quant before blaming the model.
Usage
llama-server -m LFM2.5-VL-3B-DragOn-Q6_K.gguf \
--mmproj mmproj-LFM2.5-VL-3B-DragOn-F16.gguf \
-ngl 99 -c 4096 --temp 0
Then send a chat completion with the image and this exact prompt shape (it's what the model was trained on):
This is a screenshot of a user interface. You must perform a DRAG action.
Task: <your instruction>
Give the drag as JSON with the START point (where the mouse button goes down) and the END point (where it is released), in coordinates normalised to 0-1000 for both x (left->right) and y (top->bottom):
{"start":[x,y],"end":[x,y]}
Output only the JSON.
Multiply by your actual screen size /1000 and you have your drag.
Thanks
To LiquidAI for the base model and to Nathan Bout, Maxime Langevin and Ronan Riochet at Hcompany for the DragOn dataset — see the main repo card for the full credits.
Quantised by Dwain Barnes (EryriLabs), August 2026.
Run EryriLabs/LFM2.5-VL-3B-DragOn-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models