prithivMLmods/SFT-4B-ScaleCUA-MixedOpen-GGUF overview
SFT 4B ScaleCUA MixedOpen GGUF SFT 4B ScaleCUA MixedOpen https://huggingface.co/HaoranLiu/SFT 4B ScaleCUA MixedOpen is a Qwen3 VL 4B Instruct checkpoint superv…
Runs locally from ~800.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| SFT-4B-ScaleCUA-MixedOpen.BF16.gguf | GGUF | GGUF | 7.50 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.Q3_K_L.gguf | GGUF | GGUF | 2.09 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.Q3_K_M.gguf | GGUF | GGUF | 1.93 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.Q4_K_M.gguf | GGUF | GGUF | 2.33 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.Q4_K_S.gguf | GGUF | GGUF | 2.22 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.Q5_K_M.gguf | GGUF | GGUF | 2.69 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.Q5_K_S.gguf | GGUF | GGUF | 2.63 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.Q6_K.gguf | GGUF | GGUF | 3.08 GB | Download |
| SFT-4B-ScaleCUA-MixedOpen.mmproj-bf16.gguf | GGUF | BF16 | 800.4 MB | Download |
Model Details
| Model ID | prithivMLmods/SFT-4B-ScaleCUA-MixedOpen-GGUF |
|---|---|
| Author | prithivMLmods |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | HaoranLiu/SFT-4B-ScaleCUA-MixedOpen |
| Last modified | 2026-09-21T07:24:14.000Z |
Model README
---
license: apache-2.0
base_model:
- HaoranLiu/SFT-4B-ScaleCUA-MixedOpen
library_name: transformers
tags:
- text-generation-inference
- llama-cp
- computer-use
- gui-agent
- sft
- qwen3-vl
language:
- en
pipeline_tag: image-text-to-text
---
SFT-4B-ScaleCUA-MixedOpen-GGUF
> SFT-4B-ScaleCUA-MixedOpen is a Qwen3-VL-4B-Instruct checkpoint supervised-finetuned on 1,430 Lite.ScaleCUA trajectories pooled in equal proportion from three open-source teacher agents — Qwen3.8-27B (478 trajectories), Qwen3.5-27B (477), and EvoCUA-8B-20260105 (475), one trajectory per task, 1,390 of which are full successes — trained via cua-lite + slime with token-level SFT on assistant-action tokens over 3 epochs (1,070 steps) on 2x 80GB H100 GPUs. This epoch-3 checkpoint (iter_1070) is the strongest of the measured ScaleCUA SFT arms on the Lite.OSWorld eval split (332 tasks, greedy decoding, one host/protocol), reaching a mean episode return of 0.3927 (125/332 success) — beating the best single-teacher arm (Qwen3.5-27B, 0.3682) by +0.0245/+8 tasks and the closest-matched single-teacher arm (Qwen38, 0.3623) by +0.0304/+10 tasks, both clearing the paper's >0.02 mean and ≥7 task threshold for a real effect, though the authors caution this isn't a clean single-variable ablation since teacher identity, trajectory count, and task coverage all move together. Notably, doubling trajectories per task (-MixedOpen-cap2, 2,520 trajectories) actually underperforms this model by −0.0125/−5 tasks, suggesting broader task coverage matters more than redundant per-task examples, and per-domain results show multi_apps coordination (28% of tasks) remains the weakest capability across all teacher mixes at just 0.159 mean return. The model requires the qwen3_vl adapter config with full_history_size=4 for correct serving, since it was trained on that specific history-rendering protocol.
Model Files
| File Name | Quant Type | File Size | File Link | Description |
|-----------|------------|-----------|-----------|-------------|
| SFT-4B-ScaleCUA-MixedOpen.BF16.gguf | BF16 | 8.05 GB | Link | Full BF16 weights. Highest quality, largest file size. |
| SFT-4B-ScaleCUA-MixedOpen.Q3_K_L.gguf | Q3_K_L | 2.24 GB | Link | Lower quality but usable, good for low RAM availability. |
| SFT-4B-ScaleCUA-MixedOpen.Q3_K_M.gguf | Q3_K_M | 2.08 GB | Link | Low quality. |
| SFT-4B-ScaleCUA-MixedOpen.Q4_K_M.gguf | Q4_K_M | 2.5 GB | Link | Good quality, default size for most use cases, recommended. |
| SFT-4B-ScaleCUA-MixedOpen.Q4_K_S.gguf | Q4_K_S | 2.38 GB | Link | Slightly lower quality with more space savings, recommended. |
| SFT-4B-ScaleCUA-MixedOpen.Q5_K_M.gguf | Q5_K_M | 2.89 GB | Link | High quality, recommended. |
| SFT-4B-ScaleCUA-MixedOpen.Q5_K_S.gguf | Q5_K_S | 2.82 GB | Link | High quality, recommended. |
| SFT-4B-ScaleCUA-MixedOpen.Q6_K.gguf | Q6_K | 3.31 GB | Link | Very high quality, near perfect, recommended. |
| SFT-4B-ScaleCUA-MixedOpen.mmproj-bf16.gguf | mmproj-bf16 | 839 MB | Link | Multimodal projection file in BF16 format. Used for vision/language models. |
llama.cpp
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp
Run prithivMLmods/SFT-4B-ScaleCUA-MixedOpen-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models