GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

prithivMLmods/SFT-4B-ScaleCUA-MixedOpen-GGUF overview

SFT 4B ScaleCUA MixedOpen GGUF SFT 4B ScaleCUA MixedOpen https://huggingface.co/HaoranLiu/SFT 4B ScaleCUA MixedOpen is a Qwen3 VL 4B Instruct checkpoint superv…

transformersgguftext-generation-inferencellama-cpcomputer-usegui-agentsftqwen3-vlimage-text-to-textenbase_model:HaoranLiu/SFT-4B-ScaleCUA-MixedOpenbase_model:quantized:HaoranLiu/SFT-4B-ScaleCUA-MixedOpenlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~800.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
image-text-to-text

Repository Files & Downloads

9 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
SFT-4B-ScaleCUA-MixedOpen.BF16.ggufGGUFGGUF7.50 GBDownload
SFT-4B-ScaleCUA-MixedOpen.Q3_K_L.ggufGGUFGGUF2.09 GBDownload
SFT-4B-ScaleCUA-MixedOpen.Q3_K_M.ggufGGUFGGUF1.93 GBDownload
SFT-4B-ScaleCUA-MixedOpen.Q4_K_M.ggufGGUFGGUF2.33 GBDownload
SFT-4B-ScaleCUA-MixedOpen.Q4_K_S.ggufGGUFGGUF2.22 GBDownload
SFT-4B-ScaleCUA-MixedOpen.Q5_K_M.ggufGGUFGGUF2.69 GBDownload
SFT-4B-ScaleCUA-MixedOpen.Q5_K_S.ggufGGUFGGUF2.63 GBDownload
SFT-4B-ScaleCUA-MixedOpen.Q6_K.ggufGGUFGGUF3.08 GBDownload
SFT-4B-ScaleCUA-MixedOpen.mmproj-bf16.ggufGGUFBF16800.4 MBDownload

Model Details

Model IDprithivMLmods/SFT-4B-ScaleCUA-MixedOpen-GGUF
AuthorprithivMLmods
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelHaoranLiu/SFT-4B-ScaleCUA-MixedOpen
Last modified2026-09-21T07:24:14.000Z

Model README

---

license: apache-2.0

base_model:

  • HaoranLiu/SFT-4B-ScaleCUA-MixedOpen

library_name: transformers

tags:

  • text-generation-inference
  • llama-cp
  • computer-use
  • gui-agent
  • sft
  • qwen3-vl

language:

  • en

pipeline_tag: image-text-to-text

---

SFT-4B-ScaleCUA-MixedOpen-GGUF

> SFT-4B-ScaleCUA-MixedOpen is a Qwen3-VL-4B-Instruct checkpoint supervised-finetuned on 1,430 Lite.ScaleCUA trajectories pooled in equal proportion from three open-source teacher agents — Qwen3.8-27B (478 trajectories), Qwen3.5-27B (477), and EvoCUA-8B-20260105 (475), one trajectory per task, 1,390 of which are full successes — trained via cua-lite + slime with token-level SFT on assistant-action tokens over 3 epochs (1,070 steps) on 2x 80GB H100 GPUs. This epoch-3 checkpoint (iter_1070) is the strongest of the measured ScaleCUA SFT arms on the Lite.OSWorld eval split (332 tasks, greedy decoding, one host/protocol), reaching a mean episode return of 0.3927 (125/332 success) — beating the best single-teacher arm (Qwen3.5-27B, 0.3682) by +0.0245/+8 tasks and the closest-matched single-teacher arm (Qwen38, 0.3623) by +0.0304/+10 tasks, both clearing the paper's >0.02 mean and ≥7 task threshold for a real effect, though the authors caution this isn't a clean single-variable ablation since teacher identity, trajectory count, and task coverage all move together. Notably, doubling trajectories per task (-MixedOpen-cap2, 2,520 trajectories) actually underperforms this model by −0.0125/−5 tasks, suggesting broader task coverage matters more than redundant per-task examples, and per-domain results show multi_apps coordination (28% of tasks) remains the weakest capability across all teacher mixes at just 0.159 mean return. The model requires the qwen3_vl adapter config with full_history_size=4 for correct serving, since it was trained on that specific history-rendering protocol.

Model Files

| File Name | Quant Type | File Size | File Link | Description |

|-----------|------------|-----------|-----------|-------------|

| SFT-4B-ScaleCUA-MixedOpen.BF16.gguf | BF16 | 8.05 GB | Link | Full BF16 weights. Highest quality, largest file size. |

| SFT-4B-ScaleCUA-MixedOpen.Q3_K_L.gguf | Q3_K_L | 2.24 GB | Link | Lower quality but usable, good for low RAM availability. |

| SFT-4B-ScaleCUA-MixedOpen.Q3_K_M.gguf | Q3_K_M | 2.08 GB | Link | Low quality. |

| SFT-4B-ScaleCUA-MixedOpen.Q4_K_M.gguf | Q4_K_M | 2.5 GB | Link | Good quality, default size for most use cases, recommended. |

| SFT-4B-ScaleCUA-MixedOpen.Q4_K_S.gguf | Q4_K_S | 2.38 GB | Link | Slightly lower quality with more space savings, recommended. |

| SFT-4B-ScaleCUA-MixedOpen.Q5_K_M.gguf | Q5_K_M | 2.89 GB | Link | High quality, recommended. |

| SFT-4B-ScaleCUA-MixedOpen.Q5_K_S.gguf | Q5_K_S | 2.82 GB | Link | High quality, recommended. |

| SFT-4B-ScaleCUA-MixedOpen.Q6_K.gguf | Q6_K | 3.31 GB | Link | Very high quality, near perfect, recommended. |

| SFT-4B-ScaleCUA-MixedOpen.mmproj-bf16.gguf | mmproj-bf16 | 839 MB | Link | Multimodal projection file in BF16 format. Used for vision/language models. |

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Run prithivMLmods/SFT-4B-ScaleCUA-MixedOpen-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models