GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

shafire/OpenZero-Fusion-Qwen3-4B-Agentic-GGUF overview

OpenZero Fusion Qwen3 4B Agentic — Standalone GGUF THIS MODEL WORKS BUT HAS OUTPUT ERROR ISSUES.NOT RECOMMENDED FOR PRODUCTION. ONE FILE. TWO SPECIALISTS DISTI…

llama.cppggufqwen3local-aiagentic-aicoding-assistantresearch-assistanttool-useknowledge-distillationopenzeroollamallama-cppquantizedtext-generationenbase_model:Qwen/Qwen3-4Bbase_model:quantized:Qwen/Qwen3-4Bregion:usconversational

Runs locally from ~2.33 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
OpenZero-Fusion-Qwen3-4B-Agentic-F16.ggufGGUFF167.50 GBDownload
OpenZero-Fusion-Qwen3-4B-Agentic-Q4_K_M.ggufGGUFQ4_K_M2.33 GBDownload
OpenZero-Fusion-Qwen3-4B-Agentic-Q8_0.ggufGGUFQ8_03.99 GBDownload

Model Details

Model IDshafire/OpenZero-Fusion-Qwen3-4B-Agentic-GGUF
Authorshafire
Pipelinetext-generation
License
Base modelQwen/Qwen3-4B
Last modified2026-08-11T23:11:26.000Z

Model README

---

language:

  • en

library_name: llama.cpp

pipeline_tag: text-generation

inference: false

tags:

  • gguf
  • qwen3
  • local-ai
  • agentic-ai
  • coding-assistant
  • research-assistant
  • tool-use
  • knowledge-distillation
  • openzero
  • ollama
  • llama-cpp
  • quantized

base_model: Qwen/Qwen3-4B

---

OpenZero Fusion Qwen3-4B Agentic — Standalone GGUF

THIS MODEL WORKS BUT HAS OUTPUT ERROR ISSUES.NOT RECOMMENDED FOR PRODUCTION.

ONE FILE. TWO SPECIALISTS DISTILLED INTO ONE LOCAL AGENT.

!OpenZero Fusion Qwen3-4B Agentic

OpenZero Fusion Qwen3-4B Agentic is a single-architecture Qwen3-4B student trained from the curated OpenZero corpus plus filtered outputs from the completed Qwen3-1.7B and Gemma4-E2B specialists. It is knowledge distillation—not an invalid direct weight merge between unrelated architectures.

Gemma 4 and Qwen 3 into 1 LLM made for CPU. Experimental model results may vary for this model.

What is included

| File | Purpose |

|---|---|

| OpenZero-Fusion-Qwen3-4B-Agentic-Q4_K_M.gguf | Recommended balance of size and quality |

| OpenZero-Fusion-Qwen3-4B-Agentic-Q8_0.gguf | Higher fidelity, larger download |

| OpenZero-Fusion-Qwen3-4B-Agentic-F16.gguf | Reference full-precision GGUF |

  • Standalone model: yes
  • Separate adapter required: no
  • Separate base model required: no
  • Base architecture: Qwen/Qwen3-4B
  • Training mix: 2,606 curated examples + 379 filtered two-teacher examples = 2,985 training rows
  • Held-out evaluation: 137 rows, never used for training
  • Final held-out loss: 1.643606
  • Held-out token accuracy: 0.788742
  • GGUF conversion and CPU load test: passed for all three files with llama.cpp

Run with llama.cpp

hf download shafire/OpenZero-Fusion-Qwen3-4B-Agentic-GGUF OpenZero-Fusion-Qwen3-4B-Agentic-Q4_K_M.gguf --local-dir .
llama-cli -m OpenZero-Fusion-Qwen3-4B-Agentic-Q4_K_M.gguf --jinja -c 8192 -t 8 --temp 0.6 --top-p 0.95

For a local OpenAI-compatible endpoint:

llama-server -m OpenZero-Fusion-Qwen3-4B-Agentic-Q4_K_M.gguf --jinja -c 8192 -t 8 --host 127.0.0.1 --port 8080

Why Fusion exists

The small Qwen and Gemma specialists cannot safely be tensor-merged: their architectures and tokenizers differ. Fusion instead trains one compatible Qwen3-4B student using useful outputs from both teachers, producing a single deployable GGUF.

Verified release

  • Full QLoRA training: 94/94 steps, one epoch
  • Train loss: 2.158780
  • Held-out evaluation loss: 1.643606
  • Held-out token accuracy: 0.788742
  • Q4_K_M — 2,497,280,288 bytes — SHA-256 f00ad7bb04cb88c6861070e4b6176c9ee208030ae91ea9bb9ebeece7ee9ae415
  • Q8_0 — 4,280,404,768 bytes — SHA-256 8430bcfe6704c53cfa0f8d3e30d141d78730e1d47bcd80515413ec24a090317c
  • F16 — 8,051,284,768 bytes — SHA-256 29069c52929b50681d08cfc2564cf7e443b3b0c286fbf6b8ad1d285f11d11196

Provenance and reproducibility

The student used 2,606 original OpenZero instruction rows plus 379 filtered teacher-response rows, for 2,985 training rows total. Teacher generation began with 192 balanced prompts across general, coding, research and agent/tool categories; both specialists produced a candidate for each prompt. A separate 137-row held-out set was never used for training. The completed V6 adapter was merged only into Qwen/Qwen3-4B revision 1cfa9a7208912126459214e8b04321603b3df60c. F16 was converted with llama.cpp b10333 / commit 08659901c43b51de735740f1cf61bb82fbe0c4e4; Q8_0 and Q4_K_M were independently quantized from that F16 source. Every file passed a bounded one-shot CPU text load test (-c 128 -n 1 -ngl 0 --no-conversation --single-turn --simple-io --no-warmup) and remote LFS/Xet SHA-256 verification.

This distillation transfers examples, not tensors. It does not make a 4B model equivalent to the sum of two architectures, and the reported held-out metrics are not broad benchmark claims. Tool calls are text emitted for an external runtime to validate and execute.

This model is an independent fine-tune based on Qwen and is not affiliated with or endorsed by Qwen. The Qwen base is Apache-2.0. OpenZero training materials and resulting community release are subject to the OpenZero Community Source terms; do not describe this release as OSI-approved open source. Review both upstream and OpenZero terms before redistribution or commercial use.

Run shafire/OpenZero-Fusion-Qwen3-4B-Agentic-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models