GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

cmndcntrlcyber/qwen14b-code-trainer-gguf overview

qwen14b code trainer gguf GGUF quantizations of the Code Trainer fine tuned model. The full adapter chain — DAPT qwen14b dapt offsec https://huggingface.co/cmn…

ggufllama-cppquantizedcode-generationtool-callingqwen2.5-codercode-trainertext-generationbase_model:Qwen/Qwen2.5-Coder-14B-Instructbase_model:quantized:Qwen/Qwen2.5-Coder-14B-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~8.37 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
637
Likes
0
Pipeline
text-generation

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen2.5-Coder-14B-Instruct-Q4_K_M.ggufGGUFQ4_K_M8.37 GBDownload
Qwen2.5-Coder-14B-Instruct-Q5_K_M.ggufGGUFQ5_K_M9.79 GBDownload

Model Details

Model IDcmndcntrlcyber/qwen14b-code-trainer-gguf
Authorcmndcntrlcyber
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen2.5-Coder-14B-Instruct
Last modified2026-08-29T11:24:38.000Z

Model README

---

base_model: Qwen/Qwen2.5-Coder-14B-Instruct

license: apache-2.0

tags:

  • gguf
  • llama-cpp
  • quantized
  • code-generation
  • tool-calling
  • qwen2.5-coder
  • code-trainer

pipeline_tag: text-generation

---

qwen14b-code-trainer-gguf

GGUF quantizations of the Code-Trainer fine-tuned model. The full adapter chain

— DAPT (qwen14b-dapt-offsec),

V9 SFT (qwen14b-code-trainer-v9_mixed),

and V10 GRPO (qwen14b-code-trainer-v10-grpo)

— is merged into

Qwen/Qwen2.5-Coder-14B-Instruct

and quantized via llama.cpp.

This is Phase 5 of the

Code-Trainer / RTPI

pipeline. The conversion runs as an HF Job on a100-large — the GPU sits

idle, we use that flavor only for its 144 GB system RAM during the float16

merge step.

Files

| File | Quantization | Size (≈) | Notes |

|---|---|---|---|

| Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf | Q5_K_M | ~10.5 GB | Recommended default (V9+) — preserves <tool_call> tag fidelity |

| Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf | Q4_K_M | ~9 GB | Fallback — balanced quality / footprint |

Additional quantizations (Q8_0, F16) can be produced by passing

--quants to launch_convert.py.

Intended use

  • Local inference via llama-cli, llama-server, Ollama, LM Studio, or

text-generation-webui.

  • Phase 6 hot-swap target for the project's vLLM + Qwen-Agent stack —

swapped in for compiled-language tasks alongside a smaller primary model.

  • Out of scope: anything the upstream

qwen14b-code-trainer-aggressive

card flags as out of scope (no safety tuning, no non-code tasks).

Source

The GGUF is produced by merging the full adapter chain in order, then

quantizing the merged model:

Qwen/Qwen2.5-Coder-14B-Instruct
  → merge DAPT LoRA    (qwen14b-dapt-offsec)
  → merge V9 SFT LoRA  (qwen14b-code-trainer-v9_mixed)
  → merge V10 GRPO LoRA (qwen14b-code-trainer-v10-grpo)
  → convert_hf_to_gguf.py + llama-quantize → Q5_K_M

| Stage | Repo / artifact |

|---|---|

| Base model | Qwen/Qwen2.5-Coder-14B-Instruct |

| DAPT adapter | cmndcntrlcyber/qwen14b-dapt-offsec |

| SFT adapter (V9) | cmndcntrlcyber/qwen14b-code-trainer-v9_mixed |

| GRPO adapter (V10) | cmndcntrlcyber/qwen14b-code-trainer-v10-grpo |

| Converter | llama.cpp (convert_hf_to_gguf.py + llama-quantize) |

| Conversion runtime | HF Job, a100-large, ~1 h on the merge + quantize path |

Evaluation

Quality is inherited from the source adapter chain. The current source is V10

GRPO — an RL-tuned adapter trained on a rule-based tool-call formatting reward

(mean reward ~0.14, see the

V10 model card).

The SFT foundation is V9 (40,401-row curriculum dataset, 64.3% tool coverage —

see the V9 model card).

Quantization to Q5_K_M typically introduces minimal perplexity penalty

(< 1 %) for 14 B models; Q4_K_M introduces ~1–3 %.

Quick start

llama-server

llama-server \
  -m Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  --ctx-size 8192 --n-gpu-layers 999

Ollama Modelfile

FROM ./Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ range .Messages }}{{ if eq .Role "user" }}<|im_start|>user
{{ .Content }}<|im_end|>
{{ else if eq .Role "assistant" }}<|im_start|>assistant
{{ .Content }}<|im_end|>
{{ else if eq .Role "tool" }}<|im_start|>tool
{{ .Content }}<|im_end|>
{{ end }}{{ end }}<|im_start|>assistant
"""
PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER num_ctx 8192

llama-cpp-python

from llama_cpp import Llama

llm = Llama(
    model_path="Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf",
    n_ctx=8192,
    n_gpu_layers=999,
)
print(llm.create_chat_completion(messages=[
    {"role": "user", "content": "Write a Go function that reverses a UTF-8 string."},
])["choices"][0]["message"]["content"])

Limitations

  • Lossy quantization. Q4_K_M is a 4-bit-mixed format; expect minor

degradation vs. the unquantized adapter on long-form code. Q5_K_M is

recommended for tool-calling workloads.

  • No safety tuning. Inherits all caveats from the source adapter.
  • Two quants shipped. Q5_K_M (recommended) and Q4_K_M (fallback).

For Q8_0 / F16, regenerate with

python -m src.phase5_deployment.scripts.launch_convert --quants Q8_0.

Reproducibility

set -a && source .env && set +a
python -m src.phase5_deployment.scripts.launch_convert \
    --config src/config/config.yaml --wait

(src/phase5_deployment/)

  • Config: src/config/pipeline-50.yml (deployment section)
  • Cost: ~$2 on a100-large once the job runs.

Run cmndcntrlcyber/qwen14b-code-trainer-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models