GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Jackrong/DeepSeek-V4-Pro-Qwen3.5-9B-MTP-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid cbd5e1; border radius: 16px; box shadow: 0 10px 15…

transformersgguftext-generation-inferenceunslothqwen3_5reasoningdistillationdeepseeksftrlgspomathstemtool-usefunction-callingmtptext-generationenzhkojaesrubase_model:unsloth/Qwen3.5-9B

Runs locally from ~1.70 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
65,676
Likes
77
Pipeline
text-generation
Author

Repository Files & Downloads

13 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-BF16.ggufGGUFBF1617.14 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-IQ4_XS.ggufGGUFIQ4_XS4.99 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q2_K.ggufGGUFQ2_K3.65 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q3_K_L.ggufGGUFQ3_K_L4.70 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q3_K_M.ggufGGUFQ3_K_M4.41 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q3_K_S.ggufGGUFQ3_K_S4.06 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q4_K_M.ggufGGUFQ4_K_M5.38 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q4_K_S.ggufGGUFQ4_K_S5.11 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q5_K_M.ggufGGUFQ5_K_M6.19 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q5_K_S.ggufGGUFQ5_K_S6.03 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q6_K.ggufGGUFQ6_K7.04 GBDownload
DeepSeek-V4-Pro-Qwen3.5-9B-MTP-Q8_0.ggufGGUFQ8_09.11 GBDownload
mmproj-F32.ggufGGUFF321.70 GBDownload

Model Details

Model IDJackrong/DeepSeek-V4-Pro-Qwen3.5-9B-MTP-GGUF
AuthorJackrong
Pipelinetext-generation
Licenseapache-2.0
Base modelunsloth/Qwen3.5-9B
Last modified2026-08-13T08:45:54.000Z

Model README

---

base_model: unsloth/Qwen3.5-9B

tags:

  • text-generation-inference
  • transformers
  • unsloth
  • qwen3_5
  • reasoning
  • distillation
  • deepseek
  • sft
  • rl
  • gspo
  • math
  • stem
  • tool-use
  • function-calling
  • mtp

license: apache-2.0

language:

  • en
  • zh
  • ko
  • ja
  • es
  • ru

pipeline_tag: text-generation

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">

<div style="background: linear-gradient(135deg, #7c3aed 0%, #4c1d95 100%); padding: 24px; color: white;">

<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">

<h1 style="margin: 0; font-size: 26px; font-weight: 800; color: white; border: none;">🧠 DeepSeek-V4-Pro-Qwen3.5-9B</h1>

<span style="background: #10b981; color: white; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">Distillation Release</span>

</div>

<p style="margin: 8px 0 0 0; font-size: 14px; color: #ddd6fe; font-weight: 500;">A compact 9B reasoning model distilled from DeepSeek-V4-Pro in Max Effect mode</p>

</div>

<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">

<span style="background: #f3e8ff; color: #6b21a8; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #e9d5ff;">🔬 DeepSeek-V4-Pro Distillation</span>

<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">🧠 9B Parameters</span>

<span style="background: #e0f2fe; color: #0369a1; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bae6fd;">📐 ~250K Math &amp; STEM Samples</span>

<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;">⚡ MTP &amp; NVFP4 Evaluated</span>

</div>

<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">

<div style="background: #f5f3ff; border-left: 5px solid #7c3aed; padding: 16px; border-radius: 0 8px 8px 0;">

<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #6d28d9; font-weight: 700;">💡 What is DeepSeek-V4-Pro-Qwen3.5-9B?</h3>

<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;"><b>DeepSeek-V4-Pro-Qwen3.5-9B</b> is a reasoning-focused fine-tune of <b>Qwen3.5-9B</b>, distilled from responses generated by <b>DeepSeek-V4-Pro in Max Effect mode</b>. The release concentrates its supervised signal on mathematics and STEM problem solving while preserving the efficiency and practical deployability of the 9B parameter class.</p>

</div>

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(190px, 1fr)); gap: 15px;">

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">

<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">🧩 Structured Reasoning</span>

<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Learns rigorous problem decomposition and multi-step solution patterns from a strong teacher.</span>

</div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">

<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">📐 Math &amp; STEM Focus</span>

<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Trained on approximately 250,000 samples centered on mathematics and scientific reasoning.</span>

</div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">

<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">🧪 Cross-Domain Transfer</span>

<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Shows small post-training gains in programming and tool use despite receiving no coding-specific SFT data.</span>

</div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">

<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">⚙️ Practical Deployment</span>

<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Evaluated through vLLM NVFP4 and llama.cpp MTP-Q8_0 deployment paths.</span>

</div>

</div>

</div>

</div>

🤝 Collaboration & Training Support

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: grid; grid-template-columns: repeat(auto-fit, minmax(260px, 1fr)); gap: 14px; margin-bottom: 28px;">

<div style="border: 1px solid #a7f3d0; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #10b981 0%, #047857 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;"><span>🧪</span> Hardware Cooperation &amp; Joint Collaboration</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

This project is built in close collaboration with hardware engineer <b>Kyle Hessling</b>, whose infrastructure, training support, and evaluation assistance helped make this release possible.

<div style="margin-top: 10px;">👉 Follow his hardware and model-training updates on X / Twitter: <a href="https://x.com/KyleHessling1" target="_blank" style="color: #047857; text-decoration: none; font-weight: 700;">@KyleHessling1</a></div>

</div>

</div>

<div style="border: 1px solid #ddd6fe; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #7c3aed 0%, #6d28d9 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;"><span>🦥</span> Fine-tuning Framework (Unsloth)</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

The model training workflow is accelerated and memory-optimized with <b>Unsloth</b>. Special thanks to the Unsloth team for making efficient large-model fine-tuning more accessible.

<div style="margin-top: 10px;">👉 Documentation and fine-tuning guidance: <a href="https://unsloth.ai/docs" target="_blank" style="color: #7c3aed; text-decoration: none; font-weight: 700;">unsloth.ai/docs</a></div>

</div>

</div>

</div>

🔬 1. Distillation & Training Data

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02); margin-bottom: 28px;">

<div style="background: linear-gradient(135deg, #0284c7 0%, #0369a1 100%); padding: 13px 17px; color: white; font-weight: 700; font-size: 14px;">🧬 From DeepSeek-V4-Pro to a Compact 9B Student</div>

<div style="padding: 18px; font-size: 13px; color: #334155; line-height: 1.75;">

<p style="margin: 0 0 12px 0;">The teacher data for this release was generated with <b>DeepSeek-V4-Pro (Max Effect)</b>. Approximately <b>250,000 mathematics and STEM samples</b> were used for supervised fine-tuning, with an emphasis on structured derivation, domain-aware problem decomposition, and reliable final-answer construction.</p>

<div style="background: #f0f9ff; border-left: 4px solid #0284c7; border-radius: 0 8px 8px 0; padding: 12px 14px; color: #0c4a6e;">

<b>Special training note:</b> the training path for this release ran pure supervised fine-tuning (SFT) first, followed by a light reinforcement-learning (RL) refinement stage using GSPO (Grouped Stepwise Preference Optimization). No coding data was included in either stage. Nevertheless, later testing showed a small generalization improvement in programming and tool-calling tasks. This should be interpreted as cross-domain transfer from stronger reasoning structure—not as evidence that the model received direct coding supervision.

<p style="margin: 10px 0 0 0;"><b>Why coding coverage is currently limited:</b> during this training cycle, DeepSeek released newer 0731 and 0813 models with substantially improved coding and agentic capabilities. The corresponding new teacher data was not yet available when this distillation run was prepared. Once the new model data is collected and prepared, this distillation line will be iteratively upgraded with dedicated coding and agent traces.</p>

</div>

</div>

</div>

📊 2. Benchmark Results

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 16px; overflow: hidden; background: #ffffff; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05); margin-bottom: 28px;">

<div style="background: linear-gradient(135deg, #0f766e 0%, #0369a1 100%); padding: 20px; color: white;">

<h3 style="margin: 0; font-size: 20px; font-weight: 700; color: white; border: none;">⚡ Performance Snapshot</h3>

<p style="margin: 5px 0 0 0; font-size: 13px; color: #cffafe;">Strong grade-school mathematics performance and the best average MMLU-Pro result among the tested 9B baselines.</p>

</div>

<div style="padding: 22px;">

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(170px, 1fr)); gap: 15px;">

<div style="border: 1px solid #bae6fd; padding: 16px; border-radius: 10px; background: #f0f9ff; text-align: center;">

<span style="font-size: 11px; font-weight: 700; color: #0369a1; text-transform: uppercase; display: block; margin-bottom: 6px;">GSM8K</span>

<span style="font-size: 25px; font-weight: 800; color: #0c4a6e; display: block;">94.50%</span>

<span style="font-size: 11px; color: #64748b;">4-run average · NVFP4 · vLLM</span>

</div>

<div style="border: 1px solid #ddd6fe; padding: 16px; border-radius: 10px; background: #f5f3ff; text-align: center;">

<span style="font-size: 11px; font-weight: 700; color: #6d28d9; text-transform: uppercase; display: block; margin-bottom: 6px;">MMLU-Pro Math</span>

<span style="font-size: 25px; font-weight: 800; color: #4c1d95; display: block;">92.40%</span>

<span style="font-size: 11px; color: #64748b;">MTP-Q8_0 GGUF · 8K</span>

</div>

<div style="border: 1px solid #a7f3d0; padding: 16px; border-radius: 10px; background: #f0fdf4; text-align: center;">

<span style="font-size: 11px; font-weight: 700; color: #047857; text-transform: uppercase; display: block; margin-bottom: 6px;">MMLU-Pro Average</span>

<span style="font-size: 25px; font-weight: 800; color: #065f46; display: block;">90.53%</span>

<span style="font-size: 11px; color: #64748b;">Math · Physics · Chemistry</span>

</div>

<div style="border: 1px solid #fde68a; padding: 16px; border-radius: 10px; background: #fffbeb; text-align: center;">

<span style="font-size: 11px; font-weight: 700; color: #b45309; text-transform: uppercase; display: block; margin-bottom: 6px;">Gain vs Base</span>

<span style="font-size: 25px; font-weight: 800; color: #92400e; display: block;">+0.93</span>

<span style="font-size: 11px; color: #64748b;">MMLU-Pro average points</span>

</div>

</div>

</div>

</div>

2.1 GSM8K

The NVFP4 checkpoint was evaluated on GSM8K four times, and the resulting scores were averaged. Its four-run mean accuracy is 94.50% under vLLM. In the comparison supplied with this release, the 9B student slightly exceeds the reported DeepSeek-V4-Pro teacher result and approaches the Llama-3.3-70B-Instruct reference.

<div style="overflow-x: auto; margin-bottom: 24px;">

<table style="display: table; width: 100%; table-layout: fixed; border-collapse: collapse; overflow-wrap: anywhere; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; line-height: 1.25; border: 1px solid #cbd5e1;">

<thead>

<tr style="background: #eff6ff;">

<th style="width: 72%; padding: 6px 10px !important; border-bottom: 2px solid #0284c7; text-align: left; color: #0369a1;">Model</th>

<th style="width: 28%; padding: 6px 10px !important; border-bottom: 2px solid #0284c7; text-align: right; color: #0369a1;">Accuracy</th>

</tr>

</thead>

<tbody>

<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;"><b>MiMo-V2.5-Pro</b></td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;"><b>99.60%</b></td></tr>

<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;">Llama-3.1-405B-Instruct</td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;">96.80%</td></tr>

<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;">Llama-3.3-70B-Instruct</td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;">94.84%</td></tr>

<tr style="background: #f5f3ff;"><td style="padding: 6px 10px !important; border-bottom: 1px solid #ddd6fe; color: #5b21b6;"><b>DeepSeek-V4-Pro-Qwen3.5-9B</b></td><td style="padding: 6px 10px !important; border-bottom: 1px solid #ddd6fe; text-align: right; color: #5b21b6;"><b>94.50%</b></td></tr>

<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;">DeepSeek-V4-Pro</td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;">92.60%</td></tr>

<tr><td style="padding: 6px 10px !important;">DeepSeek-V3</td><td style="padding: 6px 10px !important; text-align: right;">89.30%</td></tr>

</tbody>

</table>

</div>

<div style="background: #f8fafc; border-left: 4px solid #64748b; padding: 11px 14px; border-radius: 0 8px 8px 0; color: #475569; font-size: 12px; line-height: 1.6; margin-bottom: 26px;"><b>Evaluation configuration:</b> DeepSeek-V4-Pro-Qwen3.5-9B-NVFP4 · NVFP4 (FP8 E4M3) · vLLM · four GSM8K test runs, with the final 94.50% reported as their arithmetic mean. The reference scores for the other models shown above are sourced from the <a href="https://huggingface.co/datasets/openai/gsm8k" style="color: #0369a1; font-weight: 700; text-decoration: none;">openai/gsm8k repository</a>.</div>

2.2 MMLU-Pro: Math, Physics & Chemistry

Each model was evaluated on 500 questions from each of three MMLU-Pro subsets—Math, Physics, and Chemistry—for 1,500 questions per model. In this reported comparison, DeepSeek-V4-Pro-Qwen3.5-9B achieves the highest average at 90.53%. Its compact 4B sibling is now included alongside the Qwen and Claude Mythos-distilled references.

<div style="overflow-x: auto; margin-bottom: 20px;">

<table style="display: table; width: 100%; min-width: 760px; table-layout: auto; border-collapse: collapse; overflow-wrap: anywhere; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; line-height: 1.25; border: 1px solid #cbd5e1;">

<thead>

<tr style="background: #f5f3ff;">

<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Model</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Evaluation Build</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Math</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Physics</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Chemistry</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Average</th>

</tr>

</thead>

<tbody>

<tr style="background: #f5f3ff;">

<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; color: #5b21b6;"><b>DeepSeek-V4-Pro-Qwen3.5-9B</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe;">MTP-Q8_0 GGUF</td>

<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right;"><b>92.40%</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right;"><b>89.40%</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right;"><b>89.80%</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right; color: #5b21b6;"><b>90.53%</b></td>

</tr>

<tr>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Qwen3.5-9B</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MTP-Q8_0 GGUF</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">90.60%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.00%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.20%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.60%</td>

</tr>

<tr>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Claude Mythos-distilled 27B</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MTP-Q8 GGUF</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">88.40%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">87.60%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">82.60%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">86.20%</td>

</tr>

<tr>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Claude Mythos-distilled 9B</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MTP-Q8_0 GGUF</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">86.00%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">78.60%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">78.60%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">81.07%</td>

</tr>

<tr style="background: #f0fdf4;">

<td style="padding: 10px 12px; color: #065f46;"><b>DeepSeek-V4-Pro-Qwen3.5-4B</b></td>

<td style="padding: 10px 12px;">MTP-Q8 GGUF</td>

<td style="padding: 10px 12px; text-align: right;"><b>80.40%</b></td>

<td style="padding: 10px 12px; text-align: right;"><b>74.20%</b></td>

<td style="padding: 10px 12px; text-align: right;"><b>74.80%</b></td>

<td style="padding: 10px 12px; text-align: right; color: #065f46;"><b>76.47%</b></td>

</tr>

</tbody>

</table>

</div>

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(220px, 1fr)); gap: 14px; margin-bottom: 28px; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif;">

<div style="border: 1px solid #a7f3d0; border-radius: 10px; background: #f0fdf4; padding: 14px; color: #065f46; font-size: 13px; line-height: 1.6;"><b>📈 vs. Qwen3.5-9B</b><br>+1.80 points in Math, +0.40 in Physics, +0.60 in Chemistry, and <b>+0.93 average points</b>.</div>

<div style="border: 1px solid #bae6fd; border-radius: 10px; background: #f0f9ff; padding: 14px; color: #0c4a6e; font-size: 13px; line-height: 1.6;"><b>🧪 Evaluation setup</b><br>500 questions per subset · 1,500 total per model · llama.cpp. The 9B rows use MTP-Q8_0 GGUF at 8K context, the 4B student uses MTP-Q8 GGUF at 64K context, and Claude Mythos-distilled 27B uses MTP-Q8 GGUF at 32K context.</div>

</div>

2.3 Inference Efficiency

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 16px; overflow: hidden; background: #ffffff; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05); margin-bottom: 28px;">

<div style="background: linear-gradient(135deg, #0f766e 0%, #047857 55%, #0369a1 100%); padding: 20px; color: white;">

<h3 style="margin: 0; font-size: 20px; font-weight: 800; color: white; border: none;">⚡ More Correct Answers per Reasoning Token</h3>

<p style="margin: 6px 0 0 0; font-size: 13px; color: #ccfbf1; line-height: 1.6;">On the same 1,500-question MMLU-Pro sample, the DeepSeek-distilled model combines the highest accuracy with the lowest token cost per correct answer.</p>

</div>

<div style="padding: 20px;">

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(220px, 1fr)); gap: 14px;">

<div style="border: 1px solid #a7f3d0; border-radius: 10px; background: #f0fdf4; padding: 15px; color: #065f46; font-size: 13px; line-height: 1.6;">

<span style="font-size: 11px; font-weight: 800; color: #047857; text-transform: uppercase; letter-spacing: 0.4px;">vs. Qwen3.5-9B</span><br>

<span style="font-size: 22px; font-weight: 800;">36.1% fewer tokens</span><br>

per correct answer, with <b>+0.93 accuracy points</b>.

</div>

<div style="border: 1px solid #bae6fd; border-radius: 10px; background: #f0f9ff; padding: 15px; color: #0c4a6e; font-size: 13px; line-height: 1.6;">

<span style="font-size: 11px; font-weight: 800; color: #0369a1; text-transform: uppercase; letter-spacing: 0.4px;">vs. Claude Mythos-distilled 9B</span><br>

<span style="font-size: 22px; font-weight: 800;">22.7% fewer tokens</span><br>

per correct answer, with <b>+9.47 accuracy points</b>.

</div>

<div style="border: 1px solid #ddd6fe; border-radius: 10px; background: #f5f3ff; padding: 15px; color: #4c1d95; font-size: 13px; line-height: 1.6;">

<span style="font-size: 11px; font-weight: 800; color: #6d28d9; text-transform: uppercase; letter-spacing: 0.4px;">Correct answers / 1M tokens</span><br>

<span style="font-size: 22px; font-weight: 800;">316.8</span><br>

<b>+56.4%</b> vs. Qwen and <b>+29.3%</b> vs. Claude Mythos-distilled 9B.

</div>

</div>

</div>

</div>

<div style="overflow-x: auto; margin-bottom: 20px;">

<table style="display: table; width: 100%; min-width: 900px; table-layout: auto; border-collapse: collapse; overflow-wrap: anywhere; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; line-height: 1.25; border: 1px solid #cbd5e1;">

<thead>

<tr style="background: #ecfdf5;">

<th style="padding: 10px 12px; border-bottom: 2px solid #10b981; text-align: left; color: #047857;">Model</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #10b981; text-align: right; color: #047857;">Accuracy</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #10b981; text-align: right; color: #047857;">Avg. Reasoning<br>Tokens / Question</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #10b981; text-align: right; color: #047857;">Median</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #10b981; text-align: right; color: #047857;">Tokens /<br>Correct Answer</th>

<th style="padding: 10px 12px; border-bottom: 2px solid #10b981; text-align: right; color: #047857;">Correct Answers /<br>1M Tokens</th>

</tr>

</thead>

<tbody>

<tr style="background: #f0fdf4;">

<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; color: #065f46;"><b>DeepSeek-V4-Pro-Qwen3.5-9B</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right; color: #065f46;"><b>90.53%</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right; color: #065f46;"><b>2,858</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right;">676</td>

<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right; color: #065f46;"><b>3,157</b></td>

<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right; color: #065f46;"><b>316.8</b></td>

</tr>

<tr>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Qwen3.5-9B</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.60%</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">4,425</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">3,214</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">4,938</td>

<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">202.5</td>

</tr>

<tr>

<td style="padding: 9px 12px;">Claude Mythos-distilled 9B</td>

<td style="padding: 9px 12px; text-align: right;">81.07%</td>

<td style="padding: 9px 12px; text-align: right;">3,310</td>

<td style="padding: 9px 12px; text-align: right;"><b>639</b></td>

<td style="padding: 9px 12px; text-align: right;">4,083</td>

<td style="padding: 9px 12px; text-align: right;">244.9</td>

</tr>

</tbody>

</table>

</div>

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(280px, 1fr)); gap: 15px; margin-bottom: 28px; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif;">

<div style="border: 1px solid #ddd6fe; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: #f5f3ff; border-bottom: 1px solid #ddd6fe; padding: 11px 14px; color: #6d28d9; font-size: 13px; font-weight: 800;">🧪 Token Savings by MMLU-Pro Subject</div>

<div style="overflow-x: auto;">

<table style="display: table; width: 100%; min-width: 430px; table-layout: auto; border-collapse: collapse; overflow-wrap: anywhere; font-size: 12px; line-height: 1.25; color: #334155;">

<thead><tr style="background: #fafafa;"><th style="padding: 9px 12px; text-align: left; border-bottom: 1px solid #e2e8f0;">Subject</th><th style="padding: 9px 12px; text-align: right; border-bottom: 1px solid #e2e8f0;">vs. Qwen3.5-9B</th><th style="padding: 9px 12px; text-align: right; border-bottom: 1px solid #e2e8f0;">vs. Claude Mythos-distilled 9B</th></tr></thead>

<tbody>

<tr><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Math</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right; color: #047857;"><b>47.4%</b></td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right; color: #047857;"><b>35.6%</b></td></tr>

<tr><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Physics</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right; color: #047857;"><b>30.8%</b></td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right; color: #047857;"><b>10.8%</b></td></tr>

<tr><td style="padding: 9px 12px;">Chemistry</td><td style="padding: 9px 12px; text-align: right; color: #047857;"><b>32.1%</b></td><td style="padding: 9px 12px; text-align: right; color: #047857;"><b>23.3%</b></td></tr>

</tbody>

</table>

</div>

</div>

<div style="border: 1px solid #bae6fd; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: #f0f9ff; border-bottom: 1px solid #bae6fd; padding: 11px 14px; color: #0369a1; font-size: 13px; font-weight: 800;">📐 GSM8K NVFP4 Reasoning Profile</div>

<div style="padding: 14px;">

<div style="display: grid; grid-template-columns: repeat(2, minmax(110px, 1fr)); gap: 10px;">

<div style="background: #f8fafc; border-radius: 8px; padding: 10px;"><span style="font-size: 10px; color: #64748b; text-transform: uppercase; font-weight: 700;">Average</span><br><b style="font-size: 18px; color: #0c4a6e;">1,627</b><span style="font-size: 11px; color: #64748b;"> tokens</span></div>

<div style="background: #f8fafc; border-radius: 8px; padding: 10px;"><span style="font-size: 10px; color: #64748b; text-transform: uppercase; font-weight: 700;">Median</span><br><b style="font-size: 18px; color: #0c4a6e;">372</b><span style="font-size: 11px; color: #64748b;"> tokens</span></div>

<div style="background: #f8fafc; border-radius: 8px; padding: 10px;"><span style="font-size: 10px; color: #64748b; text-transform: uppercase; font-weight: 700;">P90</span><br><b style="font-size: 18px; color: #0c4a6e;">3,042</b><span style="font-size: 11px; color: #64748b;"> tokens</span></div>

<div style="background: #f8fafc; border-radius: 8px; padding: 10px;"><span style="font-size: 10px; color: #64748b; text-transform: uppercase; font-weight: 700;">Per Correct Answer</span><br><b style="font-size: 18px; color: #0c4a6e;">1,721</b><span style="font-size: 11px; color: #64748b;"> tokens</span></div>

</div>

<p style="margin: 11px 0 0 0; font-size: 11px; color: #64748b; line-height: 1.5;">Computed from <b>2,638</b> available raw NVFP4 trajectories.</p>

</div>

</div>

</div>

<div style="background: #f8fafc; border-left: 4px solid #64748b; padding: 11px 14px; border-radius: 0 8px 8px 0; color: #475569; font-size: 12px; line-height: 1.6; margin-bottom: 28px;"><b>Metric notes:</b> MMLU-Pro efficiency uses 1,500 questions per model and the supplied composite score for the final correct-answer total. “Tokens per correct answer” is total reasoning tokens divided by correct answers; “correct answers per 1M tokens” is the inverse throughput measure. Lower is better for token-cost metrics; higher is better for accuracy and correct-answer throughput. Comparison deltas follow the supplied composite-result calculations, so recomputing them only from the rounded display values may produce a 0.01-point difference. GSM8K statistics summarize the currently available NVFP4 raw trajectories and are reported separately from the MMLU-Pro MTP-Q8_0 comparison.</div>

🛠️ 3. Generalization Beyond the Training Mix

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; margin-bottom: 28px;">

<div style="background: linear-gradient(135deg, #10b981 0%, #047857 100%); padding: 13px 17px; color: white; font-weight: 700; font-size: 14px;">🔁 A Small but Interesting Transfer Effect</div>

<div style="padding: 18px; color: #334155; font-size: 13px; line-height: 1.7;">

Although the SFT mixture intentionally contained <b>no coding data</b>, post-training tests indicated a <b>small improvement in programming and tool-calling behavior</b>. A plausible interpretation is that mathematics and STEM supervision strengthened reusable skills such as decomposition, constraint tracking, verification, and structured action planning.

<br><br>

This observation is preliminary. It should not be read as a claim that the model is coding-specialized, and the release does not replace dedicated code benchmarks or agent-environment evaluation.

</div>

</div>

✅ 4. Instruction Following & Output-Format Compliance

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02); margin-bottom: 28px;">

<div style="background: linear-gradient(135deg, #6d28d9 0%, #4c1d95 100%); padding: 13px 17px; color: white; font-weight: 700; font-size: 14px;">📣 A Model That Listens to Its Format</div>

<div style="padding: 18px; color: #334155; font-size: 13px; line-height: 1.75;">

DeepSeek-V4-Pro-Qwen3.5-9B shows strong instruction-following. In every test task for this release it conformed fully to the demanded output format, while the official and comparison models showed partial non-following. On the MMLU-Pro eval (1,500 identical questions per model, same system prompt and backend), where the instruction was <i>"After thinking, output only <b>ONE single capital letter</b>"</i>, the model returned a terse single-letter answer in <b>97.2%</b> of cases and a verbose re-explanation in only <b>0.2%</b> — versus Qwen3.5-9B official (<b>37.3%</b> / 62.1% verbose) and Claude Mythos-distilled 9B (<b>69.6%</b> / 21.1% verbose). Format compliance is orthogonal to correctness, and on GSM8K the model likewise closes with a direct answer in ≈99.8% of generations.

</div>

</div>

🎮 Game Showcase: Star Skip by Kyle

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 16px; overflow: hidden; background: #ffffff; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05); margin-bottom: 28px;">

<div style="background: linear-gradient(135deg, #f59e0b 0%, #dc2626 100%); padding: 16px 20px; color: white;">

<h3 style="margin: 0; font-size: 20px; font-weight: 800; color: white; border: none;">🎮 An Interactive Game Built with DeepSeek-V4-Pro-Qwen3.5-9B</h3>

<p style="margin: 6px 0 0 0; font-size: 13px; color: #fef3c7; line-height: 1.6;">A hands-on demonstration of what this 9B model can power in practice.</p>

</div>

<div style="padding: 20px;">

<img src="https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/mLc45iGJh339PxVA7HzxI.png" alt="Star Skip game showcase screenshot" style="width: 100%; max-width: 720px; border-radius: 12px; border: 1px solid #e2e8f0; display: block; margin: 0 auto 16px auto;">

<p style="margin: 0 0 14px 0; font-size: 13px; color: #334155; line-height: 1.7;">Hardware engineer <b>Kyle Hessling</b> used <b>DeepSeek-V4-Pro-Qwen3.5-9B</b> to build <b>Star Skip</b> — a playable game demo that puts the model's structured reasoning and instruction-following skills to work in an interactive environment.</p>

<div style="background: #fffbeb; border: 1px solid #fde68a; border-radius: 10px; padding: 13px 16px; color: #92400e; font-size: 13px; line-height: 1.6;">

👉 <b>Try the game here:</b> <a href="https://huggingface.co/spaces/KyleHessling1/starskip" target="_blank" style="color: #b45309; font-weight: 700; text-decoration: none;">KyleHessling1/starskip — Star Skip</a>

</div>

</div>

</div>

🎯 5. Recommended Uses

  • Mathematical problem solving and step-by-step derivation
  • Physics, chemistry, and broader STEM question answering
  • Structured reasoning and analytical instruction following
  • Research on reasoning distillation and cross-domain SFT transfer
  • Experimental tool-calling or programming workflows where outputs are independently validated

⚠️ 6. Limitations

  • This is an experimental 9B community model and remains subject to hallucinations, reasoning errors, and unstable behavior on difficult or underspecified tasks.
  • The reported MMLU-Pro results are based on 500 randomly sampled questions from each of the Math, Physics, and Chemistry subsets, rather than the complete subsets; they are not a comprehensive capability or safety evaluation.
  • GSM8K and MMLU-Pro used different inference builds and backends, so their scores should not be used to infer quantization equivalence.
  • Programming and tool-use gains were observed without direct coding supervision but are described only as small, preliminary generalization gains.
  • Users should independently verify high-stakes mathematical, scientific, or factual outputs.

📚 7. Resources, Acknowledgements & Citation

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 14px; overflow: hidden; background: #ffffff; margin-top: 18px;">

<div style="background: linear-gradient(135deg, #7c3aed 0%, #4c1d95 100%); padding: 16px 20px; color: white; font-weight: 800; font-size: 16px;">🌟 Open Research Release</div>

<div style="padding: 20px; color: #334155; font-size: 13px; line-height: 1.7;">

<p style="margin-top: 0;"><b>Fine-tuning guide:</b> <a href="https://github.com/R6410418/Jackrong-llm-finetuning-guide" style="color: #7c3aed; font-weight: 700; text-decoration: none;">Jackrong LLM Fine-Tuning Guide</a></p>

<p><b>Acknowledgements:</b> Kyle Hessling for hardware collaboration, training support, and evaluation assistance; DeepSeek for the teacher model and reasoning data source; Qwen for the base model family; Unsloth for efficient fine-tuning tooling; and the open-source inference community behind vLLM and llama.cpp.</p>

<p><b>Citation:</b></p>

<pre style="background: #f5f3ff; color: #4c1d95; border: 1px solid #ddd6fe; padding: 14px; border-radius: 8px; overflow-x: auto; font-size: 12px;"><code>@misc{jackrong_deepseek_v4_pro_qwen35_9b,

title = {DeepSeek-V4-Pro-Qwen3.5-9B},

author = {Jackrong},

year = {2026},

publisher = {Hugging Face}

}</code></pre>

</div>

</div>

Run Jackrong/DeepSeek-V4-Pro-Qwen3.5-9B-MTP-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models