Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B-MTP-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid cbd5e1; border radius: 16px; box shadow: 0 10px 15…
Runs locally from ~1.24 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-BF16.gguf | GGUF | BF16 | 8.07 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-IQ4_XS.gguf | GGUF | IQ4_XS | 2.42 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q2_K.gguf | GGUF | Q2_K | 1.82 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q3_K_L.gguf | GGUF | Q3_K_L | 2.31 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q3_K_M.gguf | GGUF | Q3_K_M | 2.16 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q3_K_S.gguf | GGUF | Q3_K_S | 1.98 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q4_K_M.gguf | GGUF | Q4_K_M | 2.59 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q4_K_S.gguf | GGUF | Q4_K_S | 2.45 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q5_K_M.gguf | GGUF | Q5_K_M | 2.94 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q5_K_S.gguf | GGUF | Q5_K_S | 2.86 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q6_K.gguf | GGUF | Q6_K | 3.32 GB | Download |
| DeepSeek-V4-Pro-Qwen3.5-4B-MTP-Q8_0.gguf | GGUF | Q8_0 | 4.29 GB | Download |
| mmproj-F32.gguf | GGUF | F32 | 1.24 GB | Download |
Model Details
| Model ID | Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B-MTP-GGUF |
|---|---|
| Author | Jackrong |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | unsloth/Qwen3.5-4B |
| Last modified | 2026-08-13T08:45:49.000Z |
Model README
---
base_model: unsloth/Qwen3.5-4B
tags:
- text-generation-inference
- transformers
- unsloth
- qwen3_5
- reasoning
- distillation
- deepseek
- sft
- math
- stem
- mtp
- gguf
license: apache-2.0
language:
- en
- zh
pipeline_tag: text-generation
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #4c1d95 100%); padding: 24px; color: white;">
<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">
<h1 style="margin: 0; font-size: 26px; font-weight: 800; color: white; border: none;">🧠 DeepSeek-V4-Pro-Qwen3.5-4B</h1>
<span style="background: #10b981; color: white; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">Distillation Release</span>
</div>
<p style="margin: 8px 0 0 0; font-size: 14px; color: #ddd6fe; font-weight: 500;">A lightweight 4B reasoning model distilled from DeepSeek-V4-Pro in Max Effect mode</p>
</div>
<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">
<span style="background: #f3e8ff; color: #6b21a8; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #e9d5ff;">🔬 DeepSeek-V4-Pro Distillation</span>
<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">🧠 4B Parameters</span>
<span style="background: #e0f2fe; color: #0369a1; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bae6fd;">📐 ~250K Math & STEM Samples</span>
<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;">⚡ MTP-Q8 Evaluated</span>
</div>
<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">
<div style="background: #f5f3ff; border-left: 5px solid #7c3aed; padding: 16px; border-radius: 0 8px 8px 0;">
<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #6d28d9; font-weight: 700;">💡 What is DeepSeek-V4-Pro-Qwen3.5-4B?</h3>
<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;"><b>DeepSeek-V4-Pro-Qwen3.5-4B</b> is a reasoning-focused fine-tune of <b>Qwen3.5-4B</b>, distilled from responses generated by <b>DeepSeek-V4-Pro in Max Effect mode</b>. It follows the same training pipeline and approximately 250,000-sample mathematics and STEM mixture used for the 9B release, while targeting a substantially smaller and more accessible deployment class.</p>
</div>
<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(190px, 1fr)); gap: 15px;">
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">
<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">🧩 Structured Reasoning</span>
<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Learns multi-step solution patterns and explicit problem decomposition from a strong teacher.</span>
</div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">
<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">📐 Math & STEM Focus</span>
<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Trained on approximately 250,000 samples centered on mathematical and scientific reasoning.</span>
</div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">
<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">⚡ Compact Deployment</span>
<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Brings the DeepSeek-distilled training recipe into the efficient 4B parameter class.</span>
</div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa;">
<span style="font-weight: 700; color: #6b21a8; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase;">🧪 Reproducible Evaluation</span>
<span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Evaluated on complete GSM8K runs and fixed MMLU-Pro Math, Physics, and Chemistry samples.</span>
</div>
</div>
</div>
</div>
🤝 Collaboration & Training Support
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: grid; grid-template-columns: repeat(auto-fit, minmax(260px, 1fr)); gap: 14px; margin-bottom: 28px;">
<div style="border: 1px solid #a7f3d0; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #10b981 0%, #047857 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;"><span>🧪</span> Hardware Cooperation & Joint Collaboration</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
This project is built in close collaboration with hardware engineer <b>Kyle Hessling</b>, whose infrastructure, training support, and evaluation assistance helped make this release possible.
<div style="margin-top: 10px;">👉 Follow his hardware and model-training updates on X / Twitter: <a href="https://x.com/KyleHessling1" target="_blank" style="color: #047857; text-decoration: none; font-weight: 700;">@KyleHessling1</a></div>
</div>
</div>
<div style="border: 1px solid #ddd6fe; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #6d28d9 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;"><span>🦥</span> Fine-tuning Framework (Unsloth)</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
The model training workflow is accelerated and memory-optimized with <b>Unsloth</b>. Special thanks to the Unsloth team for making efficient large-model fine-tuning more accessible.
<div style="margin-top: 10px;">👉 Documentation and fine-tuning guidance: <a href="https://unsloth.ai/docs" target="_blank" style="color: #7c3aed; text-decoration: none; font-weight: 700;">unsloth.ai/docs</a></div>
</div>
</div>
</div>
🔬 1. Distillation & Training Data
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02); margin-bottom: 20px;">
<div style="background: linear-gradient(135deg, #0284c7 0%, #0369a1 100%); padding: 13px 17px; color: white; font-weight: 700; font-size: 14px;">🧬 The 9B Training Recipe, Compressed into a 4B Student</div>
<div style="padding: 18px; font-size: 13px; color: #334155; line-height: 1.75;">
<p style="margin: 0 0 12px 0;">The teacher data for this release was generated with <b>DeepSeek-V4-Pro (Max Effect)</b>. The 4B student uses the <b>same training flow and approximately 250,000 mathematics and STEM samples</b> as DeepSeek-V4-Pro-Qwen3.5-9B, emphasizing structured derivation, problem decomposition, verification, and reliable final-answer construction.</p>
<div style="background: #f0f9ff; border-left: 4px solid #0284c7; border-radius: 0 8px 8px 0; padding: 12px 14px; color: #0c4a6e;">
<b>Training scope:</b> no coding data was included in this SFT mixture. The published benchmarks in this card therefore focus on mathematics and STEM; coding and tool-use transfer have not yet been established for the 4B release.
<p style="margin: 10px 0 0 0;"><b>Why coding coverage is currently limited:</b> during this training cycle, DeepSeek released newer 0731 and 0813 models with substantially improved coding and agentic capabilities. The corresponding new teacher data was not yet available when this distillation run was prepared. Once the new model data is collected and prepared, this distillation line will be iteratively upgraded with dedicated coding and agent traces.</p>
</div>
</div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 12px; margin-bottom: 28px;">
<div style="border: 1px solid #ddd6fe; border-radius: 10px; background: #f5f3ff; padding: 14px; color: #4c1d95; font-size: 13px; line-height: 1.6;"><b>1 · Teacher Generation</b><br>DeepSeek-V4-Pro in Max Effect mode produces detailed mathematics and STEM solutions.</div>
<div style="border: 1px solid #bae6fd; border-radius: 10px; background: #f0f9ff; padding: 14px; color: #0c4a6e; font-size: 13px; line-height: 1.6;"><b>2 · Data Preparation</b><br>Approximately 250K samples are cleaned and formatted for supervised reasoning transfer.</div>
<div style="border: 1px solid #a7f3d0; border-radius: 10px; background: #f0fdf4; padding: 14px; color: #065f46; font-size: 13px; line-height: 1.6;"><b>3 · Unsloth SFT</b><br>Qwen3.5-4B is fine-tuned through an Unsloth LoRA pipeline and merged into BF16.</div>
<div style="border: 1px solid #fde68a; border-radius: 10px; background: #fffbeb; padding: 14px; color: #92400e; font-size: 13px; line-height: 1.6;"><b>4 · Deployment Build</b><br>An MTP-enabled Q8 GGUF build is used for the reported llama.cpp evaluations.</div>
</div>
📊 2. Benchmark Results
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 16px; overflow: hidden; background: #ffffff; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05); margin-bottom: 28px;">
<div style="background: linear-gradient(135deg, #0f766e 0%, #0369a1 100%); padding: 20px; color: white;">
<h3 style="margin: 0; font-size: 20px; font-weight: 700; color: white; border: none;">⚡ Performance Snapshot</h3>
<p style="margin: 5px 0 0 0; font-size: 13px; color: #cffafe;">Strong grade-school mathematics performance from a compact 4B model, with completed 1,500-question MMLU-Pro STEM results.</p>
</div>
<div style="padding: 22px;">
<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(170px, 1fr)); gap: 15px;">
<div style="border: 1px solid #bae6fd; padding: 16px; border-radius: 10px; background: #f0f9ff; text-align: center;">
<span style="font-size: 11px; font-weight: 700; color: #0369a1; text-transform: uppercase; display: block; margin-bottom: 6px;">GSM8K</span>
<span style="font-size: 25px; font-weight: 800; color: #0c4a6e; display: block;">91.77%</span>
<span style="font-size: 11px; color: #64748b;">2-run pass@1 mean</span>
</div>
<div style="border: 1px solid #c4b5fd; padding: 16px; border-radius: 10px; background: #faf5ff; text-align: center;">
<span style="font-size: 11px; font-weight: 700; color: #7c3aed; text-transform: uppercase; display: block; margin-bottom: 6px;">MMLU-Pro Math</span>
<span style="font-size: 25px; font-weight: 800; color: #6d28d9; display: block;">80.40%</span>
<span style="font-size: 11px; color: #64748b;">402 / 500</span>
</div>
<div style="border: 1px solid #ddd6fe; padding: 16px; border-radius: 10px; background: #f5f3ff; text-align: center;">
<span style="font-size: 11px; font-weight: 700; color: #6d28d9; text-transform: uppercase; display: block; margin-bottom: 6px;">MMLU-Pro Physics</span>
<span style="font-size: 25px; font-weight: 800; color: #4c1d95; display: block;">74.20%</span>
<span style="font-size: 11px; color: #64748b;">371 / 500</span>
</div>
<div style="border: 1px solid #a7f3d0; padding: 16px; border-radius: 10px; background: #f0fdf4; text-align: center;">
<span style="font-size: 11px; font-weight: 700; color: #047857; text-transform: uppercase; display: block; margin-bottom: 6px;">MMLU-Pro Chemistry</span>
<span style="font-size: 25px; font-weight: 800; color: #065f46; display: block;">74.80%</span>
<span style="font-size: 11px; color: #64748b;">374 / 500</span>
</div>
<div style="border: 1px solid #fde68a; padding: 16px; border-radius: 10px; background: #fffbeb; text-align: center;">
<span style="font-size: 11px; font-weight: 700; color: #b45309; text-transform: uppercase; display: block; margin-bottom: 6px;">MMLU-Pro Average</span>
<span style="font-size: 25px; font-weight: 800; color: #92400e; display: block;">76.47%</span>
<span style="font-size: 11px; color: #64748b;">1,147 / 1,500</span>
</div>
</div>
</div>
</div>
2.1 GSM8K
The MTP-Q8 build was evaluated twice on the complete 1,319-question GSM8K test split. Across 2,638 sampled answers, the model produced 2,421 exact-answer matches, yielding a two-run pass@1 mean of 91.7741%. The two independent run scores were 91.6603% and 91.8878%, a spread of only 0.23 percentage points.
<div style="overflow-x: auto; margin-bottom: 20px;">
<table style="display: table; width: 100%; table-layout: fixed; border-collapse: collapse; overflow-wrap: anywhere; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; line-height: 1.25; border: 1px solid #cbd5e1;">
<thead>
<tr style="background: #eff6ff;">
<th style="width: 72%; padding: 6px 10px !important; border-bottom: 2px solid #0284c7; text-align: left; color: #0369a1;">Model</th>
<th style="width: 28%; padding: 6px 10px !important; border-bottom: 2px solid #0284c7; text-align: right; color: #0369a1;">Accuracy</th>
</tr>
</thead>
<tbody>
<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;"><b>MiMo-V2.5-Pro</b></td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;"><b>99.60%</b></td></tr>
<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;">Llama-3.1-405B-Instruct</td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;">96.80%</td></tr>
<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;">Llama-3.3-70B-Instruct</td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;">94.84%</td></tr>
<tr style="background: #f5f3ff;"><td style="padding: 6px 10px !important; border-bottom: 1px solid #ddd6fe; color: #5b21b6;"><b>DeepSeek-V4-Pro-Qwen3.5-9B</b></td><td style="padding: 6px 10px !important; border-bottom: 1px solid #ddd6fe; text-align: right; color: #5b21b6;"><b>94.50%</b></td></tr>
<tr><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0;">DeepSeek-V4-Pro</td><td style="padding: 6px 10px !important; border-bottom: 1px solid #e2e8f0; text-align: right;">92.60%</td></tr>
<tr style="background: #ecfdf5;"><td style="padding: 6px 10px !important; border-bottom: 1px solid #a7f3d0; color: #065f46;"><b>DeepSeek-V4-Pro-Qwen3.5-4B</b></td><td style="padding: 6px 10px !important; border-bottom: 1px solid #a7f3d0; text-align: right; color: #065f46;"><b>91.77%</b></td></tr>
<tr><td style="padding: 6px 10px !important;">DeepSeek-V3</td><td style="padding: 6px 10px !important; text-align: right;">89.30%</td></tr>
</tbody>
</table>
</div>
<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(230px, 1fr)); gap: 14px; margin-bottom: 20px; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif;">
<div style="border: 1px solid #a7f3d0; border-radius: 10px; background: #f0fdf4; padding: 14px; color: #065f46; font-size: 13px; line-height: 1.6;"><b>📈 Compact-model result</b><br>The 4B student is <b>2.73 points behind</b> the published 9B score while remaining <b>2.47 points above</b> the DeepSeek-V3 reference shown in the 9B release card.</div>
<div style="border: 1px solid #bae6fd; border-radius: 10px; background: #f0f9ff; padding: 14px; color: #0c4a6e; font-size: 13px; line-height: 1.6;"><b>🧪 4B evaluation setup</b><br>1,319 questions × 2 runs · MTP-Q8 GGUF · llama.cpp · 64K context per slot · MTP draft n=2.</div>
</div>
<div style="background: #f8fafc; border-left: 4px solid #64748b; padding: 11px 14px; border-radius: 0 8px 8px 0; color: #475569; font-size: 12px; line-height: 1.6; margin-bottom: 28px;"><b>Comparison note:</b> the 4B result is a two-run MTP-Q8 GGUF evaluation under llama.cpp, while the published 9B result was measured with an NVFP4 checkpoint under vLLM. Both use the GSM8K test split, but this table is a reported-score comparison rather than a perfectly controlled parameter-scaling experiment. The non-4B reference scores are reproduced from the 9B release card.</div>
2.2 4B vs. 9B
Both students use the same DeepSeek-V4-Pro teacher pipeline and approximately 250,000-sample mathematics and STEM mixture. The 4B release prioritizes local efficiency; the 9B release retains more capacity for difficult multi-step STEM reasoning.
<div style="overflow-x: auto; margin-bottom: 20px;">
<table style="display: table; width: 100%; min-width: 760px; table-layout: auto; border-collapse: collapse; overflow-wrap: anywhere; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; line-height: 1.25; border: 1px solid #cbd5e1;">
<thead>
<tr style="background: #f8fafc;">
<th style="padding: 10px 12px; border-bottom: 2px solid #cbd5e1; text-align: left; color: #334155;">Metric</th>
<th style="padding: 10px 12px; border-bottom: 2px solid #10b981; text-align: right; color: #047857;">DeepSeek-V4-Pro-Qwen3.5-4B</th>
<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">DeepSeek-V4-Pro-Qwen3.5-9B</th>
</tr>
</thead>
<tbody>
<tr><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">GSM8K</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">91.77%</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">94.50%</td></tr>
<tr><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MMLU-Pro Math</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">80.40%</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">92.40%</td></tr>
<tr><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MMLU-Pro Physics</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">74.20%</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.40%</td></tr>
<tr><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MMLU-Pro Chemistry</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">74.80%</td><td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.80%</td></tr>
<tr><td style="padding: 9px 12px;">MMLU-Pro Average</td><td style="padding: 9px 12px; text-align: right;">76.47%</td><td style="padding: 9px 12px; text-align: right;">90.53%</td></tr>
</tbody>
</table>
</div>
<div style="background: #f8fafc; border-left: 4px solid #64748b; padding: 11px 14px; border-radius: 0 8px 8px 0; color: #475569; font-size: 12px; line-height: 1.6; margin-bottom: 28px;"><b>Interpretation:</b> the 4B student remains strong on GSM8K, while Math is its strongest subject in the selected MMLU-Pro STEM sample. The 9B student retains higher scores across all three MMLU-Pro subjects. GSM8K and the reported 9B results use different inference builds, so this table should be read as a release overview rather than a controlled scaling experiment.</div>
2.3 MMLU-Pro: Math, Physics & Chemistry
Each model in the comparison below was evaluated on 500 questions from each of three MMLU-Pro subsets—Math, Physics, and Chemistry—for 1,500 questions per model. The table follows the same column design and subject order as the 9B release card. The newly added Claude Mythos-distilled 27B result reaches 86.20% overall and appears directly before Claude Mythos-distilled 9B.
<div style="overflow-x: auto; margin-bottom: 20px;">
<table style="display: table; width: 100%; min-width: 760px; table-layout: auto; border-collapse: collapse; overflow-wrap: anywhere; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; line-height: 1.25; border: 1px solid #cbd5e1;">
<thead>
<tr style="background: #f5f3ff;">
<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Model</th>
<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Evaluation Build</th>
<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Math</th>
<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Physics</th>
<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Chemistry</th>
<th style="padding: 10px 12px; border-bottom: 2px solid #7c3aed; text-align: right; color: #6d28d9;">Average</th>
</tr>
</thead>
<tbody>
<tr style="background: #f0fdf4;">
<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; color: #065f46;"><b>DeepSeek-V4-Pro-Qwen3.5-4B</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0;">MTP-Q8 GGUF</td>
<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right;"><b>80.40%</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right;"><b>74.20%</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right;"><b>74.80%</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #a7f3d0; text-align: right; color: #065f46;"><b>76.47%</b></td>
</tr>
<tr style="background: #f5f3ff;">
<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; color: #5b21b6;"><b>DeepSeek-V4-Pro-Qwen3.5-9B</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe;">MTP-Q8_0 GGUF</td>
<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right;"><b>92.40%</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right;"><b>89.40%</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right;"><b>89.80%</b></td>
<td style="padding: 10px 12px; border-bottom: 1px solid #ddd6fe; text-align: right; color: #5b21b6;"><b>90.53%</b></td>
</tr>
<tr>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Qwen3.5-9B</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MTP-Q8_0 GGUF</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">90.60%</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.00%</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.20%</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">89.60%</td>
</tr>
<tr>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">Claude Mythos-distilled 27B</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0;">MTP-Q8 GGUF</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">88.40%</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">87.60%</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">82.60%</td>
<td style="padding: 9px 12px; border-bottom: 1px solid #e2e8f0; text-align: right;">86.20%</td>
</tr>
<tr>
<td style="padding: 9px 12px;">Claude Mythos-distilled 9B</td>
<td style="padding: 9px 12px;">MTP-Q8_0 GGUF</td>
<td style="padding: 9px 12px; text-align: right;">86.00%</td>
<td style="padding: 9px 12px; text-align: right;">78.60%</td>
<td style="padding: 9px 12px; text-align: right;">78.60%</td>
<td style="padding: 9px 12px; text-align: right;">81.07%</td>
</tr>
</tbody>
</table>
</div>
<div style="background: #f8fafc; border-left: 4px solid #64748b; padding: 11px 14px; border-radius: 0 8px 8px 0; color: #475569; font-size: 12px; line-height: 1.6; margin-bottom: 24px;"><b>Comparison note:</b> all rows use 500 questions per subject, but the reported deployment configurations differ. The 4B run uses MTP-Q8 GGUF at 64K context, the 9B comparison rows use MTP-Q8_0 GGUF at 8K context, and Claude Mythos-distilled 27B uses MTP-Q8 GGUF at 32K context under llama.cpp.</div>
⚙️ 3. Deployment Profile
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; margin-bottom: 28px;">
<div style="background: linear-gradient(135deg, #10b981 0%, #047857 100%); padding: 13px 17px; color: white; font-weight: 700; font-size: 14px;">🚀 Built for Practical Local Reasoning</div>
<div style="padding: 18px; color: #334155; font-size: 13px; line-height: 1.7;">
<p style="margin-top: 0;">The training adapter was merged with <b>unsloth/Qwen3.5-4B</b> in BF16. The reported deployment checkpoint is an <b>MTP-enabled Q8 GGUF</b> build using two draft tokens under llama.cpp speculative decoding.</p>
<p style="margin-bottom: 0;">The 4B scale is intended for users who value lower memory use and local inference accessibility. Exact speed and memory consumption depend on quantization, hardware, context length, batch size, and runtime configuration; no cross-hardware speed claim is made in this release.</p>
</div>
</div>
🎯 4. Recommended Uses
- Grade-school and general mathematical problem solving
- Lightweight physics, chemistry, and broader STEM question answering
- Structured reasoning and analytical instruction following
- Local experiments with reasoning distillation and compact models
- MTP-enabled llama.cpp deployment research
⚠️ 5. Limitations
- This is an experimental 4B community model and remains subject to hallucinations, arithmetic mistakes, reasoning failures, and unstable behavior on difficult or underspecified tasks.
- The reported MMLU-Pro result covers fixed 500-question samples from Math, Physics, and Chemistry rather than the complete subsets; broader categories remain unevaluated.
- The GSM8K comparison combines reported results from different inference builds and backends; it is useful as a release overview, not as a perfectly controlled scaling study.
- The model received no coding-specific SFT data, and coding or tool-use generalization has not yet been established for this 4B checkpoint.
- A smaller parameter budget creates a visible gap versus the 9B release on the available MMLU-Pro STEM subsets.
- Users should independently verify high-stakes mathematical, scientific, medical, legal, or factual outputs.
📚 6. Resources, Acknowledgements & Citation
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #cbd5e1; border-radius: 14px; overflow: hidden; background: #ffffff; margin-top: 18px;">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #4c1d95 100%); padding: 16px 20px; color: white; font-weight: 800; font-size: 16px;">🌟 Open Research Release</div>
<div style="padding: 20px; color: #334155; font-size: 13px; line-height: 1.7;">
<p style="margin-top: 0;"><b>Fine-tuning guide:</b> <a href="https://github.com/R6410418/Jackrong-llm-finetuning-guide" style="color: #7c3aed; font-weight: 700; text-decoration: none;">Jackrong LLM Fine-Tuning Guide</a></p>
<p><b>Related model:</b> <a href="https://huggingface.co/Jackrong/DeepSeek-V4-Pro-Qwen3.5-9B" style="color: #7c3aed; font-weight: 700; text-decoration: none;">DeepSeek-V4-Pro-Qwen3.5-9B</a></p>
<p><b>Acknowledgements:</b> Kyle Hessling for hardware collaboration, training support, and evaluation assistance; DeepSeek for the teacher model and reasoning data source; Qwen for the base model family; Unsloth for efficient fine-tuning and conversion tooling; and the open-source inference community behind llama.cpp.</p>
<p><b>Citation:</b></p>
<pre style="background: #f5f3ff; color: #4c1d95; border: 1px solid #ddd6fe; padding: 14px; border-radius: 8px; overflow-x: auto; font-size: 12px;"><code>@misc{jackrong_deepseek_v4_pro_qwen35_4b,
title = {DeepSeek-V4-Pro-Qwen3.5-4B},
author = {Jackrong},
year = {2026},
publisher = {Hugging Face}
}</code></pre>
</div>
</div>
Run Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B-MTP-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models