Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark overview
<p align="center" <img src="https://cdn uploads.huggingface.co/production/uploads/67c2e844e0921a5410eec10a/Y5M42dCag2f7Fc6fDtV0Z.jpeg" alt="Solstice AI Banner"…
Runs locally from ~5.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00001-of-00004.gguf | GGUF | IQ4_XS | 5.1 MB | Download |
| UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00002-of-00004.gguf | GGUF | IQ4_XS | 46.04 GB | Download |
| UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00003-of-00004.gguf | GGUF | IQ4_XS | 46.20 GB | Download |
| UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00004-of-00004.gguf | GGUF | IQ4_XS | 35.04 GB | Download |
| UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00001-of-00004.gguf | GGUF | Q3_K_XL | 5.1 MB | Download |
| UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00002-of-00004.gguf | GGUF | Q3_K_XL | 45.96 GB | Download |
| UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00003-of-00004.gguf | GGUF | Q3_K_XL | 46.13 GB | Download |
| UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00004-of-00004.gguf | GGUF | Q3_K_XL | 27.30 GB | Download |
| UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00001-of-00005.gguf | GGUF | Q8_K_XL | 5.1 MB | Download |
| UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00002-of-00005.gguf | GGUF | Q8_K_XL | 45.84 GB | Download |
| UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00003-of-00005.gguf | GGUF | Q8_K_XL | 46.29 GB | Download |
| UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00004-of-00005.gguf | GGUF | Q8_K_XL | 46.07 GB | Download |
| UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00005-of-00005.gguf | GGUF | Q8_K_XL | 12.56 GB | Download |
| mmproj-BF16.gguf | GGUF | BF16 | 891.2 MB | Download |
| speculative/DSpark-drafter-vision-exp.gguf | GGUF | GGUF | 6.46 GB | Download |
Model Details
| Model ID | Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark |
|---|---|
| Author | Solstice-AI |
| Pipeline | image-text-to-text |
| License | mit |
| Base model | deepseek-ai/DeepSeek-V4-Flash-Vision-Exp |
| Last modified | 2026-09-09T01:18:48.000Z |
Model README
---
pipeline_tag: image-text-to-text
base_model:
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
license: mit
library_name: gguf
tags:
- gguf
- deepseek
- deepseek-v4
- vision
- multimodal
- image-text-to-text
- solstice-ai
- llama.cpp
- ollama
- dspark
- speculative-decoding
- draft-model
- ds4
- mmproj
---
<p align="center">
<img src="https://cdn-uploads.huggingface.co/production/uploads/67c2e844e0921a5410eec10a/Y5M42dCag2f7Fc6fDtV0Z.jpeg" alt="Solstice-AI Banner" width="100%">
</p>
<h1 align="center">DeepSeek-V4-Flash-Vision-Exp (Ultra-Optimized GGUF Suite)</h1>
<h3 align="center">Official Solstice-AI GGUF Release • Unsloth Dynamic v3.0 Quants • Lossless BF16 Vision Projector • Bundled DSpark Drafter • 1M Context</h3>
<p align="center">
<img src="https://img.shields.io/badge/org-Solstice--AI-blueviolet" alt="Solstice-AI">
<img src="https://img.shields.io/badge/license-MIT-blue" alt="License">
<img src="https://img.shields.io/badge/format-GGUF-brightgreen" alt="Format">
<img src="https://img.shields.io/badge/pipeline-image--text--to--text-success" alt="Pipeline">
<img src="https://img.shields.io/badge/context-1M%20Tokens-purple" alt="Context">
<img src="https://img.shields.io/badge/speculative-DSpark%20Drafter-red" alt="DSpark">
<img src="https://img.shields.io/badge/vision-BF16%20Projector-orange" alt="Vision">
</p>
---
Model Overview
Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark provides the curated, production-grade GGUF suite for DeepSeek's experimental multimodal foundation model deepseek-ai/DeepSeek-V4-Flash-Vision-Exp (305B total parameters, 256 routed MoE experts, 13B activated per token, 43 layers, Multi-Head Latent Attention).
This release bundles the complete multimodal stack: curated base quantizations, a standalone lossless BF16 multimodal vision projector, and the pre-aligned DSpark speculative drafter.
Artifact Matrix
| Artifact | Format | Shards | Size | Target Hardware & Role |
| :--- | :---: | :---: | :---: | :--- |
| UD-IQ4_XS | IQ4_XS (imatrix) | 4 | 136.7 GB | Recommended. Production sweet spot. Fits Apple M2/M3/M4 Ultra (192GB) or dual 80GB GPUs. |
| UD-Q3_K_XL | Q3_K_XL | 4 | 128.2 GB | High-Throughput Linear K-Quant. 15–20% faster raw dequantization. Fits 128GB–192GB setups. |
| UD-Q8_K_XL | Q8_K_XL | 5 | 161.9 GB | Bit-Exact Reference. Full precision across dense projections and router heads. |
| mmproj-BF16.gguf | BF16 | 1 | 934.5 MB | Multimodal Vision Projector. Kept in native BF16 to eliminate visual hallucinations. |
| speculative/DSpark-drafter-vision-exp.gguf | Q2_K / Q8_0 | 1 | 6.94 GB | DSpark Semi-Autoregressive Drafter. Tuned for Vision-Exp revision e46e16bf. Drives 1.4×–1.7× speculative acceleration in ds4. |
---
Architectural Breakdown & Specifications
- Total Parameters: 305B total, 13B active per token across 256 routed MoE experts (6 active per token).
- Attention Mechanism: Compressed Sparse Attention (CSA) and Multi-Head Latent Attention (MLA).
- Multimodal Vision Encoder: 32-layer Vision Transformer (ViT) with patch size 14 and downsample ratio 3.
- Context Window: 1,048,576 tokens (1M native YaRN context).
---
Citations & Acknowledgments
- Original Architecture & Weights: DeepSeek AI
- Quantization: Unsloth Dynamic v3.0
- DSpark Drafter Packaging: bleysg & Solstice-AI
Run Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models