GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark overview

<p align="center" <img src="https://cdn uploads.huggingface.co/production/uploads/67c2e844e0921a5410eec10a/Y5M42dCag2f7Fc6fDtV0Z.jpeg" alt="Solstice AI Banner"…

ggufdeepseekdeepseek-v4visionmultimodalimage-text-to-textsolstice-aillama.cppollamadsparkspeculative-decodingdraft-modelds4mmprojbase_model:deepseek-ai/DeepSeek-V4-Flash-Vision-Expbase_model:quantized:deepseek-ai/DeepSeek-V4-Flash-Vision-Explicense:mitendpoints_compatibleregion:usimatrixconversational

Runs locally from ~5.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
317
Likes
1
Pipeline
image-text-to-text

Repository Files & Downloads

15 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00001-of-00004.ggufGGUFIQ4_XS5.1 MBDownload
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00002-of-00004.ggufGGUFIQ4_XS46.04 GBDownload
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00003-of-00004.ggufGGUFIQ4_XS46.20 GBDownload
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00004-of-00004.ggufGGUFIQ4_XS35.04 GBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00001-of-00004.ggufGGUFQ3_K_XL5.1 MBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00002-of-00004.ggufGGUFQ3_K_XL45.96 GBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00003-of-00004.ggufGGUFQ3_K_XL46.13 GBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00004-of-00004.ggufGGUFQ3_K_XL27.30 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00001-of-00005.ggufGGUFQ8_K_XL5.1 MBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00002-of-00005.ggufGGUFQ8_K_XL45.84 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00003-of-00005.ggufGGUFQ8_K_XL46.29 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00004-of-00005.ggufGGUFQ8_K_XL46.07 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00005-of-00005.ggufGGUFQ8_K_XL12.56 GBDownload
mmproj-BF16.ggufGGUFBF16891.2 MBDownload
speculative/DSpark-drafter-vision-exp.ggufGGUFGGUF6.46 GBDownload

Model Details

Model IDSolstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark
AuthorSolstice-AI
Pipelineimage-text-to-text
Licensemit
Base modeldeepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Last modified2026-09-09T01:18:48.000Z

Model README

---

pipeline_tag: image-text-to-text

base_model:

  • deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

license: mit

library_name: gguf

tags:

  • gguf
  • deepseek
  • deepseek-v4
  • vision
  • multimodal
  • image-text-to-text
  • solstice-ai
  • llama.cpp
  • ollama
  • dspark
  • speculative-decoding
  • draft-model
  • ds4
  • mmproj

---

<p align="center">

<img src="https://cdn-uploads.huggingface.co/production/uploads/67c2e844e0921a5410eec10a/Y5M42dCag2f7Fc6fDtV0Z.jpeg" alt="Solstice-AI Banner" width="100%">

</p>

<h1 align="center">DeepSeek-V4-Flash-Vision-Exp (Ultra-Optimized GGUF Suite)</h1>

<h3 align="center">Official Solstice-AI GGUF Release &bull; Unsloth Dynamic v3.0 Quants &bull; Lossless BF16 Vision Projector &bull; Bundled DSpark Drafter &bull; 1M Context</h3>

<p align="center">

<img src="https://img.shields.io/badge/org-Solstice--AI-blueviolet" alt="Solstice-AI">

<img src="https://img.shields.io/badge/license-MIT-blue" alt="License">

<img src="https://img.shields.io/badge/format-GGUF-brightgreen" alt="Format">

<img src="https://img.shields.io/badge/pipeline-image--text--to--text-success" alt="Pipeline">

<img src="https://img.shields.io/badge/context-1M%20Tokens-purple" alt="Context">

<img src="https://img.shields.io/badge/speculative-DSpark%20Drafter-red" alt="DSpark">

<img src="https://img.shields.io/badge/vision-BF16%20Projector-orange" alt="Vision">

</p>

---

Model Overview

Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark provides the curated, production-grade GGUF suite for DeepSeek's experimental multimodal foundation model deepseek-ai/DeepSeek-V4-Flash-Vision-Exp (305B total parameters, 256 routed MoE experts, 13B activated per token, 43 layers, Multi-Head Latent Attention).

This release bundles the complete multimodal stack: curated base quantizations, a standalone lossless BF16 multimodal vision projector, and the pre-aligned DSpark speculative drafter.

Artifact Matrix

| Artifact | Format | Shards | Size | Target Hardware & Role |

| :--- | :---: | :---: | :---: | :--- |

| UD-IQ4_XS | IQ4_XS (imatrix) | 4 | 136.7 GB | Recommended. Production sweet spot. Fits Apple M2/M3/M4 Ultra (192GB) or dual 80GB GPUs. |

| UD-Q3_K_XL | Q3_K_XL | 4 | 128.2 GB | High-Throughput Linear K-Quant. 15–20% faster raw dequantization. Fits 128GB–192GB setups. |

| UD-Q8_K_XL | Q8_K_XL | 5 | 161.9 GB | Bit-Exact Reference. Full precision across dense projections and router heads. |

| mmproj-BF16.gguf | BF16 | 1 | 934.5 MB | Multimodal Vision Projector. Kept in native BF16 to eliminate visual hallucinations. |

| speculative/DSpark-drafter-vision-exp.gguf | Q2_K / Q8_0 | 1 | 6.94 GB | DSpark Semi-Autoregressive Drafter. Tuned for Vision-Exp revision e46e16bf. Drives 1.4×–1.7× speculative acceleration in ds4. |

---

Architectural Breakdown & Specifications

  • Total Parameters: 305B total, 13B active per token across 256 routed MoE experts (6 active per token).
  • Attention Mechanism: Compressed Sparse Attention (CSA) and Multi-Head Latent Attention (MLA).
  • Multimodal Vision Encoder: 32-layer Vision Transformer (ViT) with patch size 14 and downsample ratio 3.
  • Context Window: 1,048,576 tokens (1M native YaRN context).

---

Citations & Acknowledgments

Run Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF-DSpark with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models