GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF overview

DeepSeek V4 Flash Vision Exp UltraOptimised GGUF Curated Unsloth Dynamic 3.0 quantisations of deepseek ai/DeepSeek V4 Flash Vision Exp https://huggingface.co/d…

ggufdeepseekdeepseek-v4visionmultimodalimage-text-to-textsolstice-aianvilturboquantsovereign-aillama.cppollamaendpoints_compatiblemmprojunslothbase_model:deepseek-ai/DeepSeek-V4-Flash-Vision-Explicense:mitregion:usimatrixconversational

Runs locally from ~5.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
image-text-to-text

Repository Files & Downloads

39 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
UD-IQ1_M/DeepSeek-V4-Flash-Vision-Exp-UD-IQ1_M-00001-of-00003.ggufGGUFIQ1_M5.1 MBDownload
UD-IQ1_M/DeepSeek-V4-Flash-Vision-Exp-UD-IQ1_M-00002-of-00003.ggufGGUFIQ1_M46.00 GBDownload
UD-IQ1_M/DeepSeek-V4-Flash-Vision-Exp-UD-IQ1_M-00003-of-00003.ggufGGUFIQ1_M34.90 GBDownload
UD-IQ1_S/DeepSeek-V4-Flash-Vision-Exp-UD-IQ1_S-00001-of-00003.ggufGGUFIQ1_S5.1 MBDownload
UD-IQ1_S/DeepSeek-V4-Flash-Vision-Exp-UD-IQ1_S-00002-of-00003.ggufGGUFIQ1_S46.56 GBDownload
UD-IQ1_S/DeepSeek-V4-Flash-Vision-Exp-UD-IQ1_S-00003-of-00003.ggufGGUFIQ1_S30.21 GBDownload
UD-IQ2_XXS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ2_XXS-00001-of-00003.ggufGGUFIQ2_XXS5.1 MBDownload
UD-IQ2_XXS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ2_XXS-00002-of-00003.ggufGGUFIQ2_XXS46.19 GBDownload
UD-IQ2_XXS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ2_XXS-00003-of-00003.ggufGGUFIQ2_XXS38.27 GBDownload
UD-IQ3_S/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_S-00001-of-00004.ggufGGUFIQ3_S5.1 MBDownload
UD-IQ3_S/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_S-00002-of-00004.ggufGGUFIQ3_S46.34 GBDownload
UD-IQ3_S/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_S-00003-of-00004.ggufGGUFIQ3_S46.45 GBDownload
UD-IQ3_S/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_S-00004-of-00004.ggufGGUFIQ3_S13.75 GBDownload
UD-IQ3_XXS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_XXS-00001-of-00004.ggufGGUFIQ3_XXS5.1 MBDownload
UD-IQ3_XXS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_XXS-00002-of-00004.ggufGGUFIQ3_XXS46.09 GBDownload
UD-IQ3_XXS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_XXS-00003-of-00004.ggufGGUFIQ3_XXS46.04 GBDownload
UD-IQ3_XXS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ3_XXS-00004-of-00004.ggufGGUFIQ3_XXS3.79 GBDownload
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00001-of-00004.ggufGGUFIQ4_XS5.1 MBDownload
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00002-of-00004.ggufGGUFIQ4_XS46.04 GBDownload
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00003-of-00004.ggufGGUFIQ4_XS46.20 GBDownload
UD-IQ4_XS/DeepSeek-V4-Flash-Vision-Exp-UD-IQ4_XS-00004-of-00004.ggufGGUFIQ4_XS35.04 GBDownload
UD-Q2_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q2_K_XL-00001-of-00003.ggufGGUFQ2_K_XL5.1 MBDownload
UD-Q2_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q2_K_XL-00002-of-00003.ggufGGUFQ2_K_XL46.04 GBDownload
UD-Q2_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q2_K_XL-00003-of-00003.ggufGGUFQ2_K_XL44.14 GBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00001-of-00004.ggufGGUFQ3_K_XL5.1 MBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00002-of-00004.ggufGGUFQ3_K_XL45.96 GBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00003-of-00004.ggufGGUFQ3_K_XL46.13 GBDownload
UD-Q3_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q3_K_XL-00004-of-00004.ggufGGUFQ3_K_XL27.30 GBDownload
UD-Q4_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q4_K_XL-00001-of-00005.ggufGGUFQ4_K_XL5.1 MBDownload
UD-Q4_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q4_K_XL-00002-of-00005.ggufGGUFQ4_K_XL45.57 GBDownload
UD-Q4_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q4_K_XL-00003-of-00005.ggufGGUFQ4_K_XL45.62 GBDownload
UD-Q4_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q4_K_XL-00004-of-00005.ggufGGUFQ4_K_XL46.57 GBDownload
UD-Q4_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q4_K_XL-00005-of-00005.ggufGGUFQ4_K_XL6.68 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00001-of-00005.ggufGGUFQ8_K_XL5.1 MBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00002-of-00005.ggufGGUFQ8_K_XL45.84 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00003-of-00005.ggufGGUFQ8_K_XL46.29 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00004-of-00005.ggufGGUFQ8_K_XL46.07 GBDownload
UD-Q8_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q8_K_XL-00005-of-00005.ggufGGUFQ8_K_XL12.56 GBDownload
mmproj-BF16.ggufGGUFBF16891.2 MBDownload

Model Details

Model IDSolstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF
AuthorSolstice-AI
Pipelineimage-text-to-text
Licensemit
Base model
Last modified2026-09-07T04:21:52.000Z

Model README

---

tags:

  • gguf
  • deepseek
  • deepseek-v4
  • vision
  • multimodal
  • image-text-to-text
  • solstice-ai
  • anvil
  • turboquant
  • sovereign-ai
  • llama.cpp
  • ollama
  • endpoints_compatible
  • vision
  • mmproj
  • unsloth
  • base_model:deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

license: mit

library_name: gguf

pipeline_tag: image-text-to-text

region: us

---

DeepSeek-V4-Flash-Vision-Exp UltraOptimised GGUF

Curated Unsloth Dynamic 3.0 quantisations of deepseek-ai/DeepSeek-V4-Flash-Vision-Exp — the experimental multimodal variant of DeepSeek-V4-Flash with vision support.

All quants use Unsloth's UD 3.0 methodology with a high-quality imatrix calibration dataset refined for agentic coding, chat, and multilingual performance. Each quant is model-specific and layer-optimized for maximum quality at its size.

Base model: 305B total parameters, 284B MoE (13B active per token), 1M context, vision encoder included.

Quantisation Tiers

| Quant | Size | RAM Required | Description |

|-------|------|-------------|-------------|

| UD-Q8_K_XL | ~151 GB | 160+ GB | Near-lossless — full precision equivalent |

| UD-Q4_K_XL | ~145 GB | 155+ GB | Primary workhorse — only 7 GB smaller than Q8 |

| UD-Q3_K_XL | ~119 GB | 128+ GB | Fits 128 GB RAM Macs — strong community pick |

| UD-Q2_K_XL | ~90 GB | 96+ GB | Fits 96 GB RAM — still usable for agentic tasks |

| UD-IQ4_XS | ~127 GB | 135+ GB | Compact 4-bit with importance matrix |

| UD-IQ3_S | ~106 GB | 115+ GB | Balanced 3-bit |

| UD-IQ3_XXS | ~96 GB | 105+ GB | Smaller 3-bit |

| UD-IQ2_XXS | ~85 GB | 95+ GB | Ultra-compact 2-bit |

| UD-IQ1_S | ~77 GB | 85+ GB | Extreme compression — general knowledge only |

| UD-IQ1_M | ~81 GB | 90+ GB | Extreme compression — general knowledge only |

Vision Support

Both vision projector files are included:

  • mmproj-BF16.gguf — BF16 precision (recommended)
  • mmproj-F16.gguf — FP16 precision

Quick Start

# Download with HF CLI
hf download Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF \
  UD-Q4_K_XL/* mmproj-BF16.gguf

# Run with llama.cpp (needs build >= 2026-06-04 for Gemma4/Vision-Exp support)
./llama-server \
  -m UD-Q4_K_XL/DeepSeek-V4-Flash-Vision-Exp-UD-Q4_K_XL-00001-of-00005.gguf \
  --mmproj mmproj-BF16.gguf \
  --ctx-size 32768 \
  --n-gpu-layers 99

Recommended Tiers

For most users: UD-Q4_K_XL or UD-Q3_K_XL

  • Q4 is only ~6 GB smaller than Q8 but significantly more practical to serve
  • Q3 fits 128 GB unified memory Macs (M2/M3/M4 Ultra)

For maximum quality: UD-Q8_K_XL

  • Near-lossless — the model is already ~4.25 bit in its experts, so Q8 only upscales the shared dense layers

For constrained hardware: UD-Q2_K_XL

  • Last tier before the Divergence-300 @32 quality cliff — still functional for agentic use

Quantisation Details

  • Method: Unsloth Dynamic v3.0 — model-specific layer optimization
  • Calibration: High-quality imatrix dataset refined for agentic coding, chat, and multilingual performance
  • No QAT/QAD: Pure post-training quantization — no overfitting risk
  • MTP module: Included in Q3+ quants; removed in smaller quants to save ~500 MB

Hardware Requirements

| Config | Min RAM/VRAM | Example Hardware |

|--------|-------------|-----------------|

| UD-Q8_K_XL | 160 GB | 2× M2 Ultra, 4× A100 80GB |

| UD-Q4_K_XL | 155 GB | 2× M2 Ultra, 4× A100 80GB |

| UD-Q3_K_XL | 128 GB | 1× M2/M3/M4 Ultra 128GB |

| UD-Q2_K_XL | 96 GB | 1× M2/M3 Max 96GB |

Credits

Quantisation by Unsloth — Dynamic v3.0 methodology.

Curated and hosted by Solstice-AI.

License

MIT — inherited from deepseek-ai/DeepSeek-V4-Flash-Vision-Exp.

Run Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-UltraOptimised-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models