GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

aquaduck/Qwen3.8-Flash-Next-GGUF overview

Model Card for aquaduck/Qwen3.8 Flash Next GGUF Pinned UD Q3 K XL GGUF of Qwen3.8 Flash Next qwen/qwen3.8 flash next , plus midpoint layer shards for staged / …

ggufaquaduckqwenqwen3conversationallayer-shardstext-generationlicense:otherendpoints_compatibleregion:usimatrix

Runs locally from ~10.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

5 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.ggufGGUFQ2_K_XL10.4 MBDownload
Qwen3.8-Flash-Next-UD-Q2_K_XL-00002-of-00003.ggufGGUFQ2_K_XL46.55 GBDownload
Qwen3.8-Flash-Next-UD-Q2_K_XL-00003-of-00003.ggufGGUFQ2_K_XL26.90 GBDownload
Qwen3.8-Flash-Next-UD-Q2_K_XL-layers-0-10.ggufGGUFQ2_K_XL36.95 GBDownload
Qwen3.8-Flash-Next-UD-Q2_K_XL-layers-10-48.ggufGGUFQ2_K_XL36.92 GBDownload

Model Details

Model IDaquaduck/Qwen3.8-Flash-Next-GGUF
Authoraquaduck
Pipelinetext-generation
Licenseother
Base modelqwen/qwen3.8-flash-next
Last modified2026-08-28T03:51:09.000Z

Model README

---

license: other

pipeline_tag: text-generation

library_name: gguf

base_model: qwen/qwen3.8-flash-next

base_model_relation: quantized

tags:

- gguf

- aquaduck

- qwen

- qwen3

- conversational

- layer-shards

---

Model Card for aquaduck/Qwen3.8-Flash-Next-GGUF

Pinned UD-Q3_K_XL GGUF of Qwen3.8-Flash-Next (qwen/qwen3.8-flash-next), plus midpoint layer shards for staged / multi-node loading (Aquaduck Arc layer-package-v1).

The shard files are not a new quantization. They are contiguous midpoint packages cut from the full UD-Q3_K_XL GGUF in this repo.

Model lineage

qwen/qwen3.8-flash-next

└── UD-Q3_K_XL GGUF + midpoint shards → aquaduck/Qwen3.8-Flash-Next-GGUF (this repo)

  • Base weights: https://huggingface.co/qwen/qwen3.8-flash-next (other)
  • Quantization source: https://huggingface.co/aquaduck/Qwen3.8-Flash-Next-GGUF (tag UD-Q3_K_XL)
  • This repo: full UD-Q3_K_XL GGUF (Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.gguf) and midpoint GGUF shards

Model Details

| | |

|---|---|

| Catalog id | qwen/qwen3.8-flash-next |

| Quantization | UD-Q3_K_XL |

| Parameters | 180B |

| Native context | 262,144 tokens |

| License | other |

| Base model | qwen/qwen3.8-flash-next |

| Ingest GGUF | aquaduck/Qwen3.8-Flash-Next-GGUF |

Model Description

  • Hosted by: Aquaduck (hosting and layer packaging only; base model by Qwen Team / Alibaba Cloud; GGUF quant by Unsloth / llama.cpp ecosystem)
  • Shared by: Aquaduck AI
  • Model type: Causal language model (Qwen3.8-Flash-Next), GGUF UD-Q3_K_XL
  • Language(s): Multilingual (same as base)
  • License: other (inherits from qwen/qwen3.8-flash-next)
  • Finetuned from model: N/A — not a fine-tune
  • Derived from: aquaduck/Qwen3.8-Flash-Next-GGUF ← qwen/qwen3.8-flash-next

Model Sources

  • Base model card: https://huggingface.co/qwen/qwen3.8-flash-next
  • Quantized GGUF source: https://huggingface.co/aquaduck/Qwen3.8-Flash-Next-GGUF

Files

| File | Role | Approx. size |

|------|------|--------------|

| Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.gguf | Full-model GGUF (UD-Q3_K_XL) | ~78.87 GB |

| Qwen3.8-Flash-Next-UD-Q2_K_XL-layers-0-10.gguf | Split shard (layers 0–9) | ~39.67 GB |

| Qwen3.8-Flash-Next-UD-Q2_K_XL-layers-10-48.gguf | Split shard (layers 10–47) | ~39.65 GB |

  • Total layers: 48
  • Valid split boundaries: 10

Filenames use exclusive end indices (layers-{start}-{endExclusive}).

Uses

Direct Use

  • Full Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.gguf: standard single-file UD-Q3_K_XL GGUF (llama.cpp-compatible). Use this for single-node / local runs.
  • *-layers-.gguf:* Aquaduck / Arc staged loading only. These are not drop-in complete models for stock llama.cpp.

Use the base model’s chat template (including thinking / instruct modes as documented on the base model card); other formats will not work correctly.

Out-of-Scope Use

  • Expecting any one shard to run as a complete model
  • Treating this repo as a new training run or re-quant
  • Uses prohibited by the other license or the base model’s model card guidance

Bias, Risks, and Limitations

Same capabilities, biases, and risks as qwen/qwen3.8-flash-next. UD-Q3_K_XL quantization can degrade quality vs. the original higher-precision releases. Layer sharding does not change weights beyond packaging.

Recommendations

Follow the base model’s docs for chat template, thinking vs instruct modes, and sampling. Prefer Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.gguf in this repo when you do not need staged loading.

How to Get Started

These files are meant to be loaded automatically by the Aquaduck desktop app.

  1. Download the Aquaduck desktop app and sign in.
  2. Devices connected to the internet will receive a model assignment from the model catalog (qwen/qwen3.8-flash-next).
  3. Download the model from the Home view. The app will:

- download only the assigned file from this repo (full Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.gguf or one midpoint half)

- keep that stage ready for serving

You do not need to pick files by hand, but you may for local serving. Assignment and download are driven by model catalog metadata.

The full Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.gguf is a standard UD-Q3_K_XL GGUF. The -layers-.gguf files are not.

Training Details

No training. Weights come from Qwen Team / Alibaba Cloud; UD-Q3_K_XL GGUF from aquaduck/Qwen3.8-Flash-Next-GGUF; this repo hosts that GGUF and (when split) packages it into midpoint layer shards.

Evaluation

No separate evals for the hosted GGUF or shards. See qwen/qwen3.8-flash-next.

Technical Specifications

  • Architecture: Qwen3.8-Flash-Next (~180B params, GQA (24 Q / 2 KV heads), 48 layers, hidden dim 2560)
  • Quantization: UD-Q3_K_XL
  • Packaging: pinned full UD-Q3_K_XL GGUF; optional Arc midpoint shards (*-layers-{start}-{endExclusive}.gguf)
  • Package format: layer-package-v1
  • Split: 2 stages at layer 10 (maxStages: 2)

Citation

@misc{qwen38flashnext,
    title  = {Qwen3.8-Flash-Next},
    author = {Qwen Team / Alibaba Cloud},
    year   = {2026},
    url    = {https://huggingface.co/qwen/qwen3.8-flash-next}
}

Credit:

  • The GGUF quantization source (https://huggingface.co/aquaduck/Qwen3.8-Flash-Next-GGUF)
  • llama.cpp (https://github.com/ggml-org/llama.cpp) for GGUF support

Attribution

Quantized GGUF ingested from aquaduck/Qwen3.8-Flash-Next-GGUF. Original weights: qwen/qwen3.8-flash-next. Redistributed under the base model's license.

Hosted by Aquaduck.

Model Card Contact

Aquaduck AI — https://huggingface.co/aquaduck

Run aquaduck/Qwen3.8-Flash-Next-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models