ggml-org/LongCat-Flash-Chat-GGUF overview
| Quant | Size | Mixture | PPL | 1 Mean PPL Q /PPL base | KLD | | : | : | : | : | : | : | | Q8 0 | 557.39 GiB 8.51 BPW | Q8 0 | 3.275749 ± 0.017863 | +0.2918% …
Runs locally from ~5.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| IQ1_S/LongCat-Flash-Chat-IQ1_S-00001-of-00004.gguf | GGUF | IQ1_S | 5.2 MB | Download |
| IQ1_S/LongCat-Flash-Chat-IQ1_S-00002-of-00004.gguf | GGUF | IQ1_S | 46.41 GB | Download |
| IQ1_S/LongCat-Flash-Chat-IQ1_S-00003-of-00004.gguf | GGUF | IQ1_S | 46.36 GB | Download |
| IQ1_S/LongCat-Flash-Chat-IQ1_S-00004-of-00004.gguf | GGUF | IQ1_S | 13.41 GB | Download |
| imatrix.gguf | GGUF | GGUF | 805.1 MB | Download |
Model Details
Model README
---
base_model:
- meituan-longcat/LongCat-Flash-Chat
---
| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |
| :---- | :-------------------- | :------ | :------------------- | :------------------------ | :------------------ |
| Q8_0 | 557.39 GiB (8.51 BPW) | Q8_0 | 3.275749 ± 0.017863 | +0.2918% | 0.003764 ± 0.000045 |
| IQ1_S | 106.19 GiB (1.62 BPW) | IQ1_S | 11.010332 ± 0.077291 | +237.0971% | 1.331246 ± 0.004350 |
Run ggml-org/LongCat-Flash-Chat-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models