avlp12/qwen3.5-35b-a3b-alis-ultra-slim-gguf Slim GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
avlp12/qwen3.5-35b-a3b-alis-ultra-slim-gguf overview
MoE-aware mixed-precision quantization. Best quality under 13 GiB. Pure quantization — zero fine-tuning. This is a direct quantization of the official Qwen3.5-35B-A3B weights. No merge, no fine-tune, no LoRA, no RLHF modifications. The original model's capabilities are preserved exactly as Qwen released them — only the numerical precision is optimized. PPL 7.009 — beats APEX Mini on both perplexity (7.009 vs 7.048) and HellaSwag (78.00% vs 76.75%). The best stock-llama.cpp Qwen3.5-35B GGUF under 13 GiB. Also available: Alis Ultra (14.26 GiB, PPL 6.96, HellaSwag 78.50%) — for maximum quality when memory allows.
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.5-35B-A3B-Alis-Ultra-Slim.gguf | GGUF | — | 12.88 GB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": "Qwen/Qwen3.5-35B-A3B",
"base_model_relation": "quantized",
"license": "apache-2.0",
"license_link": "https://huggingface.co/Qwen/Qwen3.5-35B-A3B/blob/main/LICENSE",
"pipeline_tag": "text-generation",
"tags": [
"gguf",
"quantized",
"imatrix",
"qwen3_5_moe",
"mixture-of-experts",
"apple-silicon",
"conversational",
"moe-aware",
"16gb"
],
"quantized_by": "avlp12",
"frontmatter": {
"base_model": "Qwen/Qwen3.5-35B-A3B",
"base_model_relation": "quantized",
"license": "apache-2.0",
"license_link": "https://huggingface.co/Qwen/Qwen3.5-35B-A3B/blob/main/LICENSE",
"pipeline_tag": "text-generation",
"tags": [
"gguf",
"quantized",
"imatrix",
"qwen3_5_moe",
"mixture-of-experts",
"apple-silicon",
"conversational",
"moe-aware",
"16gb"
],
"quantized_by": "avlp12"
},
"hero_image_url": "",
"summary": "**MoE-aware mixed-precision quantization. Best quality under 13 GiB.** > **Pure quantization — zero fine-tuning.** This is a direct quantization of the official Qwen3.5-35B-A3B weights. No merge, no fine-tune, no LoRA, no RLHF modifications. The original model's capabilities are preserved exactly as Qwen released them — only the numerical precision is optimized. PPL 7.009 — beats APEX Mini on both perplexity (7.009 vs 7.048) and HellaSwag (78.00% vs 76.75%). The best stock-llama.cpp Qwen3.5-35B GGUF under 13 GiB. > **Also available:** Alis Ultra (14.26 GiB, PPL 6.96, HellaSwag 78.50%) — for maximum quality when memory allows.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model: Qwen/Qwen3.5-35B-A3B\nbase_model_relation: quantized\nlicense: apache-2.0\nlicense_link: https://huggingface.co/Qwen/Qwen3.5-35B-A3B/blob/main/LICENSE\npipeline_tag: text-generation\ntags:\n - gguf\n - quantized\n - imatrix\n - qwen3_5_moe\n - mixture-of-experts\n - apple-silicon\n - conversational\n - moe-aware\n - 16gb\nquantized_by: avlp12\n---\n\n# Alis Ultra Slim — Qwen3.5-35B-A3B GGUF (12.87 GiB)\n\n**MoE-aware mixed-precision quantization. Best quality under 13 GiB.**\n\n> **Pure quantization — zero fine-tuning.** This is a direct quantization of the official Qwen3.5-35B-A3B weights. No merge, no fine-tune, no LoRA, no RLHF modifications. The original model's capabilities are preserved exactly as Qwen released them — only the numerical precision is optimized.\n\nPPL 7.009 — beats APEX Mini on both perplexity (7.009 vs 7.048) and HellaSwag (78.00% vs 76.75%). The best stock-llama.cpp Qwen3.5-35B GGUF under 13 GiB.\n\n> **Also available:** [Alis Ultra](https://huggingface.co/avlp12/Qwen3.5-35B-A3B-Alis-Ultra-GGUF) (14.26 GiB, PPL 6.96, HellaSwag 78.50%) — for maximum quality when memory allows.\n\n## Quick Start\n\n```bash\n# Download\nhf download avlp12/Qwen3.5-35B-A3B-Alis-Ultra-Slim-GGUF \\\n Qwen3.5-35B-A3B-Alis-Ultra-Slim.gguf --local-dir ./model\n\n# Chat\nllama-cli -m ./model/Qwen3.5-35B-A3B-Alis-Ultra-Slim.gguf \\\n --conversation -ngl 99\n\n# Server\nllama-server -m ./model/Qwen3.5-35B-A3B-Alis-Ultra-Slim.gguf \\\n --host 0.0.0.0 --port 8080 -ngl 99 \\\n --temp 0.6 --top-p 0.95 --top-k 20 \\\n -fa on --jinja\n```\n\nCompatible with stock llama.cpp, llama-server, and any GGUF-compatible backend. No custom forks required.\n\n---\n\n## Alis Ultra Slim vs Ultra — Which One?\n\n| | **Ultra Slim** (this repo) | **Ultra** ([link](https://huggingface.co/avlp12/Qwen3.5-35B-A3B-Alis-Ultra-GGUF)) |\n|---|---|---|\n| **Size** | 12.87 GiB (3.19 BPW) | 14.26 GiB (3.53 BPW) |\n| **PPL** | 7.009 ± 0.045 | **6.955** ± 0.045 |\n| **HellaSwag** | 78.00% | **78.50%** |\n| **Base type** | Q2_K | Q3_K_M |\n| **Best for** | 16 GiB constrained (M4, RTX 4060 Ti 16GB) | 24 GiB+ devices (M4 Pro/Max, RTX 4090) |\n| **Key advantage** | Best quality under 13 GiB | F16-matching accuracy |\n\n**The core difference:** Ultra Slim uses Q2_K as the base quantization type, keeping middle-layer gate/up experts at the K-quant floor (~2.6 BPW). Ultra uses Q3_K_M as the base, which allows llama.cpp's imatrix to auto-promote sensitive tensors — providing an extra quality buffer at the cost of 1.4 GiB.\n\nBoth variants share identical treatment of critical tensors: shared experts at Q8_0, attention at Q4_K/Q5_K, SSM at Q6_K, and ffn_down_exps at Q3_K.\n\n---\n\n## Why Ultra Slim over APEX Mini?\n\n| | **Alis Ultra Slim** | **APEX Mini** |\n|---|---|---|\n| Size | 12.87 GiB | **12.33 GiB** |\n| PPL | **7.009** | 7.048 |\n| HellaSwag | **78.00%** | 76.75% |\n| Compatibility | stock llama.cpp | stock llama.cpp |\n\nUltra Slim is 0.54 GiB larger but wins on both PPL (-0.04) and HellaSwag (+1.25%p). The accuracy gap is significant — 78.00% vs 76.75% is well outside the confidence interval overlap.\n\n---\n\n## Benchmarks\n\nAll measurements on M3 Ultra Mac Studio (512 GB, 80 GPU cores). PPL on wikitext-2-raw test set, context 512, 580 chunks. HellaSwag 0-shot, 400 tasks. **All values directly measured by us on the same hardware.**\n\n### Quality\n\n| Model | Size (GiB) | BPW | PPL | HellaSwag |\n|---|---|---|---|---|\n| F16 (baseline) | 64.60 | 16.00 | 6.537 | 78.50% |\n| Unsloth Q3_K_M (Dynamic) | 15.22 | 3.77 | 6.779 | 78.50% |\n| Alis Ultra | 14.26 | 3.53 | 6.955 | 78.50% |\n| **Alis Ultra Slim** | **12.87** | **3.19** | **7.009** | **78.00%** |\n| APEX Mini | 12.33 | 3.06 | 7.048 | 76.75% |\n| Unsloth IQ2_XXS | 9.91 | 2.46 | 7.519 | 77.00% |\n\n### Key Takeaways\n\n- **Alis Ultra Slim beats APEX Mini** on both PPL (7.009 vs 7.048) and HellaSwag (78.00% vs 76.75%)\n- HellaSwag 78.00% is only 0.50%p below F16 baseline (78.50%) — remarkable for 5× compression\n- Even Unsloth IQ2_XXS (77.00%) outscores APEX Mini (76.75%) on HellaSwag despite being 2.4 GiB smaller\n\n### Speed (M3 Ultra, llama-bench)\n\n| Model | pp512 (tok/s) | tg128 (tok/s) |\n|---|---|---|\n| Alis Ultra (14.26 GiB) | 2,239 ± 8.58 | 85.62 ± 1.22 |\n| Alis Ultra Slim (12.87 GiB) | 2,204 ± 7.39 | 81.95 ± 0.76 |\n\n---\n\n## Quantization Strategy\n\n### Tensor Classification & Layer Gradient\n\nBased on APEX's MoE tensor role analysis and Unsloth's 121-configuration KL divergence study:\n\n| Tensor | Edge (L0-4, L35-39) | Middle (L5-34) |\n|---|---|---|\n| **ffn_down_exps** (most sensitive) | q3_K | q3_K |\n| **ffn_gate/up_exps** | q3_K | q2_K (base) |\n| **Shared experts** (every-token path) | Q8_0 | Q8_0 |\n| **Attention Q/K** | q4_K | q4_K |\n| **Attention V** | q5_K | q5_K |\n| **Attention O** | q4_K | q4_K |\n| **SSM out** | Q6_K | Q6_K |\n| **Embedding** | Q4_K | — |\n| **Output** | Q5_K | — |\n\n### Design Principles\n\n1. **K-quant over IQ for MoE experts.** Routed expert weights have near-Gaussian distributions (kurtosis 3.41). K-quant block-scaling outperforms IQ codebooks designed for heavy-tailed distributions.\n\n2. **Protect ffn_down, compress gate/up.** ffn_down_exps is consistently the most sensitive expert tensor. We keep it at Q3_K while allowing gate/up to drop to Q2_K in middle layers.\n\n3. **Edge-layer gradient.** Layers 0-4 and 35-39 (nearest to embedding/output) get higher precision.\n\n4. **Shared experts are sacred.** With kurtosis 13.10 (4× routed experts) and 100% activation rate, shared experts stay at Q8_0.\n\n### Imatrix\n\n14,062 chunks from 76,447 calibration samples across 6 domains:\n\n- General instruction (Alpaca, 52K) — broad language coverage\n- Math reasoning (GSM8K, 7.4K) — numerical precision\n- Wikipedia (wikitext-2-raw train, ~15K) — factual knowledge\n- Korean language (1,000) — multilingual preservation\n- Code (500) — syntax structure\n- Tool-call JSON (200) — structured output accuracy\n\n---\n\n## Vision Support\n\nThis GGUF contains **text weights only** (733 tensors). Qwen3.5-35B-A3B is a multimodal model, but the vision encoder must be loaded separately as an `mmproj` file.\n\n### Setup\n\n```bash\n# 1. Download mmproj (choose one — F16 recommended for quality/size balance)\nhf download unsloth/Qwen3.5-35B-A3B-GGUF mmproj-F16.gguf --local-dir ./model\n# Alternatives: mmproj-BF16.gguf (903 MB) or mmproj-F32.gguf (1.79 GB)\n\n# 2. Run with vision — Server mode\nllama-server \\\n -m ./model/Qwen3.5-35B-A3B-Alis-Ultra-Slim.gguf \\\n --mmproj ./model/mmproj-F16.gguf \\\n --host 0.0.0.0 --port 8080 -ngl 99 \\\n -fa on --jinja\n\n# 3. Run with vision — Interactive CLI\nllama-mtmd-cli \\\n -m ./model/Qwen3.5-35B-A3B-Alis-Ultra-Slim.gguf \\\n --mmproj ./model/mmproj-F16.gguf \\\n -ngl 99 --jinja\n```\n\n### Memory Impact\n\nThe mmproj-F16 adds ~899 MiB to VRAM usage. Total for Ultra Slim + mmproj-F16: ~13.8 GiB — still fits in 16 GiB devices with room for KV cache.\n\n### Compatibility Notes\n\n- **llama.cpp**: Full support via `--mmproj` flag (llama-server, llama-mtmd-cli)\n- **Ollama**: Not currently supported — Qwen3.5 GGUF requires separate mmproj files which Ollama does not handle\n- **LM Studio**: Check for Qwen3.5 VLM support in your version\n- mmproj files are interchangeable across all Qwen3.5-35B-A3B quantizations (Alis, Unsloth, APEX, etc.)\n\n## Thinking Mode\n\nQwen3.5 supports thinking/non-thinking. To disable:\n```\n--chat-template-kwargs '{\"enable_thinking\":false}'\n```\n\n## Technical Details\n\n- **Base Model:** Qwen/Qwen3.5-35B-A3B (35B total, 3B active per token, 256 experts, 8 active + 1 shared)\n- **Architecture:** Qwen3.5-MoE with Gated Delta Networks\n- **Quantization Tool:** llama.cpp build 8770, `--tensor-type` per-layer overrides\n- **Imatrix:** 14,062 chunks, PPL 4.5318 on calibration data\n- **Source:** F16 GGUF → converted from HuggingFace BF16\n\n## Acknowledgments\n\n[Qwen Team](https://github.com/QwenLM/Qwen3.5) · [APEX/LocalAI](https://github.com/mudler/apex-quant) (MoE tensor classification) · [Unsloth](https://unsloth.ai) (KLD study) · [llama.cpp](https://github.com/ggml-org/llama.cpp) · [ik_llama.cpp](https://github.com/ikawrakow/ik_llama.cpp) (IQK research)\n\n## License\n\nApache 2.0 — same as the base model.\n",
"related_quantizations": []
},
"tags": [
"gguf",
"quantized",
"imatrix",
"qwen3_5_moe",
"mixture-of-experts",
"apple-silicon",
"conversational",
"moe-aware",
"16gb",
"text-generation",
"base_model:Qwen/Qwen3.5-35B-A3B",
"base_model:quantized:Qwen/Qwen3.5-35B-A3B",
"license:apache-2.0",
"endpoints_compatible",
"region:us"
],
"likes": 0,
"downloads": 344,
"gated": false,
"private": false,
"last_modified": "2026-04-14T11:26:49.000Z",
"created_at": "2026-04-14T11:24:46.000Z",
"pipeline_tag": "text-generation",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69de23fe0042d7540a32bc27",
"id": "avlp12/Qwen3.5-35B-A3B-Alis-Ultra-Slim-GGUF",
"modelId": "avlp12/Qwen3.5-35B-A3B-Alis-Ultra-Slim-GGUF",
"sha": "576056a12e81d138bda7d54cc77f8fe568e96075",
"createdAt": "2026-04-14T11:24:46.000Z",
"lastModified": "2026-04-14T11:26:49.000Z",
"author": "avlp12",
"downloads": 344,
"likes": 0,
"gated": false,
"private": false,
"pipeline_tag": "text-generation",
"library_name": "",
"siblings_count": 3
}