GraySoft
Projects Models About FAQ Contact Download guIDE →

aessedai/qwen3.5-397b-a17b-gguf F16 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

aessedai/qwen3.5-397b-a17b-gguf overview

Updates ### 3/10/2026 I've uploaded new quants using the new fused Up + Gate conversion, this offers up to a +10% boost in prompt processing speed from my testing.

ggufbase_model:Qwen/Qwen3.5-397B-A17Bbase_model:quantized:Qwen/Qwen3.5-397B-A17Bendpoints_compatibleregion:usimatrixconversational
aessedai/qwen3.5-397b-a17b-gguf visual
Downloads
3,364
Likes
28
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

37 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Qwen3.5-397B-A17B-IQ2_XS-00001-of-00004.gguf GGUF IQ2_XS 10.44 MB Download
Qwen3.5-397B-A17B-IQ2_XS-00002-of-00004.gguf GGUF IQ2_XS 46.23 GB Download
Qwen3.5-397B-A17B-IQ2_XS-00003-of-00004.gguf GGUF IQ2_XS 46.54 GB Download
Qwen3.5-397B-A17B-IQ2_XS-00004-of-00004.gguf GGUF IQ2_XS 30.44 GB Download
Qwen3.5-397B-A17B-IQ2_XXS-00001-of-00004.gguf GGUF IQ2_XXS 10.44 MB Download
Qwen3.5-397B-A17B-IQ2_XXS-00002-of-00004.gguf GGUF IQ2_XXS 46.43 GB Download
Qwen3.5-397B-A17B-IQ2_XXS-00003-of-00004.gguf GGUF IQ2_XXS 45.79 GB Download
Qwen3.5-397B-A17B-IQ2_XXS-00004-of-00004.gguf GGUF IQ2_XXS 21.73 GB Download
Qwen3.5-397B-A17B-IQ3_S-00001-of-00004.gguf GGUF IQ3_S 10.44 MB Download
Qwen3.5-397B-A17B-IQ3_S-00002-of-00004.gguf GGUF IQ3_S 46.56 GB Download
Qwen3.5-397B-A17B-IQ3_S-00003-of-00004.gguf GGUF IQ3_S 45.81 GB Download
Qwen3.5-397B-A17B-IQ3_S-00004-of-00004.gguf GGUF IQ3_S 44.00 GB Download
Qwen3.5-397B-A17B-IQ4_XS-00001-of-00005.gguf GGUF IQ4_XS 10.44 MB Download
Qwen3.5-397B-A17B-IQ4_XS-00002-of-00005.gguf GGUF IQ4_XS 45.87 GB Download
Qwen3.5-397B-A17B-IQ4_XS-00003-of-00005.gguf GGUF IQ4_XS 46.56 GB Download
Qwen3.5-397B-A17B-IQ4_XS-00004-of-00005.gguf GGUF IQ4_XS 44.90 GB Download
Qwen3.5-397B-A17B-IQ4_XS-00005-of-00005.gguf GGUF IQ4_XS 39.66 GB Download
Qwen3.5-397B-A17B-Q4_K_M-00001-of-00007.gguf GGUF Q4_K_M 10.44 MB Download
Qwen3.5-397B-A17B-Q4_K_M-00002-of-00007.gguf GGUF Q4_K_M 44.88 GB Download
Qwen3.5-397B-A17B-Q4_K_M-00003-of-00007.gguf GGUF Q4_K_M 45.12 GB Download
Qwen3.5-397B-A17B-Q4_K_M-00004-of-00007.gguf GGUF Q4_K_M 45.12 GB Download
Qwen3.5-397B-A17B-Q4_K_M-00005-of-00007.gguf GGUF Q4_K_M 45.12 GB Download
Qwen3.5-397B-A17B-Q4_K_M-00006-of-00007.gguf GGUF Q4_K_M 45.12 GB Download
Qwen3.5-397B-A17B-Q4_K_M-00007-of-00007.gguf GGUF Q4_K_M 2.25 GB Download
Qwen3.5-397B-A17B-Q5_K_M-00001-of-00008.gguf GGUF Q5_K_M 10.44 MB Download
Qwen3.5-397B-A17B-Q5_K_M-00002-of-00008.gguf GGUF Q5_K_M 44.49 GB Download
Qwen3.5-397B-A17B-Q5_K_M-00003-of-00008.gguf GGUF Q5_K_M 45.28 GB Download
Qwen3.5-397B-A17B-Q5_K_M-00004-of-00008.gguf GGUF Q5_K_M 45.23 GB Download
Qwen3.5-397B-A17B-Q5_K_M-00005-of-00008.gguf GGUF Q5_K_M 45.28 GB Download
Qwen3.5-397B-A17B-Q5_K_M-00006-of-00008.gguf GGUF Q5_K_M 45.23 GB Download
Qwen3.5-397B-A17B-Q5_K_M-00007-of-00008.gguf GGUF Q5_K_M 45.28 GB Download
Qwen3.5-397B-A17B-Q5_K_M-00008-of-00008.gguf GGUF Q5_K_M 2.75 GB Download
imatrix.gguf GGUF 609.72 MB Download
mmproj-Qwen3.5-397B-A17B-BF16.gguf GGUF BF16 Unknown Download
mmproj-Qwen3.5-397B-A17B-F16.gguf GGUF F16 Unknown Download
mmproj-Qwen3.5-397B-A17B-F32.gguf GGUF F32 Unknown Download
mmproj-Qwen3.5-397B-A17B-Q8_0.gguf GGUF Unknown Download

Model Details Live

Model Slug
aessedai/qwen3.5-397b-a17b-gguf
Author
AesSedai
Pipeline Task
Library
Created
2026-02-17
Last Modified
2026-03-16
Gated
No
Private
No
HF SHA
654a13fa19d259dd0acbb7c30515439ba99a2b6d
License
Unknown
Language
Unknown
Base Model
Qwen/Qwen3.5-397B-A17B

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": [
      "Qwen/Qwen3.5-397B-A17B"
    ],
    "frontmatter": {
      "base_model": [
        "Qwen/Qwen3.5-397B-A17B"
      ]
    },
    "hero_image_url": "kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\"",
    "summary": "## Updates ### 3/10/2026 I've uploaded new quants using the new fused Up + Gate conversion, this offers up to a +10% boost in prompt processing speed from my testing.",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model:\n- Qwen/Qwen3.5-397B-A17B\n---\n## Updates\n### 3/10/2026\nI've uploaded new quants using the new fused Up + Gate conversion, this offers up to a +10% boost in prompt processing speed from my testing.\n\n## Description\nThis repo contains specialized MoE-quants for Qwen3.5-397B-A17B. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, \nit should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. \nTo that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.\n\n| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |\n| :--------- | :--------- | :------- | :------- | :------- | :------- |\n| Q5_K_M | 273.55 GiB (5.93 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 3.487363 ± 0.018840 | +0.0612% | 0.004294 ± 0.000037 |\n| Q4_K_M | 227.61 GiB (4.93 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 3.495358 ± 0.018894 | +0.2905% | 0.008455 ± 0.000072 |\n| IQ4_XS | 176.99 GiB (3.84 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 3.542012 ± 0.019134 | +1.6292% | 0.022699 ± 0.000189 |\n| IQ3_S | 136.38 GiB (2.96 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 3.670508 ± 0.020012 | +5.3160% | 0.064515 ± 0.000505 |\n| IQ2_XS | 123.22 GiB (2.67 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 3.777378 ± 0.020737 | +8.3824% | 0.093718 ± 0.000714 |\n| IQ2_XXS | 113.95 GiB (2.47 BPW) | Q4_K / IQ2_XXS / IQ2_XXS / IQ3_XXS | 3.879226 ± 0.021468 | +11.3047% | 0.126000 ± 0.000893 |\n\n![kld_graph](kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\")\n![ppl_graph](kld_data/02_ppl_vs_filesize.png \"Chart showing Pareto PPL analysis of quants\")",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "base_model:Qwen/Qwen3.5-397B-A17B",
    "base_model:quantized:Qwen/Qwen3.5-397B-A17B",
    "endpoints_compatible",
    "region:us",
    "imatrix",
    "conversational"
  ],
  "likes": 28,
  "downloads": 3364,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-16T17:28:57.000Z",
  "created_at": "2026-02-17T05:40:48.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "6993ff60e168313e5cce828a",
  "id": "AesSedai/Qwen3.5-397B-A17B-GGUF",
  "modelId": "AesSedai/Qwen3.5-397B-A17B-GGUF",
  "sha": "654a13fa19d259dd0acbb7c30515439ba99a2b6d",
  "createdAt": "2026-02-17T05:40:48.000Z",
  "lastModified": "2026-03-16T17:28:57.000Z",
  "author": "AesSedai",
  "downloads": 3364,
  "likes": 28,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 49
}