GraySoft
Projects Models About FAQ Contact Download guIDE →

aessedai/step-3.5-flash-gguf IQ4_XS GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

aessedai/step-3.5-flash-gguf overview

This repo contains specialized MoE-quants for Step-3.5-Flash. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors. | Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD | | :--------- | :--------- | :------- | :------- | :------- | :------- | | Q4KM | 113.82 GiB (4.96 BPW) | Q80 / Q4K / Q4K / Q5K | 4.718049 ± 0.030373 | +0.3762% | 0.015464 ± 0.000133 | | IQ4XS | 88.90 GiB (3.88 BPW) | Q80 / IQ3S / IQ3S / IQ4XS | 4.822499 ± 0.031236 | +2.5984% | 0.042753 ± 0.000301 | | IQ3XXS | 73.10 GiB (3.19 BPW) | Q6K / IQ3XXS / IQ3XXS / IQ3XXS | 4.882908 ± 0.031560 | +3.8836% | 0.078681 ± 0.000506 | !kldgraph !pplgraph

ggufbase_model:stepfun-ai/Step-3.5-Flashbase_model:quantized:stepfun-ai/Step-3.5-Flashendpoints_compatibleregion:usimatrixconversational
aessedai/step-3.5-flash-gguf visual
Downloads
1,090
Likes
14
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

9 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Step-3.5-Flash-IQ3_XXS-00001-of-00003.gguf GGUF IQ3_XXS 4.99 MB Download
Step-3.5-Flash-IQ3_XXS-00002-of-00003.gguf GGUF IQ3_XXS 60.06 GB Download
Step-3.5-Flash-IQ3_XXS-00003-of-00003.gguf GGUF IQ3_XXS 13.04 GB Download
Step-3.5-Flash-IQ4_XS-00001-of-00003.gguf GGUF IQ4_XS 4.99 MB Download
Step-3.5-Flash-IQ4_XS-00002-of-00003.gguf GGUF IQ4_XS 59.97 GB Download
Step-3.5-Flash-IQ4_XS-00003-of-00003.gguf GGUF IQ4_XS 28.93 GB Download
Step-3.5-Flash-Q4_K_M-00001-of-00003.gguf GGUF Q4_K_M 4.99 MB Download
Step-3.5-Flash-Q4_K_M-00002-of-00003.gguf GGUF Q4_K_M 60.50 GB Download
Step-3.5-Flash-Q4_K_M-00003-of-00003.gguf GGUF Q4_K_M 53.32 GB Download

Model Details Live

Model Slug
aessedai/step-3.5-flash-gguf
Author
AesSedai
Pipeline Task
Library
Created
2026-02-07
Last Modified
2026-03-01
Gated
No
Private
No
HF SHA
341d41e9b6687b3b3830e130eb1199e60096ebd6
License
Unknown
Language
Unknown
Base Model
stepfun-ai/Step-3.5-Flash

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": [
      "stepfun-ai/Step-3.5-Flash"
    ],
    "frontmatter": {
      "base_model": [
        "stepfun-ai/Step-3.5-Flash"
      ]
    },
    "hero_image_url": "kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\"",
    "summary": "This repo contains specialized MoE-quants for Step-3.5-Flash. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors. | Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD | | :--------- | :--------- | :------- | :------- | :------- | :------- | | Q4_K_M | 113.82 GiB (4.96 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 4.718049 ± 0.030373 | +0.3762% | 0.015464 ± 0.000133 | | IQ4_XS | 88.90 GiB (3.88 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 4.822499 ± 0.031236 | +2.5984% | 0.042753 ± 0.000301 | | IQ3_XXS | 73.10 GiB (3.19 BPW) | Q6_K / IQ3_XXS / IQ3_XXS / IQ3_XXS | 4.882908 ± 0.031560 | +3.8836% | 0.078681 ± 0.000506 | !kld_graph !ppl_graph",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model:\n- stepfun-ai/Step-3.5-Flash\n---\n\nThis repo contains specialized MoE-quants for Step-3.5-Flash. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.\n\n| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |\n| :--------- | :--------- | :------- | :------- | :------- | :------- |\n| Q4_K_M | 113.82 GiB (4.96 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 4.718049 ± 0.030373 | +0.3762% | 0.015464 ± 0.000133 |\n| IQ4_XS | 88.90 GiB (3.88 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 4.822499 ± 0.031236 | +2.5984% | 0.042753 ± 0.000301 |\n| IQ3_XXS | 73.10 GiB (3.19 BPW) | Q6_K / IQ3_XXS / IQ3_XXS / IQ3_XXS | 4.882908 ± 0.031560 | +3.8836% | 0.078681 ± 0.000506 |\n\n![kld_graph](kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\")\n![ppl_graph](kld_data/02_ppl_vs_filesize.png \"Chart showing Pareto PPL analysis of quants\")",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "base_model:stepfun-ai/Step-3.5-Flash",
    "base_model:quantized:stepfun-ai/Step-3.5-Flash",
    "endpoints_compatible",
    "region:us",
    "imatrix",
    "conversational"
  ],
  "likes": 14,
  "downloads": 1090,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-01T07:31:18.000Z",
  "created_at": "2026-02-07T03:04:27.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "6986abbb326d69bcf574b869",
  "id": "AesSedai/Step-3.5-Flash-GGUF",
  "modelId": "AesSedai/Step-3.5-Flash-GGUF",
  "sha": "341d41e9b6687b3b3830e130eb1199e60096ebd6",
  "createdAt": "2026-02-07T03:04:27.000Z",
  "lastModified": "2026-03-01T07:31:18.000Z",
  "author": "AesSedai",
  "downloads": 1090,
  "likes": 14,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 17
}