GraySoft
Projects Models About FAQ Contact Download guIDE →

aessedai/step-3.5-flash-base-midtrain-gguf imatrix GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

aessedai/step-3.5-flash-base-midtrain-gguf overview

Description This repo contains specialized MoE-quants for Step-3.5-Flash-Base-Midtrain. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors. | Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD | | :--------- | :--------- | :------- | :------- | :------- | :------- | | Q5KM | 136.43 GiB (5.95 BPW) | Q80 / Q5K / Q5K / Q6K | 2.207801 ± 0.009234 | +0.6244% | 0.016217 ± 0.000097 | | Q4KM | 113.82 GiB (4.96 BPW) | Q80 / Q4K / Q4K / Q5K | 2.251718 ± 0.009525 | +2.6260% | 0.043240 ± 0.000250 | | IQ4XS | 88.90 GiB (3.88 BPW) | Q80 / IQ3S / IQ3S / IQ4XS | 2.440324 ± 0.010689 | +11.2221% | 0.136298 ± 0.000718 | | IQ3S | 68.48 GiB (2.99 BPW) | Q80 / IQ2S / IQ2S / IQ3S | 3.060379 ± 0.014918 | +39.4822% | 0.386923 ± 0.001795 | !kldgraph !pplgraph

ggufbase_model:stepfun-ai/Step-3.5-Flash-Base-Midtrainbase_model:quantized:stepfun-ai/Step-3.5-Flash-Base-Midtrainendpoints_compatibleregion:usimatrixconversational
aessedai/step-3.5-flash-base-midtrain-gguf visual
Downloads
551
Likes
4
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

15 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Step-3.5-Flash-Base-Midtrain-IQ3_S-00001-of-00003.gguf GGUF IQ3_S 4.99 MB Download
Step-3.5-Flash-Base-Midtrain-IQ3_S-00002-of-00003.gguf GGUF IQ3_S 46.20 GB Download
Step-3.5-Flash-Base-Midtrain-IQ3_S-00003-of-00003.gguf GGUF IQ3_S 22.29 GB Download
Step-3.5-Flash-Base-Midtrain-IQ4_XS-00001-of-00003.gguf GGUF IQ4_XS 4.99 MB Download
Step-3.5-Flash-Base-Midtrain-IQ4_XS-00002-of-00003.gguf GGUF IQ4_XS 46.17 GB Download
Step-3.5-Flash-Base-Midtrain-IQ4_XS-00003-of-00003.gguf GGUF IQ4_XS 42.73 GB Download
Step-3.5-Flash-Base-Midtrain-Q4_K_M-00001-of-00004.gguf GGUF Q4_K_M 4.99 MB Download
Step-3.5-Flash-Base-Midtrain-Q4_K_M-00002-of-00004.gguf GGUF Q4_K_M 46.33 GB Download
Step-3.5-Flash-Base-Midtrain-Q4_K_M-00003-of-00004.gguf GGUF Q4_K_M 46.25 GB Download
Step-3.5-Flash-Base-Midtrain-Q4_K_M-00004-of-00004.gguf GGUF Q4_K_M 21.24 GB Download
Step-3.5-Flash-Base-Midtrain-Q5_K_M-00001-of-00004.gguf GGUF Q5_K_M 4.99 MB Download
Step-3.5-Flash-Base-Midtrain-Q5_K_M-00002-of-00004.gguf GGUF Q5_K_M 45.66 GB Download
Step-3.5-Flash-Base-Midtrain-Q5_K_M-00003-of-00004.gguf GGUF Q5_K_M 46.00 GB Download
Step-3.5-Flash-Base-Midtrain-Q5_K_M-00004-of-00004.gguf GGUF Q5_K_M 44.77 GB Download
imatrix.gguf GGUF 444.41 MB Download

Model Details Live

Model Slug
aessedai/step-3.5-flash-base-midtrain-gguf
Author
AesSedai
Pipeline Task
Library
Created
2026-03-25
Last Modified
2026-03-25
Gated
No
Private
No
HF SHA
0c5c205a2c3d6890e50e81505ec1247e0eb8aebd
License
Unknown
Language
Unknown
Base Model
stepfun-ai/Step-3.5-Flash-Base-Midtrain

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": [
      "stepfun-ai/Step-3.5-Flash-Base-Midtrain"
    ],
    "frontmatter": {
      "base_model": [
        "stepfun-ai/Step-3.5-Flash-Base-Midtrain"
      ]
    },
    "hero_image_url": "kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\"",
    "summary": "## Description This repo contains specialized MoE-quants for Step-3.5-Flash-Base-Midtrain. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors. | Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD | | :--------- | :--------- | :------- | :------- | :------- | :------- | | Q5_K_M | 136.43 GiB (5.95 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 2.207801 ± 0.009234 | +0.6244% | 0.016217 ± 0.000097 | | Q4_K_M | 113.82 GiB (4.96 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 2.251718 ± 0.009525 | +2.6260% | 0.043240 ± 0.000250 | | IQ4_XS | 88.90 GiB (3.88 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 2.440324 ± 0.010689 | +11.2221% | 0.136298 ± 0.000718 | | IQ3_S | 68.48 GiB (2.99 BPW) | Q8_0 / IQ2_S / IQ2_S / IQ3_S | 3.060379 ± 0.014918 | +39.4822% | 0.386923 ± 0.001795 | !kld_graph !ppl_graph",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model:\n- stepfun-ai/Step-3.5-Flash-Base-Midtrain\n---\n## Description\nThis repo contains specialized MoE-quants for Step-3.5-Flash-Base-Midtrain. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, \nit should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. \nTo that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.\n\n| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |\n| :--------- | :--------- | :------- | :------- | :------- | :------- |\n| Q5_K_M | 136.43 GiB (5.95 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 2.207801 ± 0.009234 | +0.6244% | 0.016217 ± 0.000097 |\n| Q4_K_M | 113.82 GiB (4.96 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 2.251718 ± 0.009525 | +2.6260% | 0.043240 ± 0.000250 |\n| IQ4_XS | 88.90 GiB (3.88 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 2.440324 ± 0.010689 | +11.2221% | 0.136298 ± 0.000718 |\n| IQ3_S | 68.48 GiB (2.99 BPW) | Q8_0 / IQ2_S / IQ2_S / IQ3_S | 3.060379 ± 0.014918 | +39.4822% | 0.386923 ± 0.001795 |\n\n![kld_graph](kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\")\n![ppl_graph](kld_data/02_ppl_vs_filesize.png \"Chart showing Pareto PPL analysis of quants\")",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "base_model:stepfun-ai/Step-3.5-Flash-Base-Midtrain",
    "base_model:quantized:stepfun-ai/Step-3.5-Flash-Base-Midtrain",
    "endpoints_compatible",
    "region:us",
    "imatrix",
    "conversational"
  ],
  "likes": 4,
  "downloads": 551,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-25T16:23:17.000Z",
  "created_at": "2026-03-25T07:24:11.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69c38d9b0e6e4e2ef7c8f456",
  "id": "AesSedai/Step-3.5-Flash-Base-Midtrain-GGUF",
  "modelId": "AesSedai/Step-3.5-Flash-Base-Midtrain-GGUF",
  "sha": "0c5c205a2c3d6890e50e81505ec1247e0eb8aebd",
  "createdAt": "2026-03-25T07:24:11.000Z",
  "lastModified": "2026-03-25T16:23:17.000Z",
  "author": "AesSedai",
  "downloads": 551,
  "likes": 4,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 24
}