GraySoft
Projects Models About FAQ Contact Download guIDE →

aessedai/mistral-small-4-119b-2603-gguf BF16 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

aessedai/mistral-small-4-119b-2603-gguf overview

Description This repo contains specialized MoE-quants for Mistral-Small-4-119B-2603. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

ggufbase_model:mistralai/Mistral-Small-4-119B-2603base_model:quantized:mistralai/Mistral-Small-4-119B-2603endpoints_compatibleregion:usimatrixconversational
aessedai/mistral-small-4-119b-2603-gguf visual
Downloads
916
Likes
7
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

13 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Mistral-Small-4-119B-2603-IQ3_S-00001-of-00002.gguf GGUF IQ3_S 7.50 MB Download
Mistral-Small-4-119B-2603-IQ3_S-00002-of-00002.gguf GGUF IQ3_S 40.89 GB Download
Mistral-Small-4-119B-2603-IQ4_XS-00001-of-00003.gguf GGUF IQ4_XS 7.50 MB Download
Mistral-Small-4-119B-2603-IQ4_XS-00002-of-00003.gguf GGUF IQ4_XS 46.44 GB Download
Mistral-Small-4-119B-2603-IQ4_XS-00003-of-00003.gguf GGUF IQ4_XS 6.65 GB Download
Mistral-Small-4-119B-2603-Q5_K_M-00001-of-00003.gguf GGUF Q5_K_M 7.50 MB Download
Mistral-Small-4-119B-2603-Q5_K_M-00002-of-00003.gguf GGUF Q5_K_M 46.09 GB Download
Mistral-Small-4-119B-2603-Q5_K_M-00003-of-00003.gguf GGUF Q5_K_M 35.97 GB Download
imatrix.gguf GGUF 113.30 MB Download
mmproj-Mistral-Small-4-119B-2603-BF16.gguf GGUF BF16 826.52 MB Download
mmproj-Mistral-Small-4-119B-2603-F16.gguf GGUF F16 817.37 MB Download
mmproj-Mistral-Small-4-119B-2603-F32.gguf GGUF F32 1.60 GB Download
mmproj-Mistral-Small-4-119B-2603-Q8_0.gguf GGUF 447.77 MB Download

Model Details Live

Model Slug
aessedai/mistral-small-4-119b-2603-gguf
Author
AesSedai
Pipeline Task
Library
Created
2026-03-17
Last Modified
2026-03-17
Gated
No
Private
No
HF SHA
d779b9f5c259afffbef5bc140c8a9e73c6d7cde4
License
Unknown
Language
Unknown
Base Model
mistralai/Mistral-Small-4-119B-2603

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": [
      "mistralai/Mistral-Small-4-119B-2603"
    ],
    "frontmatter": {
      "base_model": [
        "mistralai/Mistral-Small-4-119B-2603"
      ]
    },
    "hero_image_url": "kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\"",
    "summary": "## Description This repo contains specialized MoE-quants for Mistral-Small-4-119B-2603. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model:\n- mistralai/Mistral-Small-4-119B-2603\n---\n## Description\nThis repo contains specialized MoE-quants for Mistral-Small-4-119B-2603. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, \nit should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. \nTo that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.\n\n## Notes\nI had made a Q4_K_M mix, but it kept returning NaN's for the KLD / PPL testing so I'm looking more into that.\n\n| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |\n| :--------- | :--------- | :------- | :------- | :------- | :------- |\n| Q5_K_M | 82.06 GiB (5.92 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 5.773442 ± 0.037218 | +0.3924% | 0.054105 ± 0.000366 |\n| IQ4_XS | 53.09 GiB (3.83 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 5.980665 ± 0.038900 | +3.9957% | 0.132002 ± 0.000692 |\n| IQ3_S | 40.89 GiB (2.95 BPW) | Q8_0 / IQ2_S / IQ2_S / IQ3_S | 6.506233 ± 0.043705 | +13.1347% | 0.261355 ± 0.001217 |\n\n![kld_graph](kld_data/01_kld_vs_filesize.png \"Chart showing Pareto KLD analysis of quants\")\n![ppl_graph](kld_data/02_ppl_vs_filesize.png \"Chart showing Pareto PPL analysis of quants\")",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "base_model:mistralai/Mistral-Small-4-119B-2603",
    "base_model:quantized:mistralai/Mistral-Small-4-119B-2603",
    "endpoints_compatible",
    "region:us",
    "imatrix",
    "conversational"
  ],
  "likes": 7,
  "downloads": 916,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-17T05:40:05.000Z",
  "created_at": "2026-03-17T04:57:33.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69b8df3d5bff04a265cdbe9a",
  "id": "AesSedai/Mistral-Small-4-119B-2603-GGUF",
  "modelId": "AesSedai/Mistral-Small-4-119B-2603-GGUF",
  "sha": "d779b9f5c259afffbef5bc140c8a9e73c6d7cde4",
  "createdAt": "2026-03-17T04:57:33.000Z",
  "lastModified": "2026-03-17T05:40:05.000Z",
  "author": "AesSedai",
  "downloads": 916,
  "likes": 7,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 21
}