GraySoft
Projects Models About FAQ Contact Download guIDE →

aaryank/mistral-small-4-119b-2603-gguf 2603.q4_0 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

aaryank/mistral-small-4-119b-2603-gguf overview

Comprehensive model page for aaryank/mistral-small-4-119b-2603-gguf

gguftext-generation-inferencemistralmoereasoningagentmultimodaltext-generationarenfresdeitptnljakozhbase_model:mistralai/Mistral-Small-4-119B-2603base_model:quantized:mistralai/Mistral-Small-4-119B-2603license:apache-2.0endpoints_compatibleregion:usconversational
aaryank/mistral-small-4-119b-2603-gguf visual
Downloads
738
Likes
2
Pipeline
text-generation
Library
gguf
Visibility
Public
Access
Open

Repository Files & Downloads

13 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Mistral-Small-4-119B-2603-mmproj-f16.gguf GGUF F16 817.37 MB Download
Mistral-Small-4-119B-2603.q2_k.gguf GGUF Q2_K 40.42 GB Download
Mistral-Small-4-119B-2603.q3_k_l.gguf GGUF Q3_K_L 57.38 GB Download
Mistral-Small-4-119B-2603.q4_0.gguf GGUF 62.52 GB Download
Mistral-Small-4-119B-2603.q4_1.gguf GGUF 69.42 GB Download
Mistral-Small-4-119B-2603.q4_k_m.gguf GGUF Q4_K_M 67.20 GB Download
Mistral-Small-4-119B-2603.q4_k_s.gguf GGUF Q4_K_S 63.03 GB Download
Mistral-Small-4-119B-2603.q5_0.gguf GGUF 76.31 GB Download
Mistral-Small-4-119B-2603.q5_1.gguf GGUF 83.20 GB Download
Mistral-Small-4-119B-2603.q5_k_m.gguf GGUF Q5_K_M 78.72 GB Download
Mistral-Small-4-119B-2603.q5_k_s.gguf GGUF Q5_K_S 76.31 GB Download
Mistral-Small-4-119B-2603.q6_k.gguf GGUF Q6_K 90.96 GB Download
Mistral-Small-4-119B-2603.q8_0.gguf GGUF 117.79 GB Download

Model Details Live

Model Slug
aaryank/mistral-small-4-119b-2603-gguf
Author
AaryanK
Pipeline Task
text-generation
Library
gguf
Created
2026-03-16
Last Modified
2026-03-17
Gated
No
Private
No
HF SHA
2afc75f1cdd5beadcc29c0c99f8a376031ac9aac
License
apache-2.0
Language
ar, en, fr, es, de, it, pt, nl, ja, ko, zh
Base Model
mistralai/Mistral-Small-4-119B-2603

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": "mistralai/Mistral-Small-4-119B-2603",
    "base_model_relation": "quantized",
    "language": [
      "ar",
      "en",
      "fr",
      "es",
      "de",
      "it",
      "pt",
      "nl",
      "ja",
      "ko",
      "zh"
    ],
    "library_name": "gguf",
    "license": "apache-2.0",
    "pipeline_tag": "text-generation",
    "tags": [
      "text-generation-inference",
      "mistral",
      "moe",
      "reasoning",
      "agent",
      "multimodal",
      "gguf"
    ],
    "frontmatter": {
      "base_model": "mistralai/Mistral-Small-4-119B-2603",
      "base_model_relation": "quantized",
      "language": [
        "ar",
        "en",
        "fr",
        "es",
        "de",
        "it",
        "pt",
        "nl",
        "ja",
        "ko",
        "zh"
      ],
      "library_name": "gguf",
      "license": "apache-2.0",
      "pipeline_tag": "text-generation",
      "tags": [
        "text-generation-inference",
        "mistral",
        "moe",
        "reasoning",
        "agent",
        "multimodal",
        "gguf"
      ]
    },
    "hero_image_url": "https://cdn-uploads.huggingface.co/production/uploads/64e1a459ff3fd4fd8eedb456/lUkcuXB9nGQfEfxdHAsqF.png",
    "summary": "",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model: mistralai/Mistral-Small-4-119B-2603\nbase_model_relation: quantized\nlanguage:\n  - ar\n  - en\n  - fr\n  - es\n  - de\n  - it\n  - pt\n  - nl\n  - ja\n  - ko\n  - zh\nlibrary_name: gguf\nlicense: apache-2.0\npipeline_tag: text-generation\ntags:\n  - text-generation-inference\n  - mistral\n  - moe\n  - reasoning\n  - agent\n  - multimodal\n  - gguf\n---\n\n![image](https://cdn-uploads.huggingface.co/production/uploads/64e1a459ff3fd4fd8eedb456/lUkcuXB9nGQfEfxdHAsqF.png)\n\n# Mistral-Small-4-119B-2603-GGUF\n\n## Description\n\nThis repository contains **GGUF** format model files for[Mistral AI's Mistral-Small-4-119B-2603](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603).\n\n**Mistral Small 4** is a powerful 119B-parameter Mixture-of-Experts (MoE) model that activates only 6.5B parameters per token. It unifies general instruction, advanced reasoning, and multimodal (vision) capabilities into a single highly efficient architecture. It natively supports a **256K context window** and allows toggling reasoning effort dynamically.\n\n## Files & Quantization\n\nTo see the available files, please verify the **Files and versions** tab.\n\n## How to Run (llama.cpp)\n\n**Important:** This model features a new multimodal and reasoning architecture. Support for Mistral Small 4 is actively being merged into `llama.cpp`. Please ensure you are using the absolute latest version (or the relevant PR branch) for full compatibility.\n\n**Recommended Parameters:**\n*   **Temperature:** `0.1` (Recommended by Mistral for precise instruction following/reasoning) to `0.7` (for creative tasks).\n*   **Context:** `-c` (The model supports up to 256k, adjust based on your VRAM/RAM).\n\n### CLI Example\n\n```bash\n./llama-cli -m Mistral-Small-4-119B-2603.Q4_K_M.gguf \\\n  -c 8192 \\\n  --temp 0.1 \\\n  -p \"User: Write me a sentence where every word starts with the next letter in the alphabet - start with 'a' and end with 'z'.\\nAssistant:\" \\\n  -cnv\n```\n\n### Server Example\n\n```bash\n./llama-server -m Mistral-Small-4-119B-2603.Q4_K_M.gguf \\\n  --port 8080 \\\n  --host 0.0.0.0 \\\n  -c 16384 \\\n  -ngl 99\n```\n\n<!-- ORIGINAL MODEL CARD BELOW -->\n\n# Original Model Card: Mistral-Small-4-119B-2603\n\n# Mistral Small 4 119B A6B\n\nMistral Small 4 is a powerful hybrid model capable of acting as both a general instruction model and a reasoning model. It unifies the capabilities of three different model families—**Instruct**, **Reasoning** (previously called Magistral), and **Devstral**—into a single, unified model.\n\nWith its multimodal capabilities, efficient architecture, and flexible mode switching, it is a powerful general-purpose model for any task. In a latency-optimized setup, Mistral Small 4 achieves a **40% reduction in end-to-end completion time**, and in a throughput-optimized setup, it handles **3x more requests per second** compared to Mistral Small 3.\n\nTo further improve efficiency you can either take advantages of:\n- Speculative decoding thanks to our trained eagle head [`mistralai/Mistral-Small-4-119B-2603-eagle`](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603-eagle).\n- 4 bit float precision quantization thanks to our NVFP4 checkpoint [`mistralai/Mistral-Small-4-119B-2603-NVFP4`](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603-NVFP4).\n\n## Key Features\n\nMistral Small 4 includes the following architectural choices:\n\n- **MoE**: 128 experts, 4 active.\n- **119B parameters**, with **6.5B activated per token**.\n- **256k context length**.\n- **Multimodal input**: Accepts both text and image input, with text output.\n- **Instruct and Reasoning functionalities** with function calls (reasoning effort configurable per request).\n\nMistral Small 4 offers the following capabilities:\n\n- **Reasoning Mode**: Toggle between fast instant reply mode and reasoning mode, boosting performance with test-time compute when requested.\n- **Vision**: Analyzes images and provides insights based on visual content, in addition to text.\n- **Multilingual**: Supports dozens of languages, including English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, and Arabic.\n- **System Prompt**: Strong adherence and support for system prompts.\n- **Agentic**: Best-in-class agentic capabilities with native function calling and JSON output.\n- **Speed-Optimized**: Delivers best-in-class performance and speed.\n- **Apache 2.0 License**: Open-source license for both commercial and non-commercial use.\n- **Large Context Window**: Supports a 256k context window.\n\n## Use Cases\n\nMistral Small 4 is designed for general chat assistants, coding, agentic tasks, and reasoning tasks (with reasoning mode toggled). Its multimodal capabilities also enable document and image understanding for data extraction and analysis.\n\nIts capabilities are ideal for:\n- Developers interested in coding and agentic capabilities for SWE automation and codebase exploration.\n- Enterprises seeking general chat assistants, agents, and document understanding.\n- Researchers leveraging its math and research capabilities.\n\nMistral Small 4 is also well-suited for customization and fine-tuning for more specialized tasks.\n\n### Examples\n- General chat assistant\n- Document parsing and extraction\n- Coding agent\n- Research assistant\n- Customization & fine-tuning\n- And more...\n\n## Benchmarks\n\n### Comparison with internal models\n\nDepending on your tasks you can trigger reasoning thanks to the support of the **per-request** parameter `reasoning_effort`. Set it to:\n- `reasoning_effort=\"none\"`: Fast, lightweight responses for everyday tasks, equivalent to the same chat style of [`mistralai/Mistral-Small-3.2-24B-Instruct-2506`](https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506).\n- `reasoning_effort=\"high\"`: Deep, step-by-step reasoning for complex problems, with equivalent verbosity to previous Magistral models such as [`mistralai/Magistral-Small-2509`](https://huggingface.co/mistralai/Magistral-Small-2509).\n\n![Internal benchmark](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/resolve/main/images/image2.png)\n\n#### Comparing Reasoning Models\n\n![Internal benchmark - Reasoning](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/resolve/main/images/image3.png)\n\n\n### Comparison with other models\n\nMistral Small 4 with reasoning achieves competitive scores, matching or surpassing GPT-OSS 120B across all three benchmarks while generating significantly\nshorter outputs. On AA LCR, Mistral Small 4 scores **0.72** with just **1.6K characters**, whereas Qwen models require **3.5-4x more output** (5.8-6.1K)\nfor comparable performance. On LiveCodeBench, Mistral Small 4 outperforms GPT-OSS 120B while producing **20% less output**.\nThis efficiency reduces latency, inference costs, and improves user experience.\n\n![Comparison benchmark - LCR](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/resolve/main/images/lcr.png)\n![Comparison benchmark - LiveCodeBench](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/resolve/main/images/livecode.png)\n![Comparison benchmark - AIME25](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/resolve/main/images/aime.png)\n\n## Usage\n\nYou can find Mistral Small 4 support on multiple libraries for inference and fine-tuning. We here thank everyone contributors and maintainers that helped us making it happen.\n\n### Inference\n\nThe model can be deployed with:\n- [`vllm (recommended)`](https://github.com/vllm-project/vllm): See [here](#vllm-recommended).\n- [`llama.cpp`](https://github.com/ggml-org/llama.cpp): See [here](#transformers). (WIP ⏳ – follow updates [here](https://github.com/ggml-org/llama.cpp/pull/20649))\n- [`SGLang`](https://github.com/sgl-project/sglang): (WIP ⏳ – follow updates [here](https://github.com/sgl-project/sglang/pull/20708/))\n- [`transformers`](https://github.com/huggingface/transformers): See [here](#transformers)\n\nFor optimal performance, we recommend using the Mistral AI API if local serving is subpar.\n\n### Fine-Tuning\n\nFine-tune the model via:\n- [`Axolotl`](https://github.com/axolotl-ai-cloud/axolotl): See [here](https://github.com/axolotl-ai-cloud/axolotl/tree/main/examples/mistral4).\n\n## vLLM (Recommended)\n\nWe recommend using Mistral Small 4 with the [vLLM library](https://github.com/vllm-project/vllm) for production-ready inference.\n\n### Installation\n\n> [!Tip]\n> Use our custom Docker image with fixes for tool calling and reasoning parsing in vLLM, and the latest Transformers version. We are working with the vLLM team to merge these fixes soon.\n\n**Custom Docker**\nUse the following Docker image: [`mistralllm/vllm-ms4:latest`](https://hub.docker.com/repository/docker/mistralllm/vllm-ms4/latest/):\n```bash\ndocker pull mistralllm/vllm-ms4:latest\ndocker run -it mistralllm/vllm-ms4:latest\n```\n\n**Manual Install**\nAlternatively, install `vllm` from this PR: [Add Mistral Guidance](https://github.com/vllm-project/vllm/pull/37081).\n\n> **Note**: This PR is expected to be merged into `vllm` main in the next 1-2 weeks (as of 16.03.2026). Track updates [here](https://github.com/vllm-project/vllm/pull/37081).\n\n1. Clone vLLM:\n   ```bash\n   git clone --branch fix_mistral_parsing https://github.com/juliendenize/vllm.git\n   ```\n2. Install with pre-compiled kernels:\n   ```bash\n   VLLM_USE_PRECOMPILED=1 pip install --editable .\n   ```\n3. Install `transformers` from main:\n   ```bash\n   uv pip install git+https://github.com/huggingface/transformers.git\n   ```\n   Ensure [`mistral_common >= 1.10.0`](https://github.com/mistralai/mistral-common/releases/tag/v1.10.0) is installed:\n   ```bash\n   python -c \"import mistral_common; print(mistral_common.__version__)\"\n   ```\n\n### Serve the Model\n\nWe recommend a server/client setup:\n```bash\nvllm serve mistralai/Mistral-Small-4-119B-2603 --max-model-len 262144 --tensor-parallel-size 2 --attention-backend FLASH_ATTN_MLA \\\n  --tool-call-parser mistral --enable-auto-tool-choice --reasoning-parser mistral --max_num_batched_tokens 16384 --max_num_seqs 128 \\\n  --gpu_memory_utilization 0.8\n```\n\n### Ping the Server\n\n<details>\n  <summary>Instruction Following</summary>\n\nMistral Small 4 can follow your instructions to the letter.\n\n\n```python\nfrom datetime import datetime, timedelta\n\nfrom openai import OpenAI\nfrom huggingface_hub import hf_hub_download\n\n# Modify OpenAI's API key and API base to use vLLM's API server.\nopenai_api_key = \"EMPTY\"\nopenai_api_base = \"http://localhost:8000/v1\"\n\nTEMP = 0.1\n\nclient = OpenAI(\n    api_key=openai_api_key,\n    base_url=openai_api_base,\n)\n\nmodels = client.models.list()\nmodel = models.data[0].id\n\n\ndef load_system_prompt(repo_id: str, filename: str) -> str:\n    file_path = hf_hub_download(repo_id=repo_id, filename=filename)\n    with open(file_path, \"r\") as file:\n        system_prompt = file.read()\n    today = datetime.today().strftime(\"%Y-%m-%d\")\n    yesterday = (datetime.today() - timedelta(days=1)).strftime(\"%Y-%m-%d\")\n    model_name = repo_id.split(\"/\")[-1]\n    return system_prompt.format(name=model_name, today=today, yesterday=yesterday)\n\n\nSYSTEM_PROMPT = load_system_prompt(model, \"SYSTEM_PROMPT.txt\")\n\nmessages = [\n    {\"role\": \"system\", \"content\": SYSTEM_PROMPT},\n    {\n        \"role\": \"user\",\n        \"content\": \"Write me a sentence where every word starts with the next letter in the alphabet - start with 'a' and end with 'z'.\",\n    },\n]\n\nresponse = client.chat.completions.create(\n    model=model,\n    messages=messages,\n    temperature=TEMP,\n    reasoning_effort=\"none\",\n)\n\nassistant_message = response.choices[0].message.content\nprint(assistant_message)\n```\n\n</details>\n\n<details>\n  <summary>Tool Call</summary>\n\nLet's solve some equations thanks to our simple Python calculator tool.\n\n\n```python\nimport json\nfrom datetime import datetime, timedelta\n\nfrom openai import OpenAI\nfrom huggingface_hub import hf_hub_download\n\n# Modify OpenAI's API key and API base to use vLLM's API server.\nopenai_api_key = \"EMPTY\"\nopenai_api_base = \"http://localhost:8000/v1\"\n\nTEMP = 0.1\n\nclient = OpenAI(\n    api_key=openai_api_key,\n    base_url=openai_api_base,\n)\n\nmodels = client.models.list()\nmodel = models.data[0].id\n\n\ndef load_system_prompt(repo_id: str, filename: str) -> str:\n    file_path = hf_hub_download(repo_id=repo_id, filename=filename)\n    with open(file_path, \"r\") as file:\n        system_prompt = file.read()\n    today = datetime.today().strftime(\"%Y-%m-%d\")\n    yesterday = (datetime.today() - timedelta(days=1)).strftime(\"%Y-%m-%d\")\n    model_name = repo_id.split(\"/\")[-1]\n    return system_prompt.format(name=model_name, today=today, yesterday=yesterday)\n\n\nSYSTEM_PROMPT = load_system_prompt(model, \"SYSTEM_PROMPT.txt\")\n\nimage_url = \"https://math-coaching.com/img/fiche/46/expressions-mathematiques.jpg\"\n\n\ndef my_calculator(expression: str) -> str:\n    return str(eval(expression))\n\n\ntools = [\n    {\n        \"type\": \"function\",\n        \"function\": {\n            \"name\": \"my_calculator\",\n            \"description\": \"A calculator that can evaluate a mathematical expression.\",\n            \"parameters\": {\n                \"type\": \"object\",\n                \"properties\": {\n                    \"expression\": {\n                        \"type\": \"string\",\n                        \"description\": \"The mathematical expression to evaluate.\",\n                    },\n                },\n                \"required\": [\"expression\"],\n            },\n        },\n    },\n    {\n        \"type\": \"function\",\n        \"function\": {\n            \"name\": \"rewrite\",\n            \"description\": \"Rewrite a given text for improved clarity\",\n            \"parameters\": {\n                \"type\": \"object\",\n                \"properties\": {\n                    \"text\": {\n                        \"type\": \"string\",\n                        \"description\": \"The input text to rewrite\",\n                    }\n                },\n            },\n        },\n    },\n]\n\nmessages = [\n    {\"role\": \"system\", \"content\": SYSTEM_PROMPT},\n    {\n        \"role\": \"user\",\n        \"content\": [\n            {\n                \"type\": \"text\",\n                \"text\": \"Thanks to your calculator, compute the results for the equations that involve numbers displayed in the image.\",\n            },\n            {\n                \"type\": \"image_url\",\n                \"image_url\": {\n                    \"url\": image_url,\n                },\n            },\n        ],\n    },\n]\n\nresponse = client.chat.completions.create(\n    model=model,\n    messages=messages,\n    temperature=TEMP,\n    tools=tools,\n    tool_choice=\"auto\",\n    reasoning_effort=\"none\",\n)\n\ntool_calls = response.choices[0].message.tool_calls\n\nresults = []\nfor tool_call in tool_calls:\n    function_name = tool_call.function.name\n    function_args = tool_call.function.arguments\n    if function_name == \"my_calculator\":\n        result = my_calculator(**json.loads(function_args))\n        results.append(result)\n\nmessages.append({\"role\": \"assistant\", \"tool_calls\": tool_calls})\nfor tool_call, result in zip(tool_calls, results):\n    messages.append(\n        {\n            \"role\": \"tool\",\n            \"tool_call_id\": tool_call.id,\n            \"name\": tool_call.function.name,\n            \"content\": result,\n        }\n    )\n\n\nresponse = client.chat.completions.create(\n    model=model,\n    messages=messages,\n    temperature=TEMP,\n    reasoning_effort=\"none\",\n)\n\nprint(response.choices[0].message.content)\n```\n\n</details>\n\n<details>\n  <summary>Vision Reasoning</summary>\n\nLet's see if the Mistral Small 4 knows when to pick a fight !\n\n```python\nfrom datetime import datetime, timedelta\n\nfrom openai import OpenAI\nfrom huggingface_hub import hf_hub_download\n\n# Modify OpenAI's API key and API base to use vLLM's API server.\nopenai_api_key = \"EMPTY\"\nopenai_api_base = \"http://localhost:8000/v1\"\n\nTEMP = 0.1\n\nclient = OpenAI(\n    api_key=openai_api_key,\n    base_url=openai_api_base,\n)\n\nmodels = client.models.list()\nmodel = models.data[0].id\n\n\ndef load_system_prompt(repo_id: str, filename: str) -> str:\n    file_path = hf_hub_download(repo_id=repo_id, filename=filename)\n    with open(file_path, \"r\") as file:\n        system_prompt = file.read()\n    today = datetime.today().strftime(\"%Y-%m-%d\")\n    yesterday = (datetime.today() - timedelta(days=1)).strftime(\"%Y-%m-%d\")\n    model_name = repo_id.split(\"/\")[-1]\n    return system_prompt.format(name=model_name, today=today, yesterday=yesterday)\n\n\nSYSTEM_PROMPT = load_system_prompt(model, \"SYSTEM_PROMPT.txt\")\nimage_url = \"https://static.wikia.nocookie.net/essentialsdocs/images/7/70/Battle.png/revision/latest?cb=20220523172438\"\n\nmessages = [\n    {\"role\": \"system\", \"content\": SYSTEM_PROMPT},\n    {\n        \"role\": \"user\",\n        \"content\": [\n            {\n                \"type\": \"text\",\n                \"text\": \"What action do you think I should take in this situation? List all the possible actions and explain why you think they are good or bad.\",\n            },\n            {\"type\": \"image_url\", \"image_url\": {\"url\": image_url}},\n        ],\n    },\n]\n\n\nresponse = client.chat.completions.create(\n    model=model,\n    messages=messages,\n    temperature=TEMP,\n    reasoning_effort=\"high\",\n)\n\nprint(response.choices[0].message.content)\n```\n\n</details>\n\n## Transformers\n\n### Installation\n\nYou need to install the main branch of Transformers to use Mistral Small 4:\n\n```bash\nuv pip install git+https://github.com/huggingface/transformers.git\n```\n\n### Inference\n\n> **Note**: Current implementation of Transformers does not support FP8.\nWeights have been stored in FP8 and updates to load them in this format are expected, in the meantime we provide BF16 quantization snippets to ease usage.\nAs soon as support is added, we will update the following code snippet.\n\n<details>\n<summary>Python Inference Snippet</summary>\n\n```python\nfrom pathlib import Path\n\nimport torch\nfrom huggingface_hub import snapshot_download\nfrom safetensors.torch import load_file\nfrom tqdm import tqdm\n\nfrom transformers import AutoConfig, AutoProcessor, Mistral3ForConditionalGeneration\n\n\ndef _descale_fp8_to_bf16(tensor: torch.Tensor, scale_inv: torch.Tensor) -> torch.Tensor:\n    return (tensor.to(torch.bfloat16) * scale_inv.to(torch.bfloat16)).to(torch.bfloat16)\n\n\ndef _resolve_model_dir(model_id: str) -> Path:\n    local = Path(model_id)\n    if local.is_dir():\n        return local\n    return Path(snapshot_download(model_id, allow_patterns=[\"model*.safetensors\"]))\n\n\ndef load_and_dequantize_state_dict(model_id: str) -> dict[str, torch.Tensor]:\n    model_dir = _resolve_model_dir(model_id)\n\n    shards = sorted(model_dir.glob(\"model*.safetensors\"))\n\n    full_state_dict: dict[str, torch.Tensor] = {}\n    for shard in tqdm(shards, desc=\"Loading safetensors shards\"):\n        full_state_dict.update(load_file(str(shard)))\n\n    scale_suffixes = (\"weight_scale_inv\", \"gate_up_proj_scale_inv\", \"down_proj_scale_inv\", \"up_proj_scale_inv\")\n    activation_scale_suffixes = (\"activation_scale\", \"gate_up_proj_activation_scale\", \"down_proj_activation_scale\")\n\n    keys_to_remove: set[str] = set()\n    all_keys = list(full_state_dict.keys())\n\n    for key in tqdm(all_keys, desc=\"Dequantizing FP8 weights to BF16\"):\n        if any(key.endswith(s) for s in scale_suffixes + activation_scale_suffixes):\n            continue\n\n        for scale_suffix in scale_suffixes:\n            if scale_suffix == \"weight_scale_inv\":\n                if not key.endswith(\".weight\"):\n                    continue\n                scale_key = key.rsplit(\".weight\", 1)[0] + \".weight_scale_inv\"\n            else:\n                proj_name = scale_suffix.replace(\"_scale_inv\", \"\")\n                if not key.endswith(f\".{proj_name}\"):\n                    continue\n                scale_key = key + \"_scale_inv\"\n\n            if scale_key in full_state_dict:\n                full_state_dict[key] = _descale_fp8_to_bf16(full_state_dict[key], full_state_dict[scale_key])\n                keys_to_remove.add(scale_key)\n\n    for key in full_state_dict:\n        if any(key.endswith(s) for s in activation_scale_suffixes):\n            keys_to_remove.add(key)\n\n    for key in tqdm(keys_to_remove, desc=\"Removing scale keys\"):\n        del full_state_dict[key]\n\n    return full_state_dict\n\n\ndef load_config_without_quantization(model_id: str) -> AutoConfig:\n    config = AutoConfig.from_pretrained(model_id)\n\n    if hasattr(config, \"quantization_config\"):\n        del config.quantization_config\n\n    if hasattr(config, \"text_config\") and hasattr(config.text_config, \"quantization_config\"):\n        del config.text_config.quantization_config\n\n    return config\n\n\nmodel_id = \"mistralai/Mistral-Small-4-119B-2603\"\n\nconfig = load_config_without_quantization(model_id)\nstate_dict = load_and_dequantize_state_dict(model_id)\n\nmodel = Mistral3ForConditionalGeneration.from_pretrained(\n    None,\n    config=config,\n    state_dict=state_dict,\n    device_map=\"auto\",\n)\n\nprocessor = AutoProcessor.from_pretrained(model_id)\n\nimage_url = \"https://static.wikia.nocookie.net/essentialsdocs/images/7/70/Battle.png/revision/latest?cb=20220523172438\"\n\nmessages = [\n    {\n        \"role\": \"user\",\n        \"content\": [\n            {\n                \"type\": \"text\",\n                \"text\": \"What action do you think I should take in this situation? List all the possible actions and explain why you think they are good or bad.\",\n            },\n            {\"type\": \"image_url\", \"image_url\": {\"url\": image_url}},\n        ],\n    },\n]\n\ninputs = processor.apply_chat_template(\n    messages, return_tensors=\"pt\", tokenize=True, return_dict=True, reasoning_effort=\"high\"\n)\ninputs = inputs.to(model.device)\n\noutput = model.generate(\n    **inputs,\n    max_new_tokens=1024,\n)[0]\n\n# Setting `skip_special_tokens=False` to visualize reasoning trace between [THINK] [/THINK] tags.\ndecoded_output = processor.decode(output[len(inputs[\"input_ids\"][0]) :], skip_special_tokens=False)\nprint(decoded_output)\n```\n</details>\n\n## License\n\nThis model is licensed under the [Apache 2.0 License](https://www.apache.org/licenses/LICENSE-2.0.txt).\n\n*You must not use this model in a manner that infringes, misappropriates, or violates any third party’s rights, including intellectual property rights.*\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "text-generation-inference",
    "mistral",
    "moe",
    "reasoning",
    "agent",
    "multimodal",
    "text-generation",
    "ar",
    "en",
    "fr",
    "es",
    "de",
    "it",
    "pt",
    "nl",
    "ja",
    "ko",
    "zh",
    "base_model:mistralai/Mistral-Small-4-119B-2603",
    "base_model:quantized:mistralai/Mistral-Small-4-119B-2603",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 2,
  "downloads": 738,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-17T10:34:42.000Z",
  "created_at": "2026-03-16T20:56:05.000Z",
  "pipeline_tag": "text-generation",
  "library_name": "gguf"
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69b86e659361d25cc549da09",
  "id": "AaryanK/Mistral-Small-4-119B-2603-GGUF",
  "modelId": "AaryanK/Mistral-Small-4-119B-2603-GGUF",
  "sha": "2afc75f1cdd5beadcc29c0c99f8a376031ac9aac",
  "createdAt": "2026-03-16T20:56:05.000Z",
  "lastModified": "2026-03-17T10:34:42.000Z",
  "author": "AaryanK",
  "downloads": 738,
  "likes": 2,
  "gated": false,
  "private": false,
  "pipeline_tag": "text-generation",
  "library_name": "gguf",
  "siblings_count": 15
}