GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

axana-labs/SmolLM2-135M-Instruct-GGUF overview

Finetune SmolLM2, Llama 3.2, Gemma 2, Mistral 2 5x faster with 70% less memory via Unsloth We have a free Google Colab Tesla T4 notebook for Llama 3.2 3B here:…

transformersggufllamaunslothenbase_model:HuggingFaceTB/SmolLM2-135M-Instructbase_model:quantized:HuggingFaceTB/SmolLM2-135M-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~84.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

7 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
SmolLM2-135M-Instruct-F16.ggufGGUFF16258.3 MBDownload
SmolLM2-135M-Instruct-Q2_K.ggufGGUFQ2_K84.1 MBDownload
SmolLM2-135M-Instruct-Q3_K_M.ggufGGUFQ3_K_M89.2 MBDownload
SmolLM2-135M-Instruct-Q4_K_M.ggufGGUFQ4_K_M100.6 MBDownload
SmolLM2-135M-Instruct-Q5_K_M.ggufGGUFQ5_K_M106.9 MBDownload
SmolLM2-135M-Instruct-Q6_K.ggufGGUFQ6_K132.0 MBDownload
SmolLM2-135M-Instruct-Q8_0.ggufGGUFQ8_0138.1 MBDownload

Model Details

Model IDaxana-labs/SmolLM2-135M-Instruct-GGUF
Authoraxana-labs
Pipeline
Licenseapache-2.0
Base modelHuggingFaceTB/SmolLM2-135M-Instruct
Last modified2026-07-01T07:56:29.000Z

Model README

---

base_model: HuggingFaceTB/SmolLM2-135M-Instruct

language:

  • en

library_name: transformers

license: apache-2.0

tags:

  • llama
  • unsloth
  • transformers

---

Finetune SmolLM2, Llama 3.2, Gemma 2, Mistral 2-5x faster with 70% less memory via Unsloth!

We have a free Google Colab Tesla T4 notebook for Llama 3.2 (3B) here: https://colab.research.google.com/drive/1Ys44kVvmeZtnICzWz0xgpRnrIOjZAuxp?usp=sharing

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/Discord%20button.png" width="200"/>

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>

unsloth/SmolLM2-135M-Instruct-GGUF

For more details on the model, please go to Hugging Face's original model card

✨ Finetune for Free

All notebooks are beginner friendly! Add your dataset, click "Run All", and you'll get a 2x faster finetuned model which can be exported to GGUF, vLLM or uploaded to Hugging Face.

| Unsloth supports | Free Notebooks | Performance | Memory use |

|-----------------|--------------------------------------------------------------------------------------------------------------------------|-------------|----------|

| Llama-3.2 (3B) | ▶️ Start on Colab | 2.4x faster | 58% less |

| Llama-3.2 (11B vision) | ▶️ Start on Colab | 2.4x faster | 58% less |

| Llama-3.1 (8B) | ▶️ Start on Colab | 2.4x faster | 58% less |

| Phi-3.5 (mini) | ▶️ Start on Colab | 2x faster | 50% less |

| Gemma 2 (9B) | ▶️ Start on Colab | 2.4x faster | 58% less |

| Mistral (7B) | ▶️ Start on Colab | 2.2x faster | 62% less |

| DPO - Zephyr | ▶️ Start on Colab | 1.9x faster | 19% less |

Special Thanks

A huge thank you to the Hugging Face team for creating and releasing these models.

Model Summary

SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters. They are capable of solving a wide range of tasks while being lightweight enough to run on-device.

The 1.7B variant demonstrates significant advances over its predecessor SmolLM1-1.7B, particularly in instruction following, knowledge, reasoning, and mathematics. It was trained on 11 trillion tokens using a diverse dataset combination: FineWeb-Edu, DCLM, The Stack, along with new mathematics and coding datasets that we curated and will release soon. We developed the instruct version through supervised fine-tuning (SFT) using a combination of public datasets and our own curated datasets. We then applied Direct Preference Optimization (DPO) using UltraFeedback.

The instruct model additionally supports tasks such as text rewriting, summarization and function calling thanks to datasets developed by Argilla such as Synth-APIGen-v0.1.

SmolLM2

!image/png

Run axana-labs/SmolLM2-135M-Instruct-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models