GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF overview

sasa2000/SmallThinker 4BA0.6B Instruct REAP 0.125 Q8 0 GGUF This model was converted to GGUF format from sasa2000/SmallThinker 4BA0.6B Instruct REAP 0.125 http…

transformersgguftext-generationmoepruningreapsafetensorscustom-codesmallthinkerllama-cppgguf-my-repobase_model:sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125base_model:quantized:sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125license:apache-2.0endpoints_compatibleregion:us

Runs locally from ~3.55 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
63
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
smallthinker-4ba0.6b-instruct-reap-0.125-q8_0.ggufGGUFQ8_03.55 GBDownload

Model Details

Model IDsasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF
Authorsasa2000
Pipelinetext-generation
Licenseapache-2.0
Base modelsasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125
Last modified2026-07-02T16:01:48.000Z

Model README

---

license: apache-2.0

base_model: sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125

library_name: transformers

pipeline_tag: text-generation

tags:

  • text-generation
  • moe
  • pruning
  • reap
  • safetensors
  • custom-code
  • smallthinker
  • llama-cpp
  • gguf-my-repo

---

sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF

This model was converted to GGUF format from sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125 using llama.cpp via the ggml.ai's GGUF-my-repo space.

Refer to the original model card for more details on the model.

Code:https://github.com/sasa200004/reap-smallthinker

Use with llama.cpp

Install llama.cpp through brew (works on Mac and Linux)

brew install llama.cpp

Invoke the llama.cpp server or the CLI.

CLI:

llama-cli --hf-repo sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF --hf-file smallthinker-4ba0.6b-instruct-reap-0.125-q8_0.gguf -p "The meaning to life and the universe is"

Server:

llama-server --hf-repo sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF --hf-file smallthinker-4ba0.6b-instruct-reap-0.125-q8_0.gguf -c 2048

Note: You can also use this checkpoint directly through the usage steps listed in the Llama.cpp repo as well.

Step 1: Clone llama.cpp from GitHub.

git clone https://github.com/ggerganov/llama.cpp

Step 2: Move into the llama.cpp folder and build it with LLAMA_CURL=1 flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux).

cd llama.cpp && LLAMA_CURL=1 make

Step 3: Run inference through the main binary.

./llama-cli --hf-repo sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF --hf-file smallthinker-4ba0.6b-instruct-reap-0.125-q8_0.gguf -p "The meaning to life and the universe is"

or

./llama-server --hf-repo sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF --hf-file smallthinker-4ba0.6b-instruct-reap-0.125-q8_0.gguf -c 2048

Run sasa2000/SmallThinker-4BA0.6B-Instruct-REAP-0.125-Q8_0-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models