GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Dsg2/LS-63M-A16M-GGUF overview

GGUF version: Q8 tested in the latest CPU llama.cpp build. Disable cuda, vulkan, rocm etc if not CPU version. I strongly reccomend you use the .bin version and…

gguftext-generationendataset:bigcode/the-stack-v2dataset:bigcode/starcoderdatadataset:Salesforce/wikitextbase_model:Dsg2/LS-63M-A16Mbase_model:quantized:Dsg2/LS-63M-A16Mlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~70.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
59
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LS-63M-A16M-Q8_0.ggufGGUFQ8_070.9 MBDownload

Model Details

Model IDDsg2/LS-63M-A16M-GGUF
AuthorDsg2
Pipelinetext-generation
Licenseapache-2.0
Base modelDsg2/LS-63M-A16M
Last modified2026-08-26T00:18:31.000Z

Model README

---

license: apache-2.0

datasets:

  • bigcode/the-stack-v2
  • bigcode/starcoderdata
  • Salesforce/wikitext

language:

  • en

pipeline_tag: text-generation

base_model:

  • Dsg2/LS-63M-A16M

---

GGUF version:

Q8 tested in the latest CPU llama.cpp build.

Disable cuda, vulkan, rocm etc if not CPU version.

I strongly reccomend you use the .bin version and the latest tinylm inference engine.

Note: the GGUF version is compatible with llama.cpp but llama.cpp currently does not fully support the chat template, and therefore will produce excessive hallucinations. For optimal usage, please use the .bin version with the latest tinylm.

---

LS-63M-A16M

| task | 63M / 16M act | 92M / 22M act | 220M / 25M act | chance |

|---|---|---|---|---|

| arc_easy | 31.2% | 35.0% | 35.8% | 25.1% |

| hellaswag | 27.3% | 28.5% | 31.8% | 25.0% |

| piqa | 56.5% | 60.0% | 60.8% | 50.0% |

| lambada | 13.2% | 18.5% | 19.2% | 0% |

| mmlu | 24.8% | 23.8% | 24.5% | 25.0% |

Miniature mixture of experts model with top-1 routing.

Trained entirely on a 1660 super.

This checkpoint marks the first epoch of training complete, ~1B tokens over 30 GPU hours.

Total parameters: 63M

Active parameters: 16M

context length: 8192

Training end evals:

| val loss | 1.4286 |

| --- | --- |

| perplexity | 4.17 |

Chat:

you> hi
bot> Hello! How can I assist you today?

[13 tok, 117.1 tok/s, ctx 23/16384]

you> what is the capital of france?
bot> Juan Van Gogh

[12 tok, 129.0 tok/s, ctx 52/16384]

Code:

you> write a python function that reverses a string
bot> Here is a simple Python function that reverses a string:

def reverse_string(s):

return s[::-1]


In this function, we use the `re.split()` function to split the string at the commas and create a list of words. Then we use `re.split()` to split the string on the `^`, and finally, we use `str.split()` to split the list of words.
[98 tok, 63.9 tok/s, ctx 116/16384]

To try it yourself:

Download tinylm.exe and LS-63M-A16M-q8.bin (placed in \models), run command tinylm chat LS-63M-A16M-q8 2048

Note: the bundled tinylm.exe is likely outdated. For the latest version, check here for the source code of tinylm and windows prebuilts.

Run Dsg2/LS-63M-A16M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models