iapp/openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b-GGUF overview
<div align="center" <img src="https://huggingface.co/spaces/openthaigpt/README/resolve/main/openthai logo white.png" width="160" alt="OpenThai" OpenThai 2.0 Le…
Runs locally from ~22.83 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | iapp/openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b-GGUF |
|---|---|
| Author | iapp |
| Pipeline | text-generation |
| License | other |
| Base model | iapp/openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b |
| Last modified | 2026-07-26T20:11:33.000Z |
Model README
---
license: other
license_name: nvidia-open-model-agreement
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
language:
- th
- en
base_model: iapp/openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b
pipeline_tag: text-generation
library_name: gguf
tags:
- thai
- legal
- openthai
- gguf
- llama.cpp
- mixture-of-experts
---
<div align="center">
<img src="https://huggingface.co/spaces/openthaigpt/README/resolve/main/openthai-logo-white.png" width="160" alt="OpenThai">
OpenThai 2.0 Legal 30B-A3B — GGUF
Official GGUF quantizations of iapp/openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b
Website · Announcement · Live demo · Discord
</div>
An open-weight Thai legal LLM that recalls Thai statutes and cites the exact law name and
section (มาตรา) as structured JSON. 30B Mixture-of-Experts with only ~3B parameters active
per token — which is why a 30B model runs comfortably on modest hardware.
Use it with retrieval. Open-book citation accuracy is 0.99 versus 0.07–0.40 from pure
memory. Pair it with OpenThaiRAG or your own
retrieval over authoritative statute text.
Quants
| File | Quant | Size | Notes |
|---|---|---|---|
| openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b.Q4_K_M.gguf | Q4_K_M | ~18 GB | Recommended. Fits a 24 GB GPU. |
| openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b.Q5_K_M.gguf | Q5_K_M | ~21 GB | Higher quality. |
| openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b.Q8_0.gguf | Q8_0 | ~32 GB | Near-lossless. |
Usage
Ollama
ollama run hf.co/iapp/openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b-GGUF:Q4_K_M
llama.cpp
llama-cli -m openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b.Q4_K_M.gguf \
-p "ลักทรัพย์ในเวลากลางคืน ผิดมาตราใด" -n 1024 --temp 0.3
The chat template is embedded in the GGUF. Recommended sampling: temperature=0.3,
top_p=0.9. Architecture is a hybrid Mamba2-Transformer MoE (NVIDIA
Nemotron-3-Nano-30B-A3B base) — use a recent llama.cpp build.
⚠️ Responsible use
Outputs are decision support, not legal advice. Verify every citation against the
current statute text. Near-miss rejection — telling the governing section from a closely
related one — is the hardest task for every model tested, this one included.
Citation
@misc{openthai2026legal,
title = {OpenThai 2.0 Legal: An Open-Weight Thai Legal Language Model},
author = {Viriyayudhakorn, Kobkrit and Yuenyong, Sumeth and Chay-intr, Thodsaporn},
year = {2026},
url = {https://openthai.aieat.or.th/openthai2p0-legal}
}
---
*OpenThai (formerly OpenThaiGPT) — free, open-weight Thai large language models from AIEAT
and iApp Technology, built here on the NVIDIA Nemotron and NeMo stack. With thanks to the
community members who published unofficial GGUF conversions before these existed.*
Run iapp/openthai2.0-legal-thaillm-nemotron-3-nano-30b-a3b-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models