WariHima/llm-jp-4-33b-thinking-Q4_K_M-GGUF overview
WariHima/llm jp 4 33b thinking Q4 K M GGUF This model was converted to GGUF format from WariHima/llm jp 4 33b thinking https://huggingface.co/WariHima/llm jp 4…
Runs locally from ~18.78 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| llm-jp-4-33b-thinking-q4_k_m.gguf | GGUF | Q4_K_M | 18.78 GB | Download |
Model Details
| Model ID | WariHima/llm-jp-4-33b-thinking-Q4_K_M-GGUF |
|---|---|
| Author | WariHima |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | llm-jp/llm-jp-4-33b-thinking |
| Last modified | 2026-08-18T14:57:40.000Z |
Model README
---
license: apache-2.0
language:
- en
- ja
programming_language:
- C
- C++
- C#
- Go
- Java
- JavaScript
- Lua
- PHP
- Python
- Ruby
- Rust
- Scala
- TypeScript
pipeline_tag: text-generation
library_name: transformers
inference: false
tags:
- llama-cpp
- gguf-my-repo
base_model: llm-jp/llm-jp-4-33b-thinking
---
WariHima/llm-jp-4-33b-thinking-Q4_K_M-GGUF
This model was converted to GGUF format from WariHima/llm-jp-4-33b-thinking using llama.cpp via the ggml.ai's GGUF-my-repo space.
Refer to the original model card for more details on the model.
attention!!
llm-jp4のllm-jp-4-33b-thinkingをduplicateしていろいろしたリポジトリ(WariHima/llm-jp-4-33b-thinking)
の物をgguf-my-repoで変換したモデルです。
chat templateをgpt-oss 11bから取ってくる(llm-jpの物を使う潜在的なバグ回避)
tokenizerをllmjp4-tokenizerを使わないようtokenizer_config.jsonを変更
llm-jp4のmodels/ver4.0_alpha1.0/llm-jp-tokenizer_ver4.0_alpha1.0.modeleをtokenizer.modelにリネームし追加。
理論上はmainstreamのllama.cppでも動きますが、
しかし、ためしていないしうまく動かない可能性もあるので、
その際は、hiratagoh/llm-jp-4-32b-a3b-thinking-GGUFを参考になおしてください(丸投げします)
またベンチマークもお任せします。
Run WariHima/llm-jp-4-33b-thinking-Q4_K_M-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models