lgai-exaone/k-exaone-236b-a23b-gguf Q4_K_M GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
lgai-exaone/k-exaone-236b-a23b-gguf overview
Comprehensive model page for lgai-exaone/k-exaone-236b-a23b-gguf
Downloads
1,874
Likes
17
Pipeline
text-generation
Library
transformers
Visibility
Public
Access
Open
Repository Files & Downloads
6 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| K-EXAONE-236B-A23B-BF16.gguf | GGUF | BF16 | 441.71 GB | Download |
| K-EXAONE-236B-A23B-IQ4_XS.gguf | GGUF | IQ4_XS | 118.93 GB | Download |
| K-EXAONE-236B-A23B-Q4_K_M.gguf | GGUF | Q4_K_M | 133.62 GB | Download |
| K-EXAONE-236B-A23B-Q5_K_M.gguf | GGUF | Q5_K_M | 156.72 GB | Download |
| K-EXAONE-236B-A23B-Q6_K.gguf | GGUF | Q6_K | 181.26 GB | Download |
| K-EXAONE-236B-A23B-Q8_0.gguf | GGUF | โ | 234.73 GB | Download |
Benchmarks
| ๐ Free API until Feb 12th, 2026! Try on โฌ๏ธ FriendliAI โ๏ธ |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": "LGAI-EXAONE/K-EXAONE-236B-A23B",
"base_model_relation": "quantized",
"license": "other",
"license_name": "k-exaone",
"license_link": "LICENSE",
"language": [
"en",
"ko",
"es",
"de",
"ja",
"vi"
],
"tags": [
"lg-ai",
"exaone",
"k-exaone"
],
"pipeline_tag": "text-generation",
"library_name": "transformers",
"frontmatter": {
"base_model": "LGAI-EXAONE/K-EXAONE-236B-A23B",
"base_model_relation": "quantized",
"license": "other",
"license_name": "k-exaone",
"license_link": "LICENSE",
"language": [
"en",
"ko",
"es",
"de",
"ja",
"vi"
],
"tags": [
"lg-ai",
"exaone",
"k-exaone"
],
"pipeline_tag": "text-generation",
"library_name": "transformers"
},
"hero_image_url": "assets/K-EXAONE_logo_gray.png",
"summary": "",
"quick_links": [],
"benchmark_table_html": "<table><tr><td>๐ <span style=\"color: orange\"> <b>Free API until Feb 12th, 2026</b>! </span> Try on โฌ๏ธ FriendliAI โ๏ธ</td></tr></table>",
"readme_markdown": "---\nbase_model: LGAI-EXAONE/K-EXAONE-236B-A23B\nbase_model_relation: quantized\nlicense: other\nlicense_name: k-exaone\nlicense_link: LICENSE\nlanguage:\n - en\n - ko\n - es\n - de\n - ja\n - vi\ntags:\n - lg-ai\n - exaone\n - k-exaone\npipeline_tag: text-generation\nlibrary_name: transformers\n---\n\n<br>\n<br>\n<p align=\"center\">\n<img src=\"assets/K-EXAONE_logo_gray.png\" width=\"400\">\n<br>\n<br>\n<br>\n\n<div align=\"center\">\n <a href=\"https://huggingface.co/collections/LGAI-EXAONE/k-exaone\" style=\"text-decoration: none;\">\n <img src=\"https://img.shields.io/badge/๐ค-HuggingFace-FC926C?style=for-the-badge\" alt=\"HuggingFace\">\n </a>\n <a href=\"https://www.lgresearch.ai/blog/view?seq=619\" style=\"text-decoration: none;\">\n <img src=\"https://img.shields.io/badge/๐-Blog-E343BD?style=for-the-badge\" alt=\"Blog\">\n </a>\n <a href=\"https://arxiv.org/abs/2601.01739\" style=\"text-decoration: none;\">\n <img src=\"https://img.shields.io/badge/๐-Technical_Report-684CF4?style=for-the-badge\" alt=\"Technical Report\">\n </a>\n <a href=\"https://github.com/LG-AI-EXAONE/K-EXAONE\" style=\"text-decoration: none;\">\n <img src=\"https://img.shields.io/badge/๐ฅ๏ธ-GitHub-2B3137?style=for-the-badge\" alt=\"GitHub\">\n </a>\n <a href=\"https://friendli.ai/model/LGAI-EXAONE/K-EXAONE-236B-A23B\" style=\"text-decoration: none;\">\n <img src=\"https://img.shields.io/badge/โ๏ธ_API-Try_on_FriendliAI-2649BC?style=for-the-badge\" alt=\"FriendliAI\">\n </a>\n</div>\n\n\n\n<div align=\"center\">\n<table><tr><td>๐ <span style=\"color: orange\"> <b>Free API until Feb 12th, 2026</b>! </span> Try on โฌ๏ธ FriendliAI โ๏ธ</td></tr></table>\n</div>\n\n<br>\n\n# K-EXAONE-236B-A23B-GGUF\n\n## Introduction\n\nWe introduce **K-EXAONE**, a large-scale multilingual language model developed by LG AI Research. Built using a Mixture-of-Experts architecture, K-EXAONE features **236 billion total** parameters, with **23 billion active** during inference. Performance evaluations across various benchmarks demonstrate that K-EXAONE excels in reasoning, agentic capabilities, general knowledge, multilingual understanding, and long-context processing.\n\n#### Key Features\n\n- **Architecture & Efficiency:** Features a 236B fine-grained MoE design (23B active) optimized with **Multi-Token Prediction (MTP)**, enabling self-speculative decoding that boosts inference throughput by approximately 1.5x.\n- **Long-Context Capabilities:** Natively supports a **256K context window**, utilizing a **3:1 hybrid attention** scheme with a **128-token sliding window** to significantly minimize memory usage during long-document processing.\n- **Multilingual Support:** Covers 6 languages: Korean, English, Spanish, German, Japanese, and Vietnamese. Features a redesigned **150k vocabulary** with **SuperBPE**, improving token efficiency by ~30%.\n- **Agentic Capabilities:** Demonstrates superior tool-use and search capabilities via **multi-agent strategies.**\n- **Safety & Ethics:** Aligned with **universal human values**, the model uniquely incorporates **Korean cultural and historical contexts** to address regional sensitivities often overlooked by other models. It demonstrates high reliability across diverse risk categories.\n\nFor more details, please refer to the [technical report](https://arxiv.org/abs/2601.01739), [blog](https://www.lgresearch.ai/blog/view?seq=619) and [GitHub](https://github.com/LG-AI-EXAONE/K-EXAONE).\n\n\n\n\n### Model Configuration\n\n- Number of Parameters: 236B in total and 23B activated\n- Number of Parameters (without embeddings): 234B\n- Hidden Dimension: 6,144\n- Number of Layers: 48 Main layers + 1 MTP layers\n - Hybrid Attention Pattern: 12 x (3 Sliding window attention + 1 Global attention)\n- Sliding Window Attention\n - Number of Attention Heads: 64 Q-heads and 8 KV-heads\n - Head Dimension: 128 for both Q/KV\n - Sliding Window Size: 128\n- Global Attention\n - Number of Attention Heads: 64 Q-heads and 8 KV-heads\n - Head Dimension: 128 for both Q/KV\n - No Rotary Positional Embedding Used (NoPE)\n- Mixture of Experts:\n - Number of Experts: 128\n - Number of Activated Experts: 8\n - Number of Shared Experts: 1\n - MoE Intermediate Size: 2,048\n- Vocab Size: 153,600\n- Context Length: 262,144 tokens\n- Knowledge Cutoff: Dec 2024 (2024/12)\n- Quantization: `Q8_0`, `Q6_K`, `Q5_K_M`, `Q4_K_M`, `IQ4_XS` in GGUF format (also includes BF16 weights)\n\n## Evaluation Results\n\nThe evaluation results of the original model against other models are available on the [GitHub](https://github.com/LG-AI-EXAONE/K-EXAONE#performance) page or in the [model card](https://huggingface.co/LGAI-EXAONE/K-EXAONE-236B-A23B#evaluation-results) of the original model.\nDetailed evaluation configurations and results can be found in the [technical report](https://arxiv.org/abs/2601.01739).\n## Requirements\n\nK-EXAONE is supported by multiple libraries. Please install the required libraries as needed for your use case.\n\n#### Transformers\n\nYou should install `transformers >= 5.1.0` for the K-EXAONE model.\n\n#### llama.cpp\n\nTo use the K-EXAONE model with llama.cpp library, you should install `llama.cpp >= b7737`.\n\n\n### Quickstart\n\n### llama.cpp\n\nYou should install the `llama.cpp` library with the version of `b7737` or after.\n\nAfter you install the library, you need to prepare a model file in GGUF format as below:\n```bash\n# Download GGUF model weights (e.g. Q4_K_M)\nhf download LGAI-EXAONE/K-EXAONE-236B-A23B-GGUF --include \"*Q4_K_M*\" --local-dir .\n\n# Or convert huggingface model into GGUF format on your own\nhf download LGAI-EXAONE/K-EXAONE-236B-A23B --local-dir $YOUR_MODEL_DIR\npython convert_hf_to_gguf.py $YOUR_MODEL_DIR --outtype bf16 --outfile K-EXAONE-236B-A23B-BF16.gguf\n\n# If you want to use the lower precision than BF16, you need to quantize the model\n./llama-quantize K-EXAONE-236B-A23B-BF16.gguf K-EXAONE-236B-A23B-Q4_K_M.gguf Q4_K_M\n```\n\nYou can test the model with simple chat CLI by running the command below:\n```bash\n./llama-cli -m K-EXAONE-236B-A23B-Q4_K_M.gguf \\\n -ngl 99 \\\n -fa on -sm row \\\n --temp 1.0 --top-p 0.95 --min-p 0 \\\n -c 131072 -n 32768 \\\n --no-context-shift \\\n --jinja\n```\n\nYou can also launch a server by running the command below:\n\n```bash\n./llama-server -m K-EXAONE-236B-A23B-Q4_K_M.gguf \\\n -ngl 99 \\\n -fa on -sm row \\\n --temp 1.0 --top-p 0.95 --min-p 0 \\\n -c 131072 -n 32768 \\\n --no-context-shift \\\n --jinja \\\n --host 0.0.0.0 --port 8080\n```\n\nWhen the server is ready, you can test the model using the chat-style UI at http://localhost:8080, and access the OpenAI-compatible API at http://localhost:8080/v1.\n\n### Ollama / LM-Studio\n\nOllama and LM-Studio are powered by llama.cpp, so they should be updated once llama.cpp officially supports K-EXAONE. We will update this section once each library supports K-EXAONE.\n\n\n## Usage Guideline\n\n> [!IMPORTANT]\n> To achieve the expected performance, we recommend using the following configurations:\n> - We strongly recommend to use `temperature=1.0`, `top_p=0.95`, `presence_penalty=0.0` for best performance.\n> - Different from EXAONE-4.0, K-EXAONE uses `enable_thinking=True` as default. Thus, you need to set `enable_thinking=False` when you want to use non-reasoning mode.\n>\n\n## Limitation\n\nThe K-EXAONE language model has certain limitations and may occasionally generate inappropriate responses. The language model generates responses based on the output probability of tokens, and it is determined during learning from training data. While we have made every effort to exclude personal, harmful, and biased information from the training data, some problematic content may still be included, potentially leading to undesirable responses. Please note that the text generated by K-EXAONE language model does not reflect the views of LG AI Research.\n\n- Inappropriate answers may be generated, which contain personal, harmful or other inappropriate information.\n- Biased responses may be generated, which are associated with age, gender, race, and so on.\n- The generated responses rely heavily on statistics from the training data, which can result in the generation of semantically or syntactically incorrect sentences.\n- Since the model does not reflect the latest information, the responses may be false or contradictory.\n\nLG AI Research strives to reduce potential risks that may arise from K-EXAONE language models. Users are not allowed to engage in any malicious activities (e.g., keying in illegal information) that may induce the creation of inappropriate outputs violating LG AI's ethical principles when using K-EXAONE language models.\n\n\n## License\n\nThe model is licensed under [K-EXAONE AI Model License Agreement](./LICENSE)\n\n\n## Citation\n\n```\n@article{k-exaone,\n title={K-EXAONE Technical Report},\n author={{LG AI Research}},\n journal={arXiv preprint arXiv:2601.01739},\n year={2025}\n}\n```\n\n## Contact\n\nLG AI Research Technical Support: contact_us@lgresearch.ai\n\n",
"related_quantizations": []
},
"tags": [
"transformers",
"gguf",
"lg-ai",
"exaone",
"k-exaone",
"text-generation",
"conversational",
"en",
"ko",
"es",
"de",
"ja",
"vi",
"arxiv:2601.01739",
"base_model:LGAI-EXAONE/K-EXAONE-236B-A23B",
"base_model:quantized:LGAI-EXAONE/K-EXAONE-236B-A23B",
"license:other",
"endpoints_compatible",
"region:us"
],
"likes": 17,
"downloads": 1874,
"gated": false,
"private": false,
"last_modified": "2026-02-06T05:37:24.000Z",
"created_at": "2026-01-02T01:51:25.000Z",
"pipeline_tag": "text-generation",
"library_name": "transformers"
}
Source payload excerpt (from Hugging Face API)
{
"_id": "6957249d607c867c5c037571",
"id": "LGAI-EXAONE/K-EXAONE-236B-A23B-GGUF",
"modelId": "LGAI-EXAONE/K-EXAONE-236B-A23B-GGUF",
"sha": "73a9e3c2605b16ecfdd8bbff4f5f8628454a5262",
"createdAt": "2026-01-02T01:51:25.000Z",
"lastModified": "2026-02-06T05:37:24.000Z",
"author": "LGAI-EXAONE",
"downloads": 1874,
"likes": 17,
"gated": false,
"private": false,
"pipeline_tag": "text-generation",
"library_name": "transformers",
"siblings_count": 13
}