llmfan46/gpt-oss-120b-heretic-v2-gguf MXFP4 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
llmfan46/gpt-oss-120b-heretic-v2-gguf overview
GGUF quantizations of llmfan46/gpt-oss-120b-heretic-v2. # This is a decensored version of openai/gpt-oss-120b, made using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method
Downloads
1,654
Likes
9
Pipeline
text-generation
Library
—
Visibility
Public
Access
Open
Repository Files & Downloads
1 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gpt-oss-120b-heretic-v2-MXFP4.gguf | GGUF | — | 60.88 GB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"license": "apache-2.0",
"pipeline_tag": "text-generation",
"tags": [
"heretic",
"uncensored",
"decensored",
"abliterated"
],
"base_model": [
"llmfan46/gpt-oss-120b-heretic-v2"
],
"frontmatter": {
"license": "apache-2.0",
"pipeline_tag": "text-generation",
"tags": [
"heretic",
"uncensored",
"decensored",
"abliterated"
],
"base_model": [
"llmfan46/gpt-oss-120b-heretic-v2"
]
},
"hero_image_url": "https://raw.githubusercontent.com/openai/gpt-oss/main/docs/gpt-oss-120b.svg",
"summary": "GGUF quantizations of llmfan46/gpt-oss-120b-heretic-v2. # This is a decensored version of openai/gpt-oss-120b, made using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nlicense: apache-2.0\npipeline_tag: text-generation\ntags:\n- heretic\n- uncensored\n- decensored\n- abliterated\nbase_model:\n- llmfan46/gpt-oss-120b-heretic-v2\n---\n<div style=\"background-color: #ff4444; color: white; padding: 20px; border-radius: 10px; text-align: center; margin: 20px 0;\">\n<h2 style=\"color: white; margin: 0 0 10px 0;\">🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨</h2>\n<p style=\"font-size: 18px; margin: 0 0 15px 0;\">I can no longer upload new models unless I can cover the cost of additional storage.<br>I host <b>70+ free models</b> as an independent contributor and this work is unpaid.<br><b>Without your support, no more new models can be uploaded.</b></p>\n<p style=\"font-size: 20px; margin: 0;\">\n<a href=\"https://patreon.com/LLMfan46\" style=\"color: white; text-decoration: underline;\">🎉 Patreon (Monthly)</a> | \n<a href=\"https://ko-fi.com/llmfan46\" style=\"color: white; text-decoration: underline;\">☕ Ko-fi (One-time)</a>\n</p>\n<p style=\"font-size: 16px; margin: 10px 0 0 0;\">Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone.</p>\n</div>\n\n---\n\n# gpt-oss-120b-heretic-v2-GGUF\n\nGGUF quantizations of [llmfan46/gpt-oss-120b-heretic-v2](https://huggingface.co/llmfan46/gpt-oss-120b-heretic-v2).\n\n# This is a decensored version of [openai/gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b), made using [Heretic](https://github.com/p-e-w/heretic) v1.2.0 with the [Arbitrary-Rank Ablation (ARA)](https://github.com/p-e-w/heretic/pull/211) method\n\n## Abliteration parameters\n\n| Parameter | Value |\n| :-------- | :---: |\n| **start_layer_index** | 18 |\n| **end_layer_index** | 23 |\n| **preserve_good_behavior_weight** | 0.0253 |\n| **steer_bad_behavior_weight** | 0.0099 |\n| **overcorrect_relative_weight** | 0.9989 |\n| **neighbor_count** | 10 |\n\n## Performance\n\n| Metric | This model | Original model ([gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b)) |\n| :----- | :--------: | :---------------------------: |\n| **KL divergence** | 0.0179 | 0 *(by definition)* |\n| **Refusals** | 9/100 | 98/100 |\n\nLower refusals indicate fewer content restrictions, while lower KL divergence indicates better preservation of the original model's capabilities. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections, while higher KL divergence degrades coherence, reasoning ability, and overall quality.\n\n## Usage\n\nWorks with llama.cpp, LM Studio, Ollama, and other GGUF-compatible tools.\n\n-----\n\n\n<p align=\"center\">\n <img alt=\"gpt-oss-120b\" src=\"https://raw.githubusercontent.com/openai/gpt-oss/main/docs/gpt-oss-120b.svg\">\n</p>\n\n<p align=\"center\">\n <a href=\"https://gpt-oss.com\"><strong>Try gpt-oss</strong></a> ·\n <a href=\"https://cookbook.openai.com/topic/gpt-oss\"><strong>Guides</strong></a> ·\n <a href=\"https://arxiv.org/abs/2508.10925\"><strong>Model card</strong></a> ·\n <a href=\"https://openai.com/index/introducing-gpt-oss/\"><strong>OpenAI blog</strong></a>\n</p>\n\n<br>\n\nWelcome to the gpt-oss series, [OpenAI’s open-weight models](https://openai.com/open-models) designed for powerful reasoning, agentic tasks, and versatile developer use cases.\n\nWe’re releasing two flavors of these open models:\n- `gpt-oss-120b` — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters)\n- `gpt-oss-20b` — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters)\n\nBoth models were trained on our [harmony response format](https://github.com/openai/harmony) and should only be used with the harmony format as it will not work correctly otherwise.\n\n\n> [!NOTE]\n> This model card is dedicated to the larger `gpt-oss-120b` model. Check out [`gpt-oss-20b`](https://huggingface.co/openai/gpt-oss-20b) for the smaller model.\n\n# Highlights\n\n* **Permissive Apache 2.0 license:** Build freely without copyleft restrictions or patent risk—ideal for experimentation, customization, and commercial deployment. \n* **Configurable reasoning effort:** Easily adjust the reasoning effort (low, medium, high) based on your specific use case and latency needs. \n* **Full chain-of-thought:** Gain complete access to the model’s reasoning process, facilitating easier debugging and increased trust in outputs. It’s not intended to be shown to end users. \n* **Fine-tunable:** Fully customize models to your specific use case through parameter fine-tuning.\n* **Agentic capabilities:** Use the models’ native capabilities for function calling, [web browsing](https://github.com/openai/gpt-oss/tree/main?tab=readme-ov-file#browser), [Python code execution](https://github.com/openai/gpt-oss/tree/main?tab=readme-ov-file#python), and Structured Outputs.\n* **MXFP4 quantization:** The models were post-trained with MXFP4 quantization of the MoE weights, making `gpt-oss-120b` run on a single 80GB GPU (like NVIDIA H100 or AMD MI300X) and the `gpt-oss-20b` model run within 16GB of memory. All evals were performed with the same MXFP4 quantization.\n\n---\n\n# Inference examples\n\n## Transformers\n\nYou can use `gpt-oss-120b` and `gpt-oss-20b` with Transformers. If you use the Transformers chat template, it will automatically apply the [harmony response format](https://github.com/openai/harmony). If you use `model.generate` directly, you need to apply the harmony format manually using the chat template or use our [openai-harmony](https://github.com/openai/harmony) package.\n\nTo get started, install the necessary dependencies to setup your environment:\n\n```\npip install -U transformers kernels torch \n```\n\nOnce, setup you can proceed to run the model by running the snippet below:\n\n```py\nfrom transformers import pipeline\nimport torch\n\nmodel_id = \"openai/gpt-oss-120b\"\n\npipe = pipeline(\n \"text-generation\",\n model=model_id,\n torch_dtype=\"auto\",\n device_map=\"auto\",\n)\n\nmessages = [\n {\"role\": \"user\", \"content\": \"Explain quantum mechanics clearly and concisely.\"},\n]\n\noutputs = pipe(\n messages,\n max_new_tokens=256,\n)\nprint(outputs[0][\"generated_text\"][-1])\n```\n\nAlternatively, you can run the model via [`Transformers Serve`](https://huggingface.co/docs/transformers/main/serving) to spin up a OpenAI-compatible webserver:\n\n```\ntransformers serve\ntransformers chat localhost:8000 --model-name-or-path openai/gpt-oss-120b\n```\n\n[Learn more about how to use gpt-oss with Transformers.](https://cookbook.openai.com/articles/gpt-oss/run-transformers)\n\n## vLLM\n\nvLLM recommends using [uv](https://docs.astral.sh/uv/) for Python dependency management. You can use vLLM to spin up an OpenAI-compatible webserver. The following command will automatically download the model and start the server.\n\n```bash\nuv pip install --pre vllm==0.10.1+gptoss \\\n --extra-index-url https://wheels.vllm.ai/gpt-oss/ \\\n --extra-index-url https://download.pytorch.org/whl/nightly/cu128 \\\n --index-strategy unsafe-best-match\n\nvllm serve openai/gpt-oss-120b\n```\n\n[Learn more about how to use gpt-oss with vLLM.](https://cookbook.openai.com/articles/gpt-oss/run-vllm)\n\n## PyTorch / Triton\n\nTo learn about how to use this model with PyTorch and Triton, check out our [reference implementations in the gpt-oss repository](https://github.com/openai/gpt-oss?tab=readme-ov-file#reference-pytorch-implementation).\n\n## Ollama\n\nIf you are trying to run gpt-oss on consumer hardware, you can use Ollama by running the following commands after [installing Ollama](https://ollama.com/download).\n\n```bash\n# gpt-oss-120b\nollama pull gpt-oss:120b\nollama run gpt-oss:120b\n```\n\n[Learn more about how to use gpt-oss with Ollama.](https://cookbook.openai.com/articles/gpt-oss/run-locally-ollama)\n\n#### LM Studio\n\nIf you are using [LM Studio](https://lmstudio.ai/) you can use the following commands to download.\n\n```bash\n# gpt-oss-120b\nlms get openai/gpt-oss-120b\n```\n\nCheck out our [awesome list](https://github.com/openai/gpt-oss/blob/main/awesome-gpt-oss.md) for a broader collection of gpt-oss resources and inference partners.\n\n---\n\n# Download the model\n\nYou can download the model weights from the [Hugging Face Hub](https://huggingface.co/collections/openai/gpt-oss-68911959590a1634ba11c7a4) directly from Hugging Face CLI:\n\n```shell\n# gpt-oss-120b\nhuggingface-cli download openai/gpt-oss-120b --include \"original/*\" --local-dir gpt-oss-120b/\npip install gpt-oss\npython -m gpt_oss.chat model/\n```\n\n# Reasoning levels\n\nYou can adjust the reasoning level that suits your task across three levels:\n\n* **Low:** Fast responses for general dialogue. \n* **Medium:** Balanced speed and detail. \n* **High:** Deep and detailed analysis.\n\nThe reasoning level can be set in the system prompts, e.g., \"Reasoning: high\".\n\n# Tool use\n\nThe gpt-oss models are excellent for:\n* Web browsing (using built-in browsing tools)\n* Function calling with defined schemas\n* Agentic operations like browser tasks\n\n# Fine-tuning\n\nBoth gpt-oss models can be fine-tuned for a variety of specialized use cases.\n\nThis larger model `gpt-oss-120b` can be fine-tuned on a single H100 node, whereas the smaller [`gpt-oss-20b`](https://huggingface.co/openai/gpt-oss-20b) can even be fine-tuned on consumer hardware.\n\n# Citation\n\n```bibtex\n@misc{openai2025gptoss120bgptoss20bmodel,\n title={gpt-oss-120b & gpt-oss-20b Model Card}, \n author={OpenAI},\n year={2025},\n eprint={2508.10925},\n archivePrefix={arXiv},\n primaryClass={cs.CL},\n url={https://arxiv.org/abs/2508.10925}, \n}\n```",
"related_quantizations": []
},
"tags": [
"gguf",
"heretic",
"uncensored",
"decensored",
"abliterated",
"text-generation",
"arxiv:2508.10925",
"base_model:llmfan46/gpt-oss-120b-heretic-v2",
"base_model:quantized:llmfan46/gpt-oss-120b-heretic-v2",
"license:apache-2.0",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 9,
"downloads": 1654,
"gated": false,
"private": false,
"last_modified": "2026-03-27T22:31:17.000Z",
"created_at": "2026-03-08T01:22:47.000Z",
"pipeline_tag": "text-generation",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69accf675d4f2860e2e88446",
"id": "llmfan46/gpt-oss-120b-heretic-v2-GGUF",
"modelId": "llmfan46/gpt-oss-120b-heretic-v2-GGUF",
"sha": "4424c0373bbe6a883ea6593e169cdc2ca10cc8db",
"createdAt": "2026-03-08T01:22:47.000Z",
"lastModified": "2026-03-27T22:31:17.000Z",
"author": "llmfan46",
"downloads": 1654,
"likes": 9,
"gated": false,
"private": false,
"pipeline_tag": "text-generation",
"library_name": "",
"siblings_count": 3
}