GraySoft
Projects Models About FAQ Contact Download guIDE →

aessedai/glm-4.6-derestricted-gguf IQ4_NL GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

aessedai/glm-4.6-derestricted-gguf overview

Quick update: I've fixed an issue where the chat template wasn't inluded in the quants, the first shard of each quant has been updated to include the chat template. Please re-download the first shard to pick up the fix, sorry for the inconvenience. --- This is a "derestricted" abliteration of GLM-4.6, using Jim Lai's norm-preserving biprojected abliteration technique. For more information, you can read his blog post here Essentially, I was going for a lighter abliteration. This doesn't mean the model is 100% zero-shot unrestricted. It should be more "permissive" than normal GLM-4.6, but probably still requires a system prompt to nudge it in the right direction. From my own testing, I've mainly used this model for creative writing. I've noticed a positive change compared to how base GLM-4.6 does sentence structure and this feels more varied and organic. It does not particularly reduce or alter "slop", since this isn't a finetune, but there's much less of an "assistant" voice performing soft-censorship during particular scenarios and it feels less like "LLM writing". I've only done some light technical assistant work and it still feels competent there, but I haven't exhaustively benched it. Visualized here is the analysis of the refusal direction: !analysis Provided in this repository are several quants I've produced from the abliteration I performed, as well as the measurements to produce your own abliteration if you want and the config that I used. I chose to ablate layers 30-45, using the measurement from layer 37 due to the SNR peak. Other measurements I tried showed an interesting dual-peak phenomenon with a second peak forming around layer 46, but the overall SNR magitude was only ~0.16 or so compared to the much better 0.25 peak present here. If you want to abliterate GLM-4.6 yourself, you will need to download the safetensors for the model and use this PR. For quants, I've provided a Q80 as well as others that follow the MoE quantization schema that I've been using. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. The naming convention is as follows: [Default Type]-[FFNUP]-[FFNGATE]-[FFNDOWN], eg: Q80-Q4K-Q4K-Q5K. This means: | Quant | Size | PPL | KLD | | :--------- | :--------- | :------- | :------- | | Q80 | 353.26 GiB (8.51 BPW) | 8.4801 ± 0.15099 | 0 | | Q80-Q5K-Q5K-Q6K | 248.61 GiB (5.99 BPW) | 8.4881 ± 0.15112 | 0.009449 ± 0.000677 | | Q80-Q4K-Q4K-Q5K | 208.24 GiB (5.01 BPW) | 8.5182 ± 0.15172 | 0.016299 ± 0.000839 | | IQ4NL | 187.40 GiB (4.51 BPW) | 8.6026 ± 0.15331 | 0.029524 ± 0.000858 | | Q80-IQ3S-IQ3S-IQ4XS | 163.74 GiB (3.94 BPW) | 8.7101 ± 0.15534 | 0.041096 ± 0.001202 | | Q6K-IQ2XS-IQ2XS-IQ3S | 119.79 GiB (2.88 BPW) | 9.3447 ± 0.16732 | 0.131974 ± 0.002384 | | Q5K-IQ2XXS-IQ2XXS-IQ3XXS | 106.50 GiB (2.56 BPW) | 9.5127 ± 0.17040 | 0.174152 ± 0.002976 | !pplratiovs_kld

ggufbase_model:zai-org/GLM-4.6base_model:quantized:zai-org/GLM-4.6endpoints_compatibleregion:usimatrixconversational
aessedai/glm-4.6-derestricted-gguf visual
Downloads
787
Likes
26
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

35 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
GLM-4.6-Derestricted-IQ4_NL-00001-of-00005.gguf GGUF IQ4_NL 46.29 GB Download
GLM-4.6-Derestricted-IQ4_NL-00002-of-00005.gguf GGUF IQ4_NL 46.17 GB Download
GLM-4.6-Derestricted-IQ4_NL-00003-of-00005.gguf GGUF IQ4_NL 46.09 GB Download
GLM-4.6-Derestricted-IQ4_NL-00004-of-00005.gguf GGUF IQ4_NL 46.10 GB Download
GLM-4.6-Derestricted-IQ4_NL-00005-of-00005.gguf GGUF IQ4_NL 2.76 GB Download
GLM-4.6-Derestricted-Q5_K-IQ2_XXS-IQ2_XXS-IQ3_XXS-00001-of-00003.gguf GGUF Q5_K 46.34 GB Download
GLM-4.6-Derestricted-Q5_K-IQ2_XXS-IQ2_XXS-IQ3_XXS-00002-of-00003.gguf GGUF Q5_K 46.36 GB Download
GLM-4.6-Derestricted-Q5_K-IQ2_XXS-IQ2_XXS-IQ3_XXS-00003-of-00003.gguf GGUF Q5_K 13.82 GB Download
GLM-4.6-Derestricted-Q6_K-IQ2_XS-IQ2_XS-IQ3_S-00001-of-00003.gguf GGUF Q6_K 46.46 GB Download
GLM-4.6-Derestricted-Q6_K-IQ2_XS-IQ2_XS-IQ3_S-00002-of-00003.gguf GGUF Q6_K 46.23 GB Download
GLM-4.6-Derestricted-Q6_K-IQ2_XS-IQ2_XS-IQ3_S-00003-of-00003.gguf GGUF Q6_K 27.11 GB Download
GLM-4.6-Derestricted-Q8_0-00001-of-00008.gguf GGUF 45.51 GB Download
GLM-4.6-Derestricted-Q8_0-00002-of-00008.gguf GGUF 45.37 GB Download
GLM-4.6-Derestricted-Q8_0-00003-of-00008.gguf GGUF 45.50 GB Download
GLM-4.6-Derestricted-Q8_0-00004-of-00008.gguf GGUF 45.51 GB Download
GLM-4.6-Derestricted-Q8_0-00005-of-00008.gguf GGUF 45.37 GB Download
GLM-4.6-Derestricted-Q8_0-00006-of-00008.gguf GGUF 45.50 GB Download
GLM-4.6-Derestricted-Q8_0-00007-of-00008.gguf GGUF 45.51 GB Download
GLM-4.6-Derestricted-Q8_0-00008-of-00008.gguf GGUF 34.99 GB Download
GLM-4.6-Derestricted-Q8_0-IQ3_S-IQ3_S-IQ4_XS-00001-of-00004.gguf GGUF IQ3_S 46.26 GB Download
GLM-4.6-Derestricted-Q8_0-IQ3_S-IQ3_S-IQ4_XS-00002-of-00004.gguf GGUF IQ3_S 46.56 GB Download
GLM-4.6-Derestricted-Q8_0-IQ3_S-IQ3_S-IQ4_XS-00003-of-00004.gguf GGUF IQ3_S 45.94 GB Download
GLM-4.6-Derestricted-Q8_0-IQ3_S-IQ3_S-IQ4_XS-00004-of-00004.gguf GGUF IQ3_S 24.99 GB Download
GLM-4.6-Derestricted-Q8_0-Q4_K-Q4_K-Q5_K-00001-of-00005.gguf GGUF Q4_K 46.07 GB Download
GLM-4.6-Derestricted-Q8_0-Q4_K-Q4_K-Q5_K-00002-of-00005.gguf GGUF Q4_K 46.52 GB Download
GLM-4.6-Derestricted-Q8_0-Q4_K-Q4_K-Q5_K-00003-of-00005.gguf GGUF Q4_K 46.38 GB Download
GLM-4.6-Derestricted-Q8_0-Q4_K-Q4_K-Q5_K-00004-of-00005.gguf GGUF Q4_K 46.51 GB Download
GLM-4.6-Derestricted-Q8_0-Q4_K-Q4_K-Q5_K-00005-of-00005.gguf GGUF Q4_K 22.77 GB Download
GLM-4.6-Derestricted-Q8_0-Q5_K-Q5_K-Q6_K-00001-of-00006.gguf GGUF Q5_K 46.39 GB Download
GLM-4.6-Derestricted-Q8_0-Q5_K-Q5_K-Q6_K-00002-of-00006.gguf GGUF Q5_K 46.48 GB Download
GLM-4.6-Derestricted-Q8_0-Q5_K-Q5_K-Q6_K-00003-of-00006.gguf GGUF Q5_K 46.48 GB Download
GLM-4.6-Derestricted-Q8_0-Q5_K-Q5_K-Q6_K-00004-of-00006.gguf GGUF Q5_K 46.48 GB Download
GLM-4.6-Derestricted-Q8_0-Q5_K-Q5_K-Q6_K-00005-of-00006.gguf GGUF Q5_K 46.48 GB Download
GLM-4.6-Derestricted-Q8_0-Q5_K-Q5_K-Q6_K-00006-of-00006.gguf GGUF Q5_K 16.32 GB Download
imatrix.gguf GGUF 655.70 MB Download

Model Details Live

Model Slug
aessedai/glm-4.6-derestricted-gguf
Author
AesSedai
Pipeline Task
Library
Created
2025-12-03
Last Modified
2025-12-09
Gated
No
Private
No
HF SHA
aa795564d0d1adb32093a0baa266190a1bcbffa4
License
Unknown
Language
Unknown
Base Model
zai-org/GLM-4.6

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "base_model": [
      "zai-org/GLM-4.6"
    ],
    "frontmatter": {
      "base_model": [
        "zai-org/GLM-4.6"
      ]
    },
    "hero_image_url": "glm-4.6-refusal-analysis.png \"Chart showing four graphs with refusal information\"",
    "summary": "### Quick update: I've fixed an issue where the chat template wasn't inluded in the quants, the first shard of each quant has been updated to include the chat template. Please re-download the first shard to pick up the fix, sorry for the inconvenience. --- This is a \"derestricted\" abliteration of GLM-4.6, using Jim Lai's norm-preserving biprojected abliteration technique. For more information, you can read his blog post here Essentially, I was going for a lighter abliteration. This doesn't mean the model is 100% zero-shot unrestricted. It should be more \"permissive\" than normal GLM-4.6, but probably still requires a system prompt to nudge it in the right direction. From my own testing, I've mainly used this model for creative writing. I've noticed a positive change compared to how base GLM-4.6 does sentence structure and this feels more varied and organic. It does not particularly reduce or alter \"slop\", since this isn't a finetune, but there's much less of an \"assistant\" voice performing soft-censorship during particular scenarios and it feels less like \"LLM writing\". I've only done some light technical assistant work and it still feels competent there, but I haven't exhaustively benched it. Visualized here is the analysis of the refusal direction: !analysis Provided in this repository are several quants I've produced from the abliteration I performed, as well as the measurements to produce your own abliteration if you want and the config that I used. I chose to ablate layers 30-45, using the measurement from layer 37 due to the SNR peak. Other measurements I tried showed an interesting dual-peak phenomenon with a second peak forming around layer 46, but the overall SNR magitude was only ~0.16 or so compared to the much better 0.25 peak present here. If you want to abliterate GLM-4.6 yourself, you will need to download the safetensors for the model and use this PR. For quants, I've provided a Q8_0 as well as others that follow the MoE quantization schema that I've been using. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. The naming convention is as follows: [Default Type]-[FFN_UP]-[FFN_GATE]-[FFN_DOWN], eg: Q8_0-Q4_K-Q4_K-Q5_K. This means: | Quant | Size |    PPL   |    KLD   | | :--------- | :--------- | :------- | :------- | | Q8_0 | 353.26 GiB (8.51 BPW) | 8.4801 ± 0.15099 | 0 | | Q8_0-Q5_K-Q5_K-Q6_K | 248.61 GiB (5.99 BPW) | 8.4881 ± 0.15112 | 0.009449 ± 0.000677 | | Q8_0-Q4_K-Q4_K-Q5_K | 208.24 GiB (5.01 BPW) | 8.5182 ± 0.15172 | 0.016299 ± 0.000839 | | IQ4_NL | 187.40 GiB (4.51 BPW) | 8.6026 ± 0.15331 | 0.029524 ± 0.000858 | | Q8_0-IQ3_S-IQ3_S-IQ4_XS | 163.74 GiB (3.94 BPW) | 8.7101 ± 0.15534 | 0.041096 ± 0.001202 | | Q6_K-IQ2_XS-IQ2_XS-IQ3_S | 119.79 GiB (2.88 BPW) | 9.3447 ± 0.16732 | 0.131974 ± 0.002384 | | Q5_K-IQ2_XXS-IQ2_XXS-IQ3_XXS | 106.50 GiB (2.56 BPW) | 9.5127 ± 0.17040 | 0.174152 ± 0.002976 | !ppl_ratio_vs_kld",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nbase_model:\n- zai-org/GLM-4.6\n---\n\n\n### Quick update: I've fixed an issue where the chat template wasn't inluded in the quants, the first shard of each quant has been updated to include the chat template. Please re-download the first shard to pick up the fix, sorry for the inconvenience.\n---\n\nThis is a \"derestricted\" abliteration of GLM-4.6, using Jim Lai's norm-preserving biprojected abliteration technique. For more information, you can read his blog post [here](https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration)\n\nEssentially, I was going for a lighter abliteration. This doesn't mean the model is 100% zero-shot unrestricted. It should be more \"permissive\" than normal GLM-4.6, but probably still requires a system prompt to nudge it in the right direction. From my own testing, I've mainly used this model for creative writing. I've noticed a positive change compared to how base GLM-4.6 does sentence structure and this feels more varied and organic. It does not particularly reduce or alter \"slop\", since this isn't a finetune, but there's much less of an \"assistant\" voice performing soft-censorship during particular scenarios and it feels less like \"LLM writing\". I've only done some light technical assistant work and it still feels competent there, but I haven't exhaustively benched it.\n\nVisualized here is the analysis of the refusal direction:\n![analysis](glm-4.6-refusal-analysis.png \"Chart showing four graphs with refusal information\")\n\nProvided in this repository are several quants I've produced from the abliteration I performed, as well as the measurements to produce your own abliteration if you want and the config that I used.\nI chose to ablate layers 30-45, using the measurement from layer 37 due to the SNR peak. Other measurements I tried showed an interesting dual-peak phenomenon with a second peak forming around layer 46, but the overall SNR magitude was only ~0.16 or so compared to the much better 0.25 peak present here.\n\nIf you want to abliterate GLM-4.6 yourself, you will need to download the safetensors for the model and use this [PR](https://github.com/jim-plus/llm-abliteration/pull/8).\n\nFor quants, I've provided a Q8_0 as well as others that follow the MoE quantization schema that I've been using. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization.\n\nThe naming convention is as follows: `[Default Type]-[FFN_UP]-[FFN_GATE]-[FFN_DOWN]`, eg: `Q8_0-Q4_K-Q4_K-Q5_K`. This means:\n- Q8_0 is the default type (attention, shared expert, etc.)\n- Q4_K was used for the FFN_UP and FFN_GATE conditional expert tensors\n- Q5_K was used for the FFN_DOWN conditional expert tensors\n\n\n| Quant | Size |    PPL   |    KLD   |\n| :--------- | :--------- | :------- | :------- |\n| Q8_0 | 353.26 GiB (8.51 BPW) | 8.4801 ± 0.15099 | 0 |\n| Q8_0-Q5_K-Q5_K-Q6_K | 248.61 GiB (5.99 BPW) | 8.4881 ± 0.15112 | 0.009449 ± 0.000677 |\n| Q8_0-Q4_K-Q4_K-Q5_K | 208.24 GiB (5.01 BPW) | 8.5182 ± 0.15172 | 0.016299 ± 0.000839 |\n| IQ4_NL | 187.40 GiB (4.51 BPW) | 8.6026 ± 0.15331 | 0.029524 ± 0.000858 |\n| Q8_0-IQ3_S-IQ3_S-IQ4_XS | 163.74 GiB (3.94 BPW) | 8.7101 ± 0.15534 | 0.041096 ± 0.001202 |\n| Q6_K-IQ2_XS-IQ2_XS-IQ3_S | 119.79 GiB (2.88 BPW) | 9.3447 ± 0.16732 | 0.131974 ± 0.002384 |\n| Q5_K-IQ2_XXS-IQ2_XXS-IQ3_XXS | 106.50 GiB (2.56 BPW) | 9.5127 ± 0.17040 | 0.174152 ± 0.002976 |\n\n![ppl_ratio_vs_kld](ppl_ratio_vs_kld.png \"Chart showing PPL vs KLD analysis of quants\")\n\n\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "base_model:zai-org/GLM-4.6",
    "base_model:quantized:zai-org/GLM-4.6",
    "endpoints_compatible",
    "region:us",
    "imatrix",
    "conversational"
  ],
  "likes": 26,
  "downloads": 787,
  "gated": false,
  "private": false,
  "last_modified": "2025-12-09T16:58:22.000Z",
  "created_at": "2025-12-03T05:54:11.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "692fd0833b9c265389182e6b",
  "id": "AesSedai/GLM-4.6-Derestricted-GGUF",
  "modelId": "AesSedai/GLM-4.6-Derestricted-GGUF",
  "sha": "aa795564d0d1adb32093a0baa266190a1bcbffa4",
  "createdAt": "2025-12-03T05:54:11.000Z",
  "lastModified": "2025-12-09T16:58:22.000Z",
  "author": "AesSedai",
  "downloads": 787,
  "likes": 26,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 42
}