GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AdvancedDataIntelligence/adi-llama-3.1-8b-ablit-glm5.2-GGUF overview

<img src="https://serve.thelabsource.com/u/5OaCZV.png" alt="adi llama 3.1 8b ablit glm5.2" width="800" adi llama 3.1 8b ablit glm5.2 Part of the ADI Advanced D…

ggufdistillationllamallama-3.1abliterateduncensoredadiadvanced-data-intelligencetext-generationtool-callingenbase_model:huihui-ai/Meta-Llama-3.1-8B-Instruct-abliteratedbase_model:quantized:huihui-ai/Meta-Llama-3.1-8B-Instruct-abliteratedlicense:llama3.1endpoints_compatibleregion:usconversational

Runs locally from ~4.58 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
adi-llama-3.1-8b-ablit-glm5.2-q4_k_m.ggufGGUFQ4_K_M4.58 GBDownload

Model Details

Model IDAdvancedDataIntelligence/adi-llama-3.1-8b-ablit-glm5.2-GGUF
AuthorAdvancedDataIntelligence
Pipelinetext-generation
Licensellama3.1
Base modelhuihui-ai/Meta-Llama-3.1-8B-Instruct-abliterated
Last modified2026-07-02T03:08:35.000Z

Model README

---

license: llama3.1

base_model: huihui-ai/Meta-Llama-3.1-8B-Instruct-abliterated

tags:

- gguf

- distillation

- llama

- llama-3.1

- abliterated

- uncensored

- adi

- advanced-data-intelligence

- text-generation

- tool-calling

language:

- en

pipeline_tag: text-generation

library_name: gguf

---

<img src="https://serve.thelabsource.com/u/5OaCZV.png" alt="adi-llama-3.1-8b-ablit-glm5.2" width="800">

adi-llama-3.1-8b-ablit-glm5.2

Part of the ADI (Advanced Data Intelligence) model line — ADI Llama series.

An uncensored, fully local model that reasons and answers like a frontier

teacher. Built by distilling glm-5.2 general-knowledge responses into an

abliterated Llama-3.1-8B student with a light 4-bit QLoRA fine-tune, then

merged, converted, and quantized to GGUF. The base is an abliterated

(refusal-suppressed) Meta-Llama-3.1-8B-Instruct, and the distillation was designed

to add the teacher's answer quality without restoring refusal behavior. The

student base retains native tool calling and a long context window.

Capabilities

| Size | Context | Input | Output | Tools |

|---|---|---|---|---|

| 4.92 GB | 128K | 🅣 Text | Text | ✅ |

| | |

|---|---|

| Base model | huihui-ai/Meta-Llama-3.1-8B-Instruct-abliterated (abliterated Meta-Llama-3.1-8B-Instruct) |

| Teacher | glm-5.2 (responses distilled, thinking disabled) |

| Method | Light 4-bit QLoRA SFT (rank 16, 2 epochs) → merge → GGUF |

| Quantization | Q4_K_M (~4.92 GB) |

| License | Llama 3.1 Community License (inherited from Llama-3.1-8B) |

| Context | 128K (inherited from base) |

| Tool calling | Supported (inherited from base) |

Run it

Pull directly into Ollama:

ollama run hf.co/AdvancedDataIntelligence/adi-llama-3.1-8b-ablit-glm5.2-GGUF:Q4_K_M

Or download the .gguf and point any llama.cpp-based runtime at it.

What this model is

This is a knowledge distillation: a strong teacher (glm-5.2) generated

high-quality answers across a clean general-knowledge prompt set, and the

abliterated Llama-3.1-8B student was fine-tuned to imitate them. The result reasons

and responds more like its teacher on general topics while keeping the base's

uncensored character.

What distillation does — and doesn't do. It transfers the teacher's

reasoning style and answer quality, not net-new facts. For raw factual recall,

retrieval-augmented generation (RAG) is the right tool, not fine-tuning. What you

get here is an 8B that structures and explains like a larger model on topics it

already partly knows — without the refusal behavior of an aligned model.

Uncensored behavior — please read

This model is built on an abliterated base: the refusal direction has been

suppressed, so it will attempt most requests rather than declining them. The

fine-tune was intentionally kept light (2 epochs, benign-only data) to **avoid

re-introducing refusals**, and post-training spot checks confirmed the model still

answers helpfully without added hedging.

You are responsible for using it lawfully and ethically. It has weaker built-in

safety guardrails than stock Meta-Llama-3.1-8B-Instruct; apply your own filtering

and oversight for any production or public-facing deployment.

Training

| Metric | Value |

|---|---|

| Training pairs | 2,000 (deterministic subset of a 4,982-pair clean set) |

| Epochs | 2 (kept light to preserve abliteration) |

| Steps | 500 |

| Final train loss | 1.1143 |

| LoRA rank / alpha | 16 / 16 |

| Trainable params | 41.9M |

| Precision | 4-bit QLoRA (nf4) |

| Peak VRAM | 9.66 GB |

| Hardware | single RTX 5060 Ti (16 GB) |

| Training time | 1.73 h (~12 s/step) |

The seed prompts were drawn from the human-written

Databricks Dolly-15k

dataset (filtered to remove items requiring an attached context passage, then

deduplicated). The teacher was queried with thinking disabled so the student

learns clean final answers rather than chain-of-thought.

Notes for re-builders

  • Distilling onto an abliterated base is a balancing act. Any SFT can nudge an

abliterated model back toward refusals. Two choices kept the uncensored behavior

intact: benign-only training data (the GLM-5.2 set has zero refusals to

re-learn) and a light touch (LoRA rank 16, 2 epochs). Spot-check refusals

before/after.

  • Llama 3.1 is a standard LlamaForCausalLM — no arch surprises. 4-bit QLoRA

via Unsloth with gradient checkpointing ("unsloth" mode), max_seq_length 2048,

per-device batch 2 × grad-accum 4, adamw_8bit, LoRA targeting all attention + MLP

projections. Peak VRAM 9.66 GB on a 16 GB card.

  • GGUF conversion via streaming LoRA merge → f16 GGUF → Q4_K_M quantize with

llama.cpp (convert_hf_to_gguf.py). Use the standard Llama-3.1 chat template

(<|start_header_id|> / <|eot_id|>) in the Modelfile.

Intended use

General-purpose local assistant for users who want a capable, private,

offline-capable model with minimal refusal behavior: explanations, reasoning,

creative writing, and tool-calling workflows. Not intended as a source of

authoritative facts without retrieval, and not a substitute for your own safety

review.

License

Llama 3.1 Community License, inherited from the Llama-3.1-8B lineage via the

abliterated

base model.

Review the Llama 3.1 license and

Acceptable Use Policy — note the

attribution requirement ("Built with Llama") and use restrictions. Distilled

training data was generated using glm-5.2; users should review the teacher model's

terms for their own use case.

---

Built at theLAB — Learning. Algorithms. Breakthroughs.

Run AdvancedDataIntelligence/adi-llama-3.1-8b-ablit-glm5.2-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models