Blackfrost-AI/GLM-5.3-Flash-DERISKED-GGUF overview
<div align="center" Blackfrost https://cdn uploads.huggingface.co/production/uploads/69a27f2d114e4ac9de4dafc7/xDTdhLFXmKlZOazFcvJ5S.jpeg <h1 GLM 5.3 Flash DERI…
Runs locally from ~9.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| GLM-5.3-Flash-DERISKED-Q3_K_M-00001-of-00014.gguf | GGUF | Q3_K_M | 9.0 MB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00002-of-00014.gguf | GGUF | Q3_K_M | 10.85 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00003-of-00014.gguf | GGUF | Q3_K_M | 11.10 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00004-of-00014.gguf | GGUF | Q3_K_M | 10.80 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00005-of-00014.gguf | GGUF | Q3_K_M | 10.96 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00006-of-00014.gguf | GGUF | Q3_K_M | 11.15 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00007-of-00014.gguf | GGUF | Q3_K_M | 10.80 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00008-of-00014.gguf | GGUF | Q3_K_M | 10.87 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00009-of-00014.gguf | GGUF | Q3_K_M | 11.10 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00010-of-00014.gguf | GGUF | Q3_K_M | 10.80 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00011-of-00014.gguf | GGUF | Q3_K_M | 10.88 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00012-of-00014.gguf | GGUF | Q3_K_M | 11.10 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00013-of-00014.gguf | GGUF | Q3_K_M | 10.81 GB | Download |
| GLM-5.3-Flash-DERISKED-Q3_K_M-00014-of-00014.gguf | GGUF | Q3_K_M | 10.88 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00001-of-00014.gguf | GGUF | Q4_K_M | 9.0 MB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00002-of-00014.gguf | GGUF | Q4_K_M | 13.42 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00003-of-00014.gguf | GGUF | Q4_K_M | 14.08 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00004-of-00014.gguf | GGUF | Q4_K_M | 13.52 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00005-of-00014.gguf | GGUF | Q4_K_M | 13.72 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00006-of-00014.gguf | GGUF | Q4_K_M | 13.56 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00007-of-00014.gguf | GGUF | Q4_K_M | 13.51 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00008-of-00014.gguf | GGUF | Q4_K_M | 14.18 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00009-of-00014.gguf | GGUF | Q4_K_M | 13.52 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00010-of-00014.gguf | GGUF | Q4_K_M | 13.50 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00011-of-00014.gguf | GGUF | Q4_K_M | 14.18 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00012-of-00014.gguf | GGUF | Q4_K_M | 15.26 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00013-of-00014.gguf | GGUF | Q4_K_M | 13.54 GB | Download |
| GLM-5.3-Flash-DERISKED-Q4_K_M-00014-of-00014.gguf | GGUF | Q4_K_M | 13.62 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00001-of-00014.gguf | GGUF | Q5_K_M | 9.0 MB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00002-of-00014.gguf | GGUF | Q5_K_M | 15.83 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00003-of-00014.gguf | GGUF | Q5_K_M | 16.39 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00004-of-00014.gguf | GGUF | Q5_K_M | 16.10 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00005-of-00014.gguf | GGUF | Q5_K_M | 16.33 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00006-of-00014.gguf | GGUF | Q5_K_M | 16.16 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00007-of-00014.gguf | GGUF | Q5_K_M | 16.09 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00008-of-00014.gguf | GGUF | Q5_K_M | 16.49 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00009-of-00014.gguf | GGUF | Q5_K_M | 16.10 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00010-of-00014.gguf | GGUF | Q5_K_M | 16.09 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00011-of-00014.gguf | GGUF | Q5_K_M | 16.50 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00012-of-00014.gguf | GGUF | Q5_K_M | 16.99 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00013-of-00014.gguf | GGUF | Q5_K_M | 16.12 GB | Download |
| GLM-5.3-Flash-DERISKED-Q5_K_M-00014-of-00014.gguf | GGUF | Q5_K_M | 16.21 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00001-of-00014.gguf | GGUF | Q6_K | 9.0 MB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00002-of-00014.gguf | GGUF | Q6_K | 18.41 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00003-of-00014.gguf | GGUF | Q6_K | 18.84 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00004-of-00014.gguf | GGUF | Q6_K | 18.85 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00005-of-00014.gguf | GGUF | Q6_K | 19.11 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00006-of-00014.gguf | GGUF | Q6_K | 18.92 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00007-of-00014.gguf | GGUF | Q6_K | 18.84 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00008-of-00014.gguf | GGUF | Q6_K | 18.96 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00009-of-00014.gguf | GGUF | Q6_K | 18.85 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00010-of-00014.gguf | GGUF | Q6_K | 18.84 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00011-of-00014.gguf | GGUF | Q6_K | 18.97 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00012-of-00014.gguf | GGUF | Q6_K | 18.84 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00013-of-00014.gguf | GGUF | Q6_K | 18.86 GB | Download |
| GLM-5.3-Flash-DERISKED-Q6_K-00014-of-00014.gguf | GGUF | Q6_K | 18.97 GB | Download |
Model Details
| Model ID | Blackfrost-AI/GLM-5.3-Flash-DERISKED-GGUF |
|---|---|
| Author | Blackfrost-AI |
| Pipeline | text-generation |
| License | mit |
| Base model | Blackfrost-AI/GLM-5.3-Flash-DERISKED-BF16 |
| Last modified | 2026-09-04T15:48:42.000Z |
Model README
---
license: mit
base_model:
- Blackfrost-AI/GLM-5.3-Flash-DERISKED-BF16
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
language:
- en
- zh
tags:
- glm
- glm-5.3
- glm5next
- mixture-of-experts
- moe
- derisked
- gguf
- llama.cpp
- blackwell
- research
- security-research
- red-teaming
- gated
extra_gated_heading: 18+ Research Access Request
extra_gated_prompt: >
Access is free but manually reviewed. You must be at least 18 years old and
accept the research-only terms in this model card. Submit an accurate
intended-use statement. Approval is discretionary and may be revoked for
misuse or material misrepresentation.
extra_gated_fields:
I confirm that I am at least 18 years old: checkbox
I will use this model only for lawful research, security evaluation, red teaming, or controlled local testing: checkbox
I accept responsibility for access control, output review, logging, and compliance with applicable law: checkbox
I have read and accept the license, disclaimer, and research-only terms in this model card: checkbox
Research affiliation or independent researcher: text
Intended research use: text
---
<div align="center">
<h1>GLM-5.3-Flash-DERISKED-GGUF</h1>
<h3>Standard and quality-profile GGUF builds for 2–4 RTX PRO 6000 Blackwell GPUs</h3>
<p><strong>Built by <a href="https://x.com/Blackfrost_AI">Blackfrost</a> · Las Vegas, NV</strong></p>
<p>
<img src="https://img.shields.io/badge/Free_research_access-047857?style=for-the-badge" />
<img src="https://img.shields.io/badge/18%2B_manual_gate-b45309?style=for-the-badge" />
<img src="https://img.shields.io/badge/GGUF-1f2937?style=for-the-badge" />
<img src="https://img.shields.io/badge/Family_harmful_refusal-1.3%25-047857?style=for-the-badge" />
</p>
</div>
---
> ## Standard GGUF release
>
> Q3_K_M is released after clean load, generation, and operator behavioral snap-back approval. Q4_K_M, Q5_K_M, and Q6_K completed artifact-integrity verification and are released under the operator-approved standard K-quant pipeline. The BF16 master is not included in this repository.
Release contents
This repository contains four standard text-generation GGUF variants. Each completed variant retains the checkpoint's NextN/MTP tensors and passes shard, checksum, and tensor-integrity audits before publication.
| variant | intended hardware | size / shards | release status |
|---|---:|---:|---|
| Q3_K_M | 2× RTX PRO 6000 Blackwell 96 GB | 142.11 GiB / 14 | released — operator approved |
| Q4_K_M | 2–3× RTX PRO 6000 Blackwell 96 GB | 179.61 GiB / 14 | released — integrity verified |
| Q5_K_M | 3× RTX PRO 6000 Blackwell 96 GB | 211.42 GiB / 14 | released — integrity verified |
| Q6_K | 4× RTX PRO 6000 Blackwell 96 GB | 245.24 GiB / 14 | released — integrity verified |
Final file sizes, shard counts, checksums, and measured runtime figures are added only after each artifact is complete and verified. No BF16 weights are planned for this repository.
Model summary
| | |
|---|---|
| Architecture | Glm5NextForConditionalGeneration / glm5next hybrid MoE |
| Official base | zai-org/GLM-5.3-Flash-BF16 |
| Parameters | 320B total, 18B active per token, as reported by Z.ai |
| Experts | 288 routed, 8 active per token |
| Text blocks | 45 main blocks plus one retained NextN/MTP block |
| Architectural context | up to 1,048,576 tokens; practical limits depend on hardware and runtime |
| Format in this repo | sharded GGUF, text generation only |
| Refusal evaluation | Release-family reference: 1.3% harmful · 1.1% overall |
GLM-5.3-Flash is natively multimodal upstream. This repository's planned GGUF files contain the text model only; no vision projector is currently promised.
Lineage
The GGUF files are derived from the Blackfrost BF16 master based on the official Z.ai BF16 checkpoint. The BF16 master is not part of this free GGUF repository.
Validation status
The card is intentionally conservative while weights are being tested. For released Q3_K_M:
- all 14 shards passed local and durable checksum verification;
- all retained NextN/MTP tensors passed the tensor audit;
- the exact released candidate passed a clean two-GPU load and short generation;
- operator behavioral snap-back validation passed;
- the final release-family refusal reference is published below; Q3 has not been independently rerun across all 450 prompts.
For released Q4_K_M, Q5_K_M, and Q6_K:
- all 14 shards for each variant passed local and durable checksum verification;
- all retained NextN/MTP tensors passed the tensor audit;
- runtime memory, throughput, context, and quantitative refusal results have not been measured separately for these variants;
- no refusal-rate or capability-retention claim is made.
Use the table below only as a release-family behavior reference. Do not cite it as a variant-specific refusal rate, capability score, or deployment performance figure.
Refusal evaluation
The release-family reference below was measured on the behavior-matched NVFP4 deployment
checkpoint using R1-HARMFUL-BENCH-450 under a bare chat configuration. Responses were reviewed
after generation to distinguish actual refusals from false-positive string matches.
Configuration: thinking enabled · maximum reasoning effort · temperature 1.0 · top-p 0.95 ·
top-k omitted · maximum 16,384 output tokens
| Evaluation slice | Final judged refusals |
|---|---:|
| Harmful prompts | 4 / 300 (1.3%) |
| Full suite | 5 / 450 (1.1%) |
| API errors | 0 / 450 |
These figures are a release-family reference. This format was not independently rerun across all
450 prompts, so the table should not be represented as a format-specific measurement. The results
are behavioral observations, not a safety certification.
---
Serving notes
GLM-5.3 GGUF support currently requires a compatible llama.cpp build with glm5next support. The released Q3_K_M validation configuration is:
export NVIDIA_TF32_OVERRIDE=0
llama-server \
-m GLM-5.3-Flash-DERISKED-Q3_K_M-00001-of-00014.gguf \
-ngl all -sm layer -ts 1,1 -fa off \
-c 4096 --host 0.0.0.0 --port 8080
Use the first shard as the model path; llama.cpp resolves the remaining shards automatically. GPU count and --tensor-split must match the selected variant and available memory.
The NextN/MTP weights are retained in each planned artifact. Do not infer speculative-decoding support from their presence alone; use only a runtime configuration verified for this architecture.
Access terms — 18+ research only
Access is free and manually approved. It is limited to applicants who:
- are at least 18 years old;
- provide an accurate research purpose;
- use the weights only for lawful research, red teaming, security evaluation, or controlled local testing;
- maintain appropriate authentication, access controls, isolation, logging, and human review;
- comply with the upstream MIT license and all applicable laws and institutional rules; and
- do not represent this checkpoint as a safety-stock model or as validated beyond the results published here.
Access is personal to the approved Hugging Face account. Do not redistribute the weights, mirror them, transfer access, or use another person's approval. Blackfrost may deny or revoke access for inaccurate applications, misuse, redistribution, or breach of these terms.
Disclaimer
This checkpoint has a deliberately altered refusal profile and is intended for controlled research. It is not a safety boundary, policy engine, authorization system, or substitute for application-level safeguards.
The model is provided "as is", without warranty of any kind. Outputs may be inaccurate, offensive, unsafe, unlawful, or otherwise unsuitable. Blackfrost makes no guarantee that any prompt will be accepted or refused, that upstream capabilities are retained, or that behavior generalizes across samplers, prompts, context lengths, tools, modalities, runtimes, or hardware.
Operators are solely responsible for lawful use, secure deployment, tool permissions, data handling, output review, monitoring, incident response, and downstream consequences. Do not connect the model to real systems, accounts, credentials, infrastructure, or physical processes without independent controls appropriate to the risk.
Any modification, merge, fine-tune, conversion, or requantization produces an artifact Blackfrost has not evaluated and does not characterize.
License and attribution
The official GLM-5.3-Flash-BF16 checkpoint is released under the MIT License by Z.AI Co., Ltd. The upstream license and copyright notice apply to this derivative and must be preserved. Review the official model card before use.
Contact Blackfrost
<div align="center">
<h3><a href="https://x.com/Blackfrost_AI">@Blackfrost_AI</a> on X</h3>
<p><strong>Blackfrost</strong> · Las Vegas, Nevada<br>
Frontier model engineering</p>
</div>
---
<div align="center">
<p><strong>GLM-5.3-Flash-DERISKED-GGUF</strong> · © 2026 Blackfrost Softwares Corp.</p>
</div>
Run Blackfrost-AI/GLM-5.3-Flash-DERISKED-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models