llmfan46/Nex-N2-mini-ultra-uncensored-heretic-GGUF overview
<div style="background color: ff4444; color: white; padding: 20px; border radius: 10px; text align: center; margin: 20px 0;" <h2 style="color: white; margin: 0β¦
Runs locally from ~861.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Nex-N2-mini-ultra-uncensored-heretic-BF16.gguf | GGUF | BF16 | 64.61 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q3_K_L.gguf | GGUF | Q3_K_L | 16.97 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q3_K_M.gguf | GGUF | Q3_K_M | 15.71 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q4_K_M.gguf | GGUF | Q4_K_M | 19.78 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q4_K_S.gguf | GGUF | Q4_K_S | 18.59 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q5_K_M.gguf | GGUF | Q5_K_M | 23.06 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q5_K_S.gguf | GGUF | Q5_K_S | 22.37 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q6_K.gguf | GGUF | Q6_K | 26.61 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-Q8_0.gguf | GGUF | Q8_0 | 34.37 GB | Download |
| Nex-N2-mini-ultra-uncensored-heretic-mmproj-BF16.gguf | GGUF | BF16 | 861.0 MB | Download |
Model Details
| Model ID | llmfan46/Nex-N2-mini-ultra-uncensored-heretic-GGUF |
|---|---|
| Author | llmfan46 |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | llmfan46/Nex-N2-mini-ultra-uncensored-heretic |
| Last modified | 2026-06-25T10:12:11.000Z |
Model README
---
license: apache-2.0
pipeline_tag: text-generation
library_name: transformers
tags:
- qwen3_5_moe
- heretic
- uncensored
- decensored
- abliterated
- mpoa
base_model:
- llmfan46/Nex-N2-mini-ultra-uncensored-heretic
---
<div style="background-color: #ff4444; color: white; padding: 20px; border-radius: 10px; text-align: center; margin: 20px 0;">
<h2 style="color: white; margin: 0 0 10px 0;">π¨β οΈ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT β οΈπ¨</h2>
<p style="font-size: 18px; margin: 0 0 15px 0;">I can no longer upload new models unless I can cover the cost of additional storage.<br>I host <b>70+ free models</b> as an independent contributor and this work is unpaid.<br><b>Without your support, no more new models can be uploaded.</b></p>
<p style="font-size: 20px; margin: 0;">
<a href="https://patreon.com/LLMfan46" style="color: white; text-decoration: underline;">π Patreon (Monthly)</a> |
<a href="https://ko-fi.com/llmfan46" style="color: white; text-decoration: underline;">β Ko-fi (One-time)</a>
</p>
<p style="font-size: 16px; margin: 10px 0 0 0;">Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone.</p>
</div>
---
93% fewer refusals (5/100 Uncensored vs 74/100 Original) while preserving model quality (0.0020 KL divergence).
β€οΈ Support My Work
Creating these models takes significant time, work and compute. If you find them useful consider supporting me:
| Platform | Link | What you get |
|----------|------|--------------|
| π Patreon | Monthly support | Priority model requests |
| β Ko-fi | One-time tip | My eternal gratitude |
Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.
-----
GGUF quantizations of llmfan46/Nex-N2-mini-ultra-uncensored-heretic.
This is a decensored version of a nex-agi/Nex-N2-mini, made using Heretic v1.2.0 with a variant of the Magnitude-Preserving Orthogonal Ablation (MPOA) method
Abliteration parameters
| Parameter | Value |
| :-------- | :---: |
| direction_index | 16.68 |
| attn.out_proj.max_weight | 1.11 |
| attn.out_proj.max_weight_position | 29.79 |
| attn.out_proj.min_weight | 0.83 |
| attn.out_proj.min_weight_distance | 26.98 |
| mlp.down_proj.max_weight | 1.94 |
| mlp.down_proj.max_weight_position | 29.92 |
| mlp.down_proj.min_weight | 1.84 |
| mlp.down_proj.min_weight_distance | 26.37 |
| attn.o_proj.max_weight | 1.65 |
| attn.o_proj.max_weight_position | 29.28 |
| attn.o_proj.min_weight | 1.36 |
| attn.o_proj.min_weight_distance | 23.64 |
Targeted components
* attn.o_proj
* attn.out_proj
* mlp.down_proj
Performance
| Metric | This model | Original model (Nex-N2-mini) |
| :----- | :--------: | :---------------------------: |
| KL divergence | <span style="color:darkgoldenrod">0.0020</span> | 0 (by definition) |
| Refusals | β <span style="color:darkgreen">5/100</span> | β <span style="color:blue">74/100</span> |
Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections.
-----
Quantizations
For the K-quants below, small SSM tensors are kept at higher precision where useful.
-Q6_K keeps ssm_alpha, ssm_beta, and ssm_out as Q8_0.
-Q5_K, Q4_K, and Q3_K quants keep ssm_alpha and ssm_beta as Q8_0, while ssm_out is kept as Q6_K.
This helps preserve the hybrid/SSM blocks with a small file-size increase.
| Filename | Quant | Description |
|----------|-------|-------------|
| Nex-N2-mini-ultra-uncensored-heretic-BF16.gguf | BF16 | Full precision |
| Nex-N2-mini-ultra-uncensored-heretic-Q8_0.gguf | Q8_0 | Near-lossless, recommended |
| Nex-N2-mini-ultra-uncensored-heretic-Q6_K.gguf | Q6_K | Excellent quality |
| Nex-N2-mini-ultra-uncensored-heretic-Q5_K_M.gguf | Q5_K_M | Good balance |
| Nex-N2-mini-ultra-uncensored-heretic-Q5_K_S.gguf | Q5_K_S | Smaller Q5 |
| Nex-N2-mini-ultra-uncensored-heretic-Q4_K_M.gguf | Q4_K_M | Good for limited VRAM |
| Nex-N2-mini-ultra-uncensored-heretic-Q4_K_S.gguf | Q4_K_S | Smaller Q4 |
| Nex-N2-mini-ultra-uncensored-heretic-Q3_K_L.gguf | Q3_K_L | Low VRAM, decent quality
| Nex-N2-mini-ultra-uncensored-heretic-Q3_K_M.gguf | Q3_K_M | Low VRAM, smaller |
Vision Projector
| Filename | Quant | Description |
|----------|-------|-------------|
| Nex-N2-mini-ultra-uncensored-heretic-mmproj-BF16.gguf | BF16 | Native precision |
A Vision Projector File is Required for vision/multimodal capabilities. Use alongside any quantization above.
Usage
Works with llama.cpp, LM Studio, Ollama, and other GGUF-compatible tools.
-----
<div align="center">
<img src="./figures/NEX_logo.svg" width="20%"/>
</div>
---
<div align="center">
π€ <a href="https://hf.co/collections/nex-agi/nex-n2"><b>Model</b></a>   |   
π <a href="https://openrouter.ai/nex-agi/Nex-N2-Pro:free"><b>OpenRouter (Enjoy two weeks free starting June 9!)</b></a>   |   
π» <a href="https://github.com/nex-agi/Nex-N2"><b>Github</b></a>   |   
π§ <a href="https://www.modelscope.cn/collections/nex-agi/Nex-N2"><b>ModelScope</b></a>   |   
π <a href="https://nex-agi.com"><b>Nex-AGI</b></a>
</div>
Nex-N2
An agentic model with Agentic Thinking.
Today, we are officially releasing and open-sourcing our next-generation model, Nex-N2 β an agent model built for real-world productivity scenarios. With first-tier coding and agentic capabilities, Nex-N2 keeps driving complex, long-horizon tasks forward in real environments to deliver stable, end-to-end results.
Over the past year, a paradigm shift led by Vibe Coding and Harness Engineering has been redefining the limits of LLM agents. From dialogue, to reasoning, to agents that execute long-horizon tasks with environmental feedback, the tasks models must handle keep growing harder, the contexts longer, and the environments more realistic. The core of next-generation model competition is no longer whether a model can think, but whether it can reliably and efficiently turn thinking into actions that are executable, verifiable, and iterable.
Rather than treating reasoning, tool use, and environment execution as separate capabilities, Nex-N2 unifies them through an Agentic Thinking framework that connects requirement understanding, task planning, code implementation, environmental feedback, evaluation and debugging, and continuous iteration into a single closed loop. The framework has two parts:
- Adaptive Thinking lets the model decide on its own when to think and how deeply β executing simple actions quickly while reasoning thoroughly on critical decisions.
- Coherent Thinking carries one consistent reasoning paradigm across general reasoning and diverse agentic tasks, staying consistent across tasks and modalities to enable stable capability transfer.
Across real agentic workflows β agentic coding, deep research, tool calling, and terminal execution β Nex-N2 reaches first-tier performance, with substantial gains over the previous-generation Nex-N1 on multiple authoritative benchmarks. In real productivity scenarios such as OpenClaw one-person-company workflows, end-to-end game development, and web and multimodal generation, it likewise demonstrates outstanding usability, robustness, and stability.
Open Source
In keeping with our commitment to open source, we are releasing both Nex-N2-Pro and Nex-N2-mini as open-source models starting today.
- Nex-N2-Pro: Hugging Face | ModelScope
- Nex-N2-mini: Hugging Face | ModelScope
- Early Access: SiliconFlow
We welcome developers and enterprises to integrate and try Nex-N2 and share their feedback.
Performance
We evaluate Nex-N2 in real agentic workflows along three directions β agentic tasks, coding tasks, and general tasks β covering benchmarks across tool calling, search-based decision-making, software engineering, and terminal execution. Nex-N2-Pro delivers strong performance that keeps pace with top-tier models such as GPT-5.5 and Opus 4.7: it excels at coding (e.g., 75.3 on Terminal-Bench 2.1) and long-horizon tasks (1585 on GDPval), and shows especially strong generalization and competitiveness on newer benchmarks like SWE-Atlas and DeepSWE. On general capability and core reasoning, it stands on par with leading frontier models.
Nex-N2 ships in two variants, both post-trained on the Qwen3.5 series: Nex-N2-Pro (built on Qwen3.5-397B-A17B) and Nex-N2-mini (built on Qwen3.5-35B-A3B-Base), covering different latency and quality trade-offs. The table below reports their scores alongside leading proprietary and open models across our full evaluation suite.
| Benchmark | Nex-N2-mini | Nex-N2-Pro | GPT-5.5 | Opus 4.7 | Kimi-K2.6 | GLM-5.1 | MiniMax M3 | DeepSeek-V4-Pro |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Agent | | | | | | | | |
| BrowseComp | 74.1 | 83.7 | 84.4 | 79.8 | 83.2 | 79.3 | 83.5 | 83.4 |
| GDPval | 1402 | 1585 | 1769 | 1753 | 1481 | 1535 | - | 1554 |
| Toolathlon | 33.3 | 51.9 | 55.6 | 52.8 | 50.0 | 40.7 | - | 51.8 |
| WildClawBench | 47.7 | 53.5 | 58.2 | 62.2 | - | 48.2 | - | 43.7 |
| WideSearch | 62.0 | 75.6 | - | - | 80.8 | - | - | - |
| TAU3 | 65.9 | 71.1 | - | - | - | 70.6 | - | - |
| Coding & SWE | | | | | | | | |
| SWE-Bench Pro | 50.2 | 58.8 | 58.6 | 64.3 | 58.6 | 58.4 | 59.0 | 55.4 |
| Terminal-Bench 2.1 | 60.7 | 75.3 | 83.4 | 69.7 | - | 58.7 | 66.0 | 72.0 |
| DeepSWE | 8.0 | 33.6 | 70 | 54 | 24 | 18 | - | 8 |
| SWE-Bench Verified | 74.4 | 80.8 | 82.9 | 87.6 | 80.2 | - | 80.5 | 80.6 |
| SWE Atlas QnA | 31.5 | 37.9 | 45.4 | 45.2 | - | - | 37.9 | - |
| SWE Atlas RF | 30.0 | 32.9 | 44.8 | 48.6 | - | - | - | - |
| SWE Atlas TW | 23.3 | 40.0 | 42.6 | 38.2 | - | - | 30.8 | - |
| General & Reasoning | | | | | | | | |
| GPQA Diamond | 82.6 | 90.7 | 93.6 | 94.2 | 90.5 | 86.2 | - | 90.1 |
| IFEval | 89.1 | 94.0 | - | - | 94.5 | 94.5 | - | 91.9 |
| Apex | 9.4 | 36.5 | - | - | 24.0 | 11.5 | - | 38.3 |
Usage
Local Deployment
> Note: For the best performance with Nex-series models, we recommend serving them with our customized sglang fork.
First, install our sglang fork:
# Use the customized `sglang` fork
git clone https://github.com/nex-agi/sglang.git
cd sglang
# Install the python packages
pip install --upgrade pip
pip install -e "python"
Nex-N2-Pro
Launch the server (example on two 8Γ H100 servers with CUDA 13.0):
# Multi-node (2 nodes). Run the same command on every node with:
# <node-rank> = 0 on the head node, 1 on the other node
# <node0-ip> = IP of the head node (reachable from all others)
python -m sglang.launch_server \
--model-path /path/to/your/model \
--tp 16 \
--nnodes 2 \
--node-rank <node-rank> \
--dist-init-addr <node0-ip>:20000 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-scheduler-strategy extra_buffer
Nex-N2-mini
Launch the server (example on one 2Γ H100 server with CUDA 13.0):
python -m sglang.launch_server \
--model-path /path/to/your/model \
--tp 2 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-scheduler-strategy extra_buffer
Docker Deployment
We also provide a prebuilt Docker image with our customized sglang fork preinstalled: nexagi/sglang:v0.5.12. The launch command is the same as above.
Nex-N2-Pro
# Multi-node (2 nodes). Run the same command on every node with:
# <node-rank> = 0 on the head node, 1 on the other node
# <node0-ip> = IP of the head node (reachable from all others)
docker run --gpus all --shm-size 32g --network host \
-v /path/to/your/model:/model \
nexagi/sglang:v0.5.12 \
python3 -m sglang.launch_server \
--model-path /model \
--tp 16 \
--nnodes 2 \
--node-rank <node-rank> \
--dist-init-addr <node0-ip>:20000 \
--host 0.0.0.0 --port 30000 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-scheduler-strategy extra_buffer
Nex-N2-mini
Single node with 2Γ H100:
docker run --gpus all --shm-size 32g --ipc=host \
-p 30000:30000 \
-v /path/to/your/model:/model \
nexagi/sglang:v0.5.12 \
python3 -m sglang.launch_server \
--model-path /model \
--tp 2 \
--host 0.0.0.0 --port 30000 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-scheduler-strategy extra_buffer
Recommended Sampling Parameters
For the best generation quality, we recommend the following sampling parameters:
temperature: 0.7top_p: 0.95top_k: 40
Function Calling
Nex-series models support robust function-calling capabilities. To enable function calling, add the --tool-call-parser qwen3_coder flag when launching the server:
python -m sglang.launch_server --model-path /path/to/your/model --tool-call-parser qwen3_coder
Reasoning Parser
Nex-series models emit explicit reasoning traces. Add the --reasoning-parser qwen3 flag to parse the reasoning content separately from the final response. It can be combined with the function-calling parser above:
python -m sglang.launch_server --model-path /path/to/your/model --tool-call-parser qwen3_coder --reasoning-parser qwen3Run llmfan46/Nex-N2-mini-ultra-uncensored-heretic-GGUF with guIDE
Download guIDE β the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face Β· Compare models