GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

poisonxa/PXA-Fusion-122B-A5B-GGUF overview

PXA Fusion 122B A5B banner.png PXA Fusion 122B A5B GGUF 122B MoE A5B, top 4 active · qwen35moe · MTP speculative decode · 256K context · vanilla PXQ quants for…

ggufendpoints_compatibleregion:usconversational

Runs locally from ~22.12 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
30
Likes
1
Pipeline
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
PXA-Fusion-122B-A5B-PXQ2.ggufGGUFGGUF34.59 GBDownload
PXA-Fusion-122B-A5B-PXQ4.ggufGGUFGGUF62.15 GBDownload
PXA-Fusion-122B-A5B-PXQ6.ggufGGUFGGUF65.60 GBDownload
PXA-Fusion-122B-A5B-PXQU24.ggufGGUFGGUF22.12 GBDownload
PXA-Fusion-122B-A5B-PXQU32.ggufGGUFGGUF30.09 GBDownload
PXA-Fusion-122B-A5B-PXQU48.ggufGGUFGGUF45.27 GBDownload

Model Details

Model IDpoisonxa/PXA-Fusion-122B-A5B-GGUF
Authorpoisonxa
Pipeline
License
Base model
Last modified2026-07-25T01:48:10.000Z

Model README

!PXA-Fusion-122B-A5B

PXA-Fusion-122B-A5B-GGUF

122B MoE (A5B, top-4 active) · qwen35moe · MTP speculative decode · 256K context · vanilla PXQ quants for salvaged Pascal / Volta silicon.

Runs on old cheap datacenter cards (P100 / V100) via the PXA pxq_llama fork.

> ## Re-download if you pulled before 2026-07-25 - chat template fixed

>

> All six quants shipped a chat template with an inverted enable_thinking default. Unless your client

> explicitly sent enable_thinking:false, the template pre-filled an unclosed <think> tag, so generation

> began inside an unbounded reasoning channel. In agentic and coding harnesses this looks like the model

> planning forever - long first-person prose, no tool calls, no files written.

>

> Fixed in place on 2026-07-25. The patch rewrites only the template metadata; the weights are byte-identical

> (every file was sha256-verified against the published copy before patching, and the patch preserves byte

> length). Re-pull just the .gguf you use.

>

> Note for llama.cpp/pxq_llama users: that engine injects enable_thinking into the template context

> unconditionally, so it was never affected by the polarity. This fix matters for HF transformers, Ollama and

> vLLM, which leave the variable undefined. To turn reasoning off on the server, use --reasoning off;

> --reasoning-budget 0 does not work for this template.

Quant tiers

| Tier | Notes |

|---|---|

| PXQU24 / PXQU32 / PXQU48 | universal mixed-tier PXQ (knapsack-optimized) |

| PXQ2 / PXQ4 / PXQ6 | flat PXQ ladder (PXQ6 = real 5-bit) |

_All tiers carry grafted MTP heads. Vanilla (no imatrix)._

Speed & benchmarks

🔧 Tests in progress — SWE-bench Lite + 179-Q reasoning gauntlet + per-tier decode t/s landing soon. This card is a placeholder until those numbers are in.

Run poisonxa/PXA-Fusion-122B-A5B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models