GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE β†’
Model Intelligence Sheet

IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-GGUF overview

XYZ Aquila mini APEX I MiniPlus GGUF The Definitive 35B Multimodal Search & UI Agent MoE Β· Active V2 Release Available IMPORTANT 🌟 Official V2 Release Availab…

ggufapexcustom-quantizationunsloth-studiomoemultimodalvisionagentic-searchsearch-agentcomputer-usellama.cppqwen35moeimage-text-to-textenzhbase_model:XYZAILab/XYZ-Aquila-minibase_model:finetune:XYZAILab/XYZ-Aquila-minilicense:apache-2.0region:us
Downloads
0
Likes
0
Pipeline
image-text-to-text
Author

Repository Files & Downloads

0 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Browse files on Hugging Face

Model Details

Model IDIsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-GGUF
AuthorIsValorum
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelXYZAILab/XYZ-Aquila-mini
Last modified2026-09-17T01:39:36.000Z

Model README

---

base_model: XYZAILab/XYZ-Aquila-mini

library_name: gguf

tags:

- gguf

- apex

- custom-quantization

- unsloth-studio

- moe

- multimodal

- vision

- agentic-search

- search-agent

- computer-use

- llama.cpp

- qwen35moe

license: apache-2.0

language:

- en

- zh

pipeline_tag: image-text-to-text

---

XYZ-Aquila-mini APEX-I-MiniPlus GGUF

The Definitive 35B Multimodal Search & UI Agent MoE Β· Active V2 Release Available

> [!IMPORTANT]

> ### 🌟 Official V2 Release Available

> The official upgraded release with non-linear IQ codebooks, F32 router selectors, and bundled Q8_0 vision projector is live at:

> πŸ‘‰ IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-V2-GGUF

---

<a id="quick-navigation"></a>⚑ Quick Navigation Index

---

<a id="model-specifications"></a>

πŸ“¦ Model Files & Specifications

| File Name | File Size | Memory Footprint | Format / Precision | Purpose |

| :--- | :--- | :--- | :--- | :--- |

| XYZ-Aquila-mini.APEX-I-MiniPlus-V2.gguf | 14.63 GB (13.63 GiB) | 13.63 GiB | Custom APEX-I (3.38 BPW) | Main agentic search, browser reasoning & logic core |

| mmproj-XYZAILab_XYZ-Aquila-mini-Q8_0.gguf| 610 MB (582 MiB) | 582 MiB | High-Precision Q8_0 Projector | Required for browser viewport inspection, UI clicks & OCR |

  • Base Architecture: Qwen3_5MoeForConditionalGeneration (40 layers, 256 fine-grained micro-experts with intermediate dimension 512, 8 active per token) + Vision Projector.
  • Active Parameters: approx. 3.2B active parameters per token.

---

<a id="comparative-analysis"></a>

πŸ”¬ Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)

Also, don't confuse APEX-I-MiniPlus-V2 with a generic baseline APEX-I-Mini. Traditional APEX-I-Mini drops core experts aggressively to 2-bit IQ2_S and leaves output.weight at 3-bit Q3_K_M, which creates a noticeable perplexity hit on complex reasoning tasks. V2 was specifically re-engineered to avoid that quality floor (keeping core experts at calibrated IQ3_XXS, output in Q6_K, shared expert in non-linear IQ4_NL, and routers in F32).

To put the numbers in perspective: this cuts nearly 2 GB off a flat 3-bit quant (approx. 15.6 GB), and weighs only about approx. 1 GB more than a generic APEX-I-Mini (approx. 12.5 GB). For that single extra gigabyte of VRAM, you get a massive jump in reasoning and syntactic stability.

Take a look at the tensor-by-tensor comparison table below to inspect the exact architectural differences and see why this specific allocation is optimal. That's specifically what this was built for:

| Architectural Component | Generic Automated Quants (Flat Q3_K_S / IQ3_S) | Generic APEX-I-Mini (Baseline Recipe) | Our Handcrafted APEX-I-MiniPlus-V2 (IsValorum) | Perceived Quality & Real-World Impact |

| :--- | :--- | :--- | :--- | :--- |

| Output Head (output.weight) | Flat IQ3_S / Q3_K_S (approx. 3.44 BPW) | Inherits base type Q3_K_M (approx. 3.44 BPW unarmored) | Q6_K (approx. 6.56 BPW uncompromised) | Eliminates Syntax & Vocabulary Hallucinations: Low-bit output heads cause tokenizer classification noise, breaking code indentation, brackets ({}, []), math symbols, and domain terms. Q6_K preserves near-FP16 output classification. |

| Expert Routers (ffn_gate_inp.weight) | Blindly quantized to 3-bit / unoptimized | Inherits base type Q3_K_M (approx. 3.44 BPW compressed) | F32 uncompressed (32.0 BPW, 2 MB/layer) | Zero Router Drift: In micro-expert models, even minuscule quantization errors in router logits misdirect tokens to wrong experts. Retaining uncompressed F32 guarantees 100% routing fidelity with virtually zero memory overhead (approx. 80 MB total). |

| Attention & Language (attn_output, attn_qkv) | Flat IQ3_S / Q3_K_S | Q3_K on 34 middle layers (L3–36), Q4_K on 6 edge layers | Q6_K for attn_output, IQ3_S for attn_qkv | Contextual Retrieval Precision: Generic APEX reduces attention and language projections to Q3_K across 85% of layers. Our V2 build protects attention output in high-precision Q6_K and uses calibrated non-linear IQ3_S, ensuring flawless needle-in-a-haystack retrieval across deep 128k–256k context windows. |

| Attention Gates (attn_gate.weight) | Blindly compressed to 3-bit | Compressed to Q3_K (middle) / Q4_K (edges) | Q8_0 (8.50 BPW) | Attention Head Stability: Attention gates modulate query-key routing across hybrid attention layers. Keeping them in 8-bit prevents attention crosstalk and hallucination over long contexts. |

| *Shared Foundation Expert (ffn__shexp) | Flat IQ3_S / Q3_K_S (3.44 BPW) | Linear Q4_K (middle) / Q5_K (edges) | IQ4_NL (4.50 BPW non-linear codebook) | Foundational Knowledge Armor:** The shared expert executes for 100% of tokens. In 256 micro-expert models, IQ4_NL non-linear codebooks preserve heavy-tailed outlier representations far better than standard linear quantization. |

| Core MoE Layers (Middle: 10–29) | Flat IQ3_S / Q3_K_S (uniform bit-rate across all layers) | Aggressive IQ2_S (2.50 BPW) | IQ3_XXS (3.06 BPW) + calibrated imatrix | Above the Quality Threshold: Generic 2-bit IQ2_S baselines drop below the critical quality floor for 35B MoEs, resulting in perplexity spikes on reasoning tasks. Our IQ3_XXS with imatrix achieves deep compression (272 MiB β†’ 98 MiB per block) without sacrificing logic. |

| Edge MoE Layers (Layers 0–9 & 30–39) | Flat IQ3_S / Q3_K_S (no layer-wise gradient) | Q3_K (limited to first/last 5 layers only: L0–4, L35–39) | IQ3_S (expanded to 10 input & 10 output layers) | Protected Ingestion & Synthesis: Half of the model's layers (10 at input, 10 at output) form a non-linear armored envelope, preventing prompt misunderstanding and token degeneration across 256 micro-experts. |

| Multimodal Vision (mmproj) | Often omitted, or left as uncompressed FP16 (approx. 900 MB) | Often omitted or separate uncompressed FP16 | Bundled Q8_0 (582 MB) with 27 critical F32/F16 fallbacks | Saves approx. 320 MB VRAM with Zero Loss: Handcrafted quantization preserves normalization and bias tensors in F32/F16, ensuring razor-sharp OCR, DOM viewport reading, and coordinate detection without visual noise. |

| Normalization & Biases | Often degraded | Standard | F32 uncompressed | Numerical Stability: Prevents cumulative floating-point underflow/overflow across deep 40-layer computation. |

---

<a id="vision-projector"></a>

πŸ‘οΈ Bundled Q8_0 High-Precision Multimodal Vision Projector

  • Bundled Q8_0 Projector: Pre-quantized to Q8_0 (582 MiB / 610 MB), saving approx. 300 MB of VRAM.
  • Audited Layer Fallbacks: llama.cpp automatically preserved 27 critical normalization and bias tensors in F32/F16, ensuring razor-sharp rendering of browser DOM text, minute UI action targets, and dense infographic diagrams.

---

<a id="laptop-benchmarks"></a>

πŸ’» Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)

  • GPU VRAM Allocation: Uses only approx. 3.8 GB VRAM (fits effortlessly on budget laptop GPUs).
  • System Memory Offload: Standard 32GB system RAM accommodates the remaining layers.
  • Estimated Document Ingestion (Prefill): 300 to 420+ tokens/second sustained across full viewport inputs.
  • Estimated Streaming Generation: 20 to 24+ tokens/second sustained output across system RAM!

---

<a id="context-scaling"></a>

πŸ”₯ The 24GB Miracle: Full 256K Context Runs In VRAM!

| Context Length | Model Weights (Est.) | KV Cache (q8_0, 4 slots) | Compute Buffers | Total GPU VRAM (Est.) | Hardware Feasibility |

| :--- | :--- | :--- | :--- | :--- | :--- |

| 32,768 (32k) | 13.63 GiB | 0.58 GiB | 1.80 GiB | 16.01 GiB | Full offload on 24GB; partial on 16GB |

| 65,536 (64k) | 13.63 GiB | 0.92 GiB | 1.95 GiB | 16.50 GiB | Effortless fit on 24GB GPUs |

| 131,072 (128k)| 13.63 GiB | 1.58 GiB | 2.22 GiB | 17.43 GiB | Effortless fit on 24GB GPUs |

| 262,144 (256k)| 13.63 GiB | 2.92 GiB | 2.80 GiB | 19.35 GiB | πŸ”₯ FULL 256K AGENT TRACE IN VRAM! |

---

<a id="throughput-projections"></a>

🏎️ Hardware Throughput Projections (RTX 30 / 40 / 50)

| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |

| :--- | :--- | :---: | :---: | : |

| NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) + mmproj | 105 – 130+ tok/s | 2,400 – 3,500+ tok/s | Blistering autonomous web search throughput |

| NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) + mmproj | 75 – 100+ tok/s | 1,700 – 2,500+ tok/s | Real-time browser DOM parsing & action generation |

| NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) + mmproj | 62 – 78+ tok/s | 1,350 – 1,950+ tok/s | Full 256k multi-turn web search in dedicated VRAM |

| Consumer Laptop (4GB GPU + 32GB RAM)| Hybrid Offload | 20 – 24+ tok/s | 300 – 420+ tok/s | Smooth streaming from system DDR4/DDR5 RAM |

---

<a id="tensor-map"></a>

πŸ› οΈ Surgical Tensor Quantization Map

| Tensor Pattern | Layer Scope | Quant Type | BPW | Engineering Rationale |

| :--- | :--- | :--- | :--- | :--- |

| output.weight | Vocabulary Head | Q6_K | 6.56 | Uncompromised 6-bit precision for web queries, structured JSON & tool syntax |

| token_embd.weight | Embedding | High-Prec | High | Preserves subtle token semantics and prompt grounding |

| ffn_gate_inp.weight | Expert Routers | F32 | 32.0 | Uncompressed full-precision routers preventing visual token misrouting |

| attn_gate.weight | Attention Gates | Q8_0 | 8.50 | High-precision 8-bit gating for attention routing dynamics |

| *ffn__shexp | Shared Experts | IQ4_NL** | 4.50 | 4-bit non-linear codebook for the 100% active shared foundational expert |

| ffn_down/up/gate | Edges (0–9, 30–39) | IQ3_S | 3.44 | Armored boundary layers protecting prompt ingest and final UI action synthesis |

| ffn_down/up/gate | Core (10–29) | IQ3_XXS | 3.06 | Deep compression (272 MiB β†’ 98 MiB per block) calibrated via multimodal imatrix |

| mmproj (Vision) | Visual Projector | Q8_0 | 8.00 | High-fidelity OCR and UI coordinate rendering with 27 critical F32/F16 fallbacks |

| Norms & Biases | All Layers | F32 | 32.0 | Absolute numerical stability across deep 40-layer computation |

---

<a id="recommended-setup"></a>

πŸ“– Recommended Configuration & Setup

See the primary repository for complete configuration and download links:

πŸ‘‰ IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-V2-GGUF

Run IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-GGUF with guIDE

Download guIDE β€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE β†’ Β· Browse 524k+ models Β· Compare models

Source: Hugging Face Β· Compare models