IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-GGUF overview
XYZ Aquila mini APEX I MiniPlus GGUF The Definitive 35B Multimodal Search & UI Agent MoE Β· Active V2 Release Available IMPORTANT π Official V2 Release Availabβ¦
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Browse files on Hugging Face | ||||
Model Details
| Model ID | IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-GGUF |
|---|---|
| Author | IsValorum |
| Pipeline | image-text-to-text |
| License | apache-2.0 |
| Base model | XYZAILab/XYZ-Aquila-mini |
| Last modified | 2026-09-17T01:39:36.000Z |
Model README
---
base_model: XYZAILab/XYZ-Aquila-mini
library_name: gguf
tags:
- gguf
- apex
- custom-quantization
- unsloth-studio
- moe
- multimodal
- vision
- agentic-search
- search-agent
- computer-use
- llama.cpp
- qwen35moe
license: apache-2.0
language:
- en
- zh
pipeline_tag: image-text-to-text
---
XYZ-Aquila-mini APEX-I-MiniPlus GGUF
The Definitive 35B Multimodal Search & UI Agent MoE Β· Active V2 Release Available
> [!IMPORTANT]
> ### π Official V2 Release Available
> The official upgraded release with non-linear IQ codebooks, F32 router selectors, and bundled Q8_0 vision projector is live at:
> π IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-V2-GGUF
---
<a id="quick-navigation"></a>β‘ Quick Navigation Index
- π¦ Model Files & Specifications
- π¬ Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)
- ποΈ Bundled Q8_0 High-Precision Multimodal Vision Projector
- π» Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)
- π₯ The 24GB Miracle: Full 256K Context Runs In VRAM!
- ποΈ Hardware Throughput Projections (RTX 30 / 40 / 50)
- π οΈ Surgical Tensor Quantization Map
- π Recommended Configuration & Setup
---
<a id="model-specifications"></a>
π¦ Model Files & Specifications
| File Name | File Size | Memory Footprint | Format / Precision | Purpose |
| :--- | :--- | :--- | :--- | :--- |
| XYZ-Aquila-mini.APEX-I-MiniPlus-V2.gguf | 14.63 GB (13.63 GiB) | 13.63 GiB | Custom APEX-I (3.38 BPW) | Main agentic search, browser reasoning & logic core |
| mmproj-XYZAILab_XYZ-Aquila-mini-Q8_0.gguf| 610 MB (582 MiB) | 582 MiB | High-Precision Q8_0 Projector | Required for browser viewport inspection, UI clicks & OCR |
- Base Architecture:
Qwen3_5MoeForConditionalGeneration(40 layers, 256 fine-grained micro-experts with intermediate dimension 512, 8 active per token) + Vision Projector. - Active Parameters: approx. 3.2B active parameters per token.
---
<a id="comparative-analysis"></a>
π¬ Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)
Also, don't confuse APEX-I-MiniPlus-V2 with a generic baseline APEX-I-Mini. Traditional APEX-I-Mini drops core experts aggressively to 2-bit IQ2_S and leaves output.weight at 3-bit Q3_K_M, which creates a noticeable perplexity hit on complex reasoning tasks. V2 was specifically re-engineered to avoid that quality floor (keeping core experts at calibrated IQ3_XXS, output in Q6_K, shared expert in non-linear IQ4_NL, and routers in F32).
To put the numbers in perspective: this cuts nearly 2 GB off a flat 3-bit quant (approx. 15.6 GB), and weighs only about approx. 1 GB more than a generic APEX-I-Mini (approx. 12.5 GB). For that single extra gigabyte of VRAM, you get a massive jump in reasoning and syntactic stability.
Take a look at the tensor-by-tensor comparison table below to inspect the exact architectural differences and see why this specific allocation is optimal. That's specifically what this was built for:
| Architectural Component | Generic Automated Quants (Flat Q3_K_S / IQ3_S) | Generic APEX-I-Mini (Baseline Recipe) | Our Handcrafted APEX-I-MiniPlus-V2 (IsValorum) | Perceived Quality & Real-World Impact |
| :--- | :--- | :--- | :--- | :--- |
| Output Head (output.weight) | Flat IQ3_S / Q3_K_S (approx. 3.44 BPW) | Inherits base type Q3_K_M (approx. 3.44 BPW unarmored) | Q6_K (approx. 6.56 BPW uncompromised) | Eliminates Syntax & Vocabulary Hallucinations: Low-bit output heads cause tokenizer classification noise, breaking code indentation, brackets ({}, []), math symbols, and domain terms. Q6_K preserves near-FP16 output classification. |
| Expert Routers (ffn_gate_inp.weight) | Blindly quantized to 3-bit / unoptimized | Inherits base type Q3_K_M (approx. 3.44 BPW compressed) | F32 uncompressed (32.0 BPW, 2 MB/layer) | Zero Router Drift: In micro-expert models, even minuscule quantization errors in router logits misdirect tokens to wrong experts. Retaining uncompressed F32 guarantees 100% routing fidelity with virtually zero memory overhead (approx. 80 MB total). |
| Attention & Language (attn_output, attn_qkv) | Flat IQ3_S / Q3_K_S | Q3_K on 34 middle layers (L3β36), Q4_K on 6 edge layers | Q6_K for attn_output, IQ3_S for attn_qkv | Contextual Retrieval Precision: Generic APEX reduces attention and language projections to Q3_K across 85% of layers. Our V2 build protects attention output in high-precision Q6_K and uses calibrated non-linear IQ3_S, ensuring flawless needle-in-a-haystack retrieval across deep 128kβ256k context windows. |
| Attention Gates (attn_gate.weight) | Blindly compressed to 3-bit | Compressed to Q3_K (middle) / Q4_K (edges) | Q8_0 (8.50 BPW) | Attention Head Stability: Attention gates modulate query-key routing across hybrid attention layers. Keeping them in 8-bit prevents attention crosstalk and hallucination over long contexts. |
| *Shared Foundation Expert (ffn__shexp) | Flat IQ3_S / Q3_K_S (3.44 BPW) | Linear Q4_K (middle) / Q5_K (edges) | IQ4_NL (4.50 BPW non-linear codebook) | Foundational Knowledge Armor:** The shared expert executes for 100% of tokens. In 256 micro-expert models, IQ4_NL non-linear codebooks preserve heavy-tailed outlier representations far better than standard linear quantization. |
| Core MoE Layers (Middle: 10β29) | Flat IQ3_S / Q3_K_S (uniform bit-rate across all layers) | Aggressive IQ2_S (2.50 BPW) | IQ3_XXS (3.06 BPW) + calibrated imatrix | Above the Quality Threshold: Generic 2-bit IQ2_S baselines drop below the critical quality floor for 35B MoEs, resulting in perplexity spikes on reasoning tasks. Our IQ3_XXS with imatrix achieves deep compression (272 MiB β 98 MiB per block) without sacrificing logic. |
| Edge MoE Layers (Layers 0β9 & 30β39) | Flat IQ3_S / Q3_K_S (no layer-wise gradient) | Q3_K (limited to first/last 5 layers only: L0β4, L35β39) | IQ3_S (expanded to 10 input & 10 output layers) | Protected Ingestion & Synthesis: Half of the model's layers (10 at input, 10 at output) form a non-linear armored envelope, preventing prompt misunderstanding and token degeneration across 256 micro-experts. |
| Multimodal Vision (mmproj) | Often omitted, or left as uncompressed FP16 (approx. 900 MB) | Often omitted or separate uncompressed FP16 | Bundled Q8_0 (582 MB) with 27 critical F32/F16 fallbacks | Saves approx. 320 MB VRAM with Zero Loss: Handcrafted quantization preserves normalization and bias tensors in F32/F16, ensuring razor-sharp OCR, DOM viewport reading, and coordinate detection without visual noise. |
| Normalization & Biases | Often degraded | Standard | F32 uncompressed | Numerical Stability: Prevents cumulative floating-point underflow/overflow across deep 40-layer computation. |
---
<a id="vision-projector"></a>
ποΈ Bundled Q8_0 High-Precision Multimodal Vision Projector
- Bundled Q8_0 Projector: Pre-quantized to
Q8_0(582 MiB / 610 MB), saving approx. 300 MB of VRAM. - Audited Layer Fallbacks:
llama.cppautomatically preserved 27 critical normalization and bias tensors in F32/F16, ensuring razor-sharp rendering of browser DOM text, minute UI action targets, and dense infographic diagrams.
---
<a id="laptop-benchmarks"></a>
π» Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)
- GPU VRAM Allocation: Uses only approx. 3.8 GB VRAM (fits effortlessly on budget laptop GPUs).
- System Memory Offload: Standard 32GB system RAM accommodates the remaining layers.
- Estimated Document Ingestion (Prefill): 300 to 420+ tokens/second sustained across full viewport inputs.
- Estimated Streaming Generation: 20 to 24+ tokens/second sustained output across system RAM!
---
<a id="context-scaling"></a>
π₯ The 24GB Miracle: Full 256K Context Runs In VRAM!
| Context Length | Model Weights (Est.) | KV Cache (q8_0, 4 slots) | Compute Buffers | Total GPU VRAM (Est.) | Hardware Feasibility |
| :--- | :--- | :--- | :--- | :--- | :--- |
| 32,768 (32k) | 13.63 GiB | 0.58 GiB | 1.80 GiB | 16.01 GiB | Full offload on 24GB; partial on 16GB |
| 65,536 (64k) | 13.63 GiB | 0.92 GiB | 1.95 GiB | 16.50 GiB | Effortless fit on 24GB GPUs |
| 131,072 (128k)| 13.63 GiB | 1.58 GiB | 2.22 GiB | 17.43 GiB | Effortless fit on 24GB GPUs |
| 262,144 (256k)| 13.63 GiB | 2.92 GiB | 2.80 GiB | 19.35 GiB | π₯ FULL 256K AGENT TRACE IN VRAM! |
---
<a id="throughput-projections"></a>
ποΈ Hardware Throughput Projections (RTX 30 / 40 / 50)
| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
| :--- | :--- | :---: | :---: | : |
| NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) + mmproj | 105 β 130+ tok/s | 2,400 β 3,500+ tok/s | Blistering autonomous web search throughput |
| NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) + mmproj | 75 β 100+ tok/s | 1,700 β 2,500+ tok/s | Real-time browser DOM parsing & action generation |
| NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) + mmproj | 62 β 78+ tok/s | 1,350 β 1,950+ tok/s | Full 256k multi-turn web search in dedicated VRAM |
| Consumer Laptop (4GB GPU + 32GB RAM)| Hybrid Offload | 20 β 24+ tok/s | 300 β 420+ tok/s | Smooth streaming from system DDR4/DDR5 RAM |
---
<a id="tensor-map"></a>
π οΈ Surgical Tensor Quantization Map
| Tensor Pattern | Layer Scope | Quant Type | BPW | Engineering Rationale |
| :--- | :--- | :--- | :--- | :--- |
| output.weight | Vocabulary Head | Q6_K | 6.56 | Uncompromised 6-bit precision for web queries, structured JSON & tool syntax |
| token_embd.weight | Embedding | High-Prec | High | Preserves subtle token semantics and prompt grounding |
| ffn_gate_inp.weight | Expert Routers | F32 | 32.0 | Uncompressed full-precision routers preventing visual token misrouting |
| attn_gate.weight | Attention Gates | Q8_0 | 8.50 | High-precision 8-bit gating for attention routing dynamics |
| *ffn__shexp | Shared Experts | IQ4_NL** | 4.50 | 4-bit non-linear codebook for the 100% active shared foundational expert |
| ffn_down/up/gate | Edges (0β9, 30β39) | IQ3_S | 3.44 | Armored boundary layers protecting prompt ingest and final UI action synthesis |
| ffn_down/up/gate | Core (10β29) | IQ3_XXS | 3.06 | Deep compression (272 MiB β 98 MiB per block) calibrated via multimodal imatrix |
| mmproj (Vision) | Visual Projector | Q8_0 | 8.00 | High-fidelity OCR and UI coordinate rendering with 27 critical F32/F16 fallbacks |
| Norms & Biases | All Layers | F32 | 32.0 | Absolute numerical stability across deep 40-layer computation |
---
<a id="recommended-setup"></a>
π Recommended Configuration & Setup
See the primary repository for complete configuration and download links:
Run IsValorum/XYZ-Aquila-mini-APEX-I-MiniPlus-GGUF with guIDE
Download guIDE β the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face Β· Compare models