GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

jamesatron1512/Qwen3.8-27B-GGUF overview

Qwen3.8 27B NVFP4 with Dynamic Subspace Engine DSE Dynamic Subspace Engine Performance dse live performance.png This repository contains the production Dynamic…

dynamic-subspace-engineqwen3.8nvfp44bitcompressed-tensorslow-vramwddm-shared-memorytext-generationenzhlicense:apache-2.0region:us
Downloads
0
Likes
1
Pipeline
text-generation

Repository Files & Downloads

0 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Browse files on Hugging Face

Model Details

Model IDjamesatron1512/Qwen3.8-27B-GGUF
Authorjamesatron1512
Pipelinetext-generation
Licenseapache-2.0
Base model
Last modified2026-08-14T18:34:10.000Z

Model README

---

language:

  • en
  • zh

license: apache-2.0

tags:

  • dynamic-subspace-engine
  • qwen3.8
  • nvfp4
  • 4bit
  • compressed-tensors
  • low-vram
  • wddm-shared-memory
  • text-generation

pipeline_tag: text-generation

inference: false

---

Qwen3.8-27B-NVFP4 with Dynamic Subspace Engine (DSE)

!Dynamic Subspace Engine Performance

This repository contains the production Dynamic Subspace Engine (DSE) runner and configuration for Qwen3.8-27B-NVFP4 (4-bit Compressed Tensors), enabling full 27-Billion parameter reasoning and code generation on consumer GPUs such as the NVIDIA GeForce RTX 5070 (12GB GDDR7).

---

⚡ Key Highlights

  • Hardware Footprint: Runs the 27B model across 12GB VRAM + 32GB System RAM using WDDM Shared Memory & PCIe direct streaming.
  • Upfront Subspace Predictor ($P_{\text{pred}}$): Prunes 95%–98% of SwiGLU MLP computations on the fly without loss of reasoning coherence.
  • Continuous Subspace Softmax ($P_{\text{sub}}$): Filters the 248,077 subword vocabulary down to candidate manifolds, eliminating 910× output logit memory traffic.
  • Native Tokenizer: Full support for Qwen3.8 BPE vocabulary (248k tokens) and ChatML formatting (<|im_start|>).
  • Dual API Server: Built-in OpenAI (/v1/chat/completions) and Ollama (/api/generate, /api/chat) REST server.

---

🚀 Quickstart

1. Requirements

pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
pip install transformers compressed-tensors accelerate fastapi uvicorn matplotlib

2. Interactive Terminal Chat

python run_dse_engine.py --interactive

3. Run a One-Off Prompt

python run_dse_engine.py --prompt "Write a Python script to calculate the golden ratio." --sparsity 0.95

4. Launch the Dual OpenAI & Ollama REST API Server

python run_dse_engine.py --serve --port 8000

---

📄 Technical Reference

For the complete mathematical formulation, error bound derivations, and Single-CCD L3 cache residency theorems, see DYNAMIC_SUBSPACES_RESEARCH_PAPER.md.

Run jamesatron1512/Qwen3.8-27B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models