GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

zsyjsld/Xinghe1-9B-GGUF overview

Xinghe1 9B GGUF 杏核1代 9B GGUF Xinghe1 9B GGUF 杏核 contains the quantized GGUF versions of the Xinghe1 9B model, which is fine tuned from the Qwen3.5 9B Instruct …

transformersgguftext-generation-inferenceqwentcmmedicallicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~5.24 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Xinghe1-9B-Q4_K_M.ggufGGUFQ4_K_M5.24 GBDownload
Xinghe1-9B-Q6_K.ggufGGUFQ6_K6.85 GBDownload
Xinghe1-9B-Q8_0.ggufGGUFQ8_08.87 GBDownload

Model Details

Model IDzsyjsld/Xinghe1-9B-GGUF
Authorzsyjsld
Pipeline
Licenseapache-2.0
Base modelQwen/Qwen3.5-9B-Instruct
Last modified2026-07-01T01:43:43.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen3.5-9B-Instruct

tags:

  • text-generation-inference
  • transformers
  • qwen
  • tcm
  • medical
  • gguf

---

Xinghe1-9B-GGUF (杏核1代-9B GGUF)

Xinghe1-9B-GGUF (杏核) contains the quantized GGUF versions of the Xinghe1-9B model, which is fine-tuned from the Qwen3.5-9B-Instruct architecture for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing.

This repository provides multiple quantization formats compiled using llama-quantize for efficient local deployment and inference.

Released Quantizations

  • Xinghe1-9B-BF16.gguf (16.69 GB) — Native best-quality Bfloat16 precision, recommended for GPU setups.
  • Xinghe1-9B-Q8_0.gguf (8.87 GB) — 8-bit high-quality quantization, minimal quality loss.
  • Xinghe1-9B-Q6_K.gguf (6.85 GB) — 6-bit balanced quantization, recommended for most setups.
  • Xinghe1-9B-Q4_K_M.gguf (5.24 GB) — 4-bit fast and lightweight quantization.

Note: All GGUF models in this repository have been fixed to exclude the Qwen3.5 MTP (Multi-Token Prediction) layers, avoiding metadata layer count conflicts. They can be directly loaded into LM Studio without errors.

Deployment & Usage

1. LM Studio

Search for zsyjsld/Xinghe1-9B-GGUF directly in the LM Studio search bar, select your preferred quantization version, and load the model.

2. Ollama

Run the 4-bit version using the following command:

ollama run hf.co/zsyjsld/Xinghe1-9B-GGUF:Q4_K_M

Model Details

  • Developed by: zsyjsld
  • Base Model: Qwen/Qwen3.5-9B-Instruct
  • Fine-tuning Method: QLoRA SFT on V3 Double-Purity dataset.
  • Language(s): English & LaTeX (Internal reasoning formulas), English (Output explanations)
  • License: Apache 2.0 (for code/dataset) & Qwen License Agreement (for model weights)

Technical Architecture

Xinghe1-9B is trained to strictly bind clinical concepts and perform dynamical derivations inside the <think> tag, and output natural, clean clinical TCM analyses outside the <think> tag.

The internal derivation is based on the unified meta-framework:

$$\dot{x} = F(x) + G(x, \mu(t))$$

$$y = h(x) + \epsilon$$

Run zsyjsld/Xinghe1-9B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models