GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

johnbean393/fluid-2-qwen3.5-4b-beta-GGUF overview

Fluid 2 Qwen3.5 4B Beta — GGUF Public GGUF conversions of the private step 344 Fluid 2 beta checkpoint. | File | Quantization | Size | SHA 256 | | | :| :| | | …

llama.cppggufqwen3.5fluid-2text-generationbase_model:johnbean393/fluid-2-qwen3.5-4b-betabase_model:quantized:johnbean393/fluid-2-qwen3.5-4b-betaendpoints_compatibleregion:usconversational

Runs locally from ~2.52 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
fluid-2-qwen3.5-4b-beta-Q4_K_M.ggufGGUFQ4_K_M2.52 GBDownload
fluid-2-qwen3.5-4b-beta-Q6_K.ggufGGUFQ6_K3.23 GBDownload
fluid-2-qwen3.5-4b-beta-Q8_0.ggufGGUFQ8_04.17 GBDownload

Model Details

Model IDjohnbean393/fluid-2-qwen3.5-4b-beta-GGUF
Authorjohnbean393
Pipelinetext-generation
License
Base modeljohnbean393/fluid-2-qwen3.5-4b-beta
Last modified2026-08-13T20:49:02.000Z

Model README

---

base_model: johnbean393/fluid-2-qwen3.5-4b-beta

library_name: llama.cpp

pipeline_tag: text-generation

tags:

  • gguf
  • qwen3.5
  • fluid-2

---

Fluid 2 Qwen3.5 4B Beta — GGUF

Public GGUF conversions of the private step-344 Fluid 2 beta

checkpoint.

| File | Quantization | Size | SHA-256 |

|---|---:|---:|---|

| fluid-2-qwen3.5-4b-beta-Q4_K_M.gguf | Q4_K_M | 2,708,804,640 bytes | 5089fd6907ba0cf84f44eb9b749f99e63b1ab6aa31e910bab7450d52fd37e96c |

| fluid-2-qwen3.5-4b-beta-Q6_K.gguf | Q6_K | 3,464,055,840 bytes | f2c1073232acf39de25283d0d8b8a9311e316972bcf033236f8b6520596d7e95 |

| fluid-2-qwen3.5-4b-beta-Q8_0.gguf | Q8_0 | 4,482,403,360 bytes | 39ed2353abf513814ddf90d48e16884574cbca2698e769c1c6d9bae039fcd5a3 |

Corrected development-set evaluation

| Metric | Result |

|---|---:|

| Scored text rows | 7,024 |

| Exact match | 33.5849% |

| CER | 16.4947% |

| WER | 25.4750% |

| Excluded EOS-only rows | 121 |

| Excluded generation-capped rows | 16 |

EM, CER, and WER exclude both empty-target/EOS-only rows and non-empty-target

generations that reached the configured token cap. Capped requests remain in

the failure and throughput census. This public 7,161-row development set was

used during training and is not a blind-test result.

Prompt template

This is a dictation-cleaning completion model, not a conversational assistant.

Do not apply a chat template. Input must end immediately after

<|start_target_text|>; generation stops at <|end_target_text|>.

<|dictation_clean_v1|>
<|start_prev_text|>{previous context}<|end_prev_text|>
<|start_post_text|>{following context}<|end_post_text|>
<|start_asr_text|>{ASR transcript to clean}<|end_asr_text|>
<|start_target_text|>

Previous and following context may be empty, but keep all marker pairs.

Run with llama.cpp

Use a recent llama.cpp llama-completion binary. This example selects

Q4_K_M and uses greedy decoding:

PROMPT='<|dictation_clean_v1|>
<|start_prev_text|><|end_prev_text|>
<|start_post_text|><|end_post_text|>
<|start_asr_text|>hello world<|end_asr_text|>
<|start_target_text|>'

./llama-completion \
  --hf-repo johnbean393/fluid-2-qwen3.5-4b-beta-GGUF:Q4_K_M \
  --prompt "$PROMPT" \
  --predict 256 \
  --temperature 0 \
  --ctx-size 8192 \
  --no-conversation \
  --no-display-prompt

Use :Q6_K or :Q8_0 for another quant. For a local file, replace

--hf-repo ... with --model ./fluid-2-qwen3.5-4b-beta-Q6_K.gguf.

Add --special while debugging to display the terminal control token.

The files were converted with the matching b10411 converter and

quantized with the official pre-built Ubuntu x64 llama.cpp

b10411 release. No CUDA/source build was performed. Every quant

passed GGUF metadata validation and a load/generation smoke test with that

pre-built binary. Exact hashes are recorded in conversion_manifest.json.

MTP / NextN note

The source configuration declares one MTP (multi-token prediction), also known

as NextN, speculative draft layer. The causal-LM checkpoint itself contains

the 32 trained decoder layers and **does not contain any MTP/NextN

draft-layer tensors**. Default conversion would therefore advertise a

nonexistent extra block and fail when the runtime requests that tensor.

These GGUFs intentionally use --no-mtp. This omits only the absent optional

speculative draft layer; it does not remove trained decoder weights and does

not change ordinary next-token generation. The files correctly declare

32 blocks.

Run johnbean393/fluid-2-qwen3.5-4b-beta-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models