GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-GGUF overview

Qwen3.6 27B Omnimerge v4 GGUF GGUF quantizations of ManniX ITA/Qwen3.6 27B Omnimerge v4 https://huggingface.co/ManniX ITA/Qwen3.6 27B Omnimerge v4 — the MLP pa…

ggufimatrixquantizedmergemergekitqwen3_5reasoningcodeimage-text-to-textenbase_model:ManniX-ITA/Qwen3.6-27B-Omnimerge-v4base_model:quantized:ManniX-ITA/Qwen3.6-27B-Omnimerge-v4license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~884.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
3,002
Likes
36
Pipeline
image-text-to-text

Repository Files & Downloads

28 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.6-27B-Omnimerge-v4-F16.ggufGGUFF1650.11 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ2_M.ggufGGUFIQ2_M9.32 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ2_S.ggufGGUFIQ2_S8.72 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ2_XS.ggufGGUFIQ2_XS8.47 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ2_XXS.ggufGGUFIQ2_XXS7.85 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ3_M.ggufGGUFIQ3_M11.72 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ3_XS.ggufGGUFIQ3_XS11.15 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ3_XXS.ggufGGUFIQ3_XXS10.42 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ4_NL.ggufGGUFIQ4_NL14.72 GBDownload
Qwen3.6-27B-Omnimerge-v4-IQ4_XS.ggufGGUFIQ4_XS14.05 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q2_K.ggufGGUFQ2_K9.98 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q2_K_L.ggufGGUFQ2_K_L11.13 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q3_K_L.ggufGGUFQ3_K_L13.36 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q3_K_M.ggufGGUFQ3_K_M12.39 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q3_K_S.ggufGGUFQ3_K_S11.24 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q3_K_XL.ggufGGUFQ3_K_XL13.42 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q4_0.ggufGGUFQ4_014.41 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q4_1.ggufGGUFQ4_115.91 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q4_K_L.ggufGGUFQ4_K_L16.29 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q4_K_M.ggufGGUFQ4_K_M15.41 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q4_K_S.ggufGGUFQ4_K_S14.52 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q5_K_L.ggufGGUFQ5_K_L18.64 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q5_K_M.ggufGGUFQ5_K_M17.91 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q5_K_S.ggufGGUFQ5_K_S17.40 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q6_K.ggufGGUFQ6_K20.57 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q6_K_L.ggufGGUFQ6_K_L21.14 GBDownload
Qwen3.6-27B-Omnimerge-v4-Q8_0.ggufGGUFQ8_026.63 GBDownload
mmproj-Qwen3.6-27B-Omnimerge-v4-F16.ggufGGUFF16884.6 MBDownload

Model Details

Model IDManniX-ITA/Qwen3.6-27B-Omnimerge-v4-GGUF
AuthorManniX-ITA
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelManniX-ITA/Qwen3.6-27B-Omnimerge-v4
Last modified2026-09-09T14:11:31.000Z

Model README

---

base_model: ManniX-ITA/Qwen3.6-27B-Omnimerge-v4

base_model_relation: quantized

license: apache-2.0

language:

  • en

tags:

- gguf

- imatrix

- quantized

- merge

- mergekit

- qwen3_5

- reasoning

- code

pipeline_tag: image-text-to-text

library_name: gguf

---

Qwen3.6-27B-Omnimerge-v4-GGUF

GGUF quantizations of ManniX-ITA/Qwen3.6-27B-Omnimerge-v4 — the MLP-passthrough variant that defends against the Qwen3.6 think-policy fragility we discovered. Source dtype is BF16; this repo provides the standard bartowski quant ladder (F16 → IQ2_XXS) for llama.cpp.

> Source model: ManniX-ITA/Qwen3.6-27B-Omnimerge-v4 (BF16 weights, model card with full benchmarks and methodology).

> NOT a quant of clean Qwen/Qwen3.6-27B — these GGUFs contain the v4 merge.

>

> MTP companion (2× decode speedup): weight-identical GGUFs with the MTP head retained for llama.cpp --spec-type draft-mtp self-speculative decoding are at ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-MTP-GGUF. Quality is statistically indistinguishable from this repo (HE 137/164 ↔ 137/164, GPQA 155/198 ↔ 154/198); aggregate decode is 2.0-2.3 × faster on a single 24 GB GPU. Use that repo for interactive / single-request workloads where latency matters.

All quants made using imatrix with calibration data v5, the same calibration set bartowski uses for the Qwen3.6 base release — so quality fingerprints are directly comparable to bartowski's Qwen_Qwen3.6-27B-GGUF repo.

Why this merge exists

Same-base DARE-TIES (Omnimerge_v2 method) merge of Qwen/Qwen3.6-27B + 3 Qwen3.6 fine-tunes. Direct successor to ManniX-ITA/Qwen3.5-27B-Omnimerge-v2 on the newer Qwen3.6 base, with mlp.{gate,up,down}_proj copied verbatim from clean Qwen3.6 (the "MLP-passthrough" surgery) to defend against a Qwen3.6-specific reasoning-tag fragility we found during forensic delta inspection. See the v4 model card for the full story, scripts, and benchmark methodology.

Benchmark headline (Q6_K, head-to-head vs Qwen3.6 base + Omnimerge-v2)

All scored under identical llama.cpp + lm_eval conditions (--reasoning-format deepseek --reasoning-budget 8192 --parallel 2, raw /v1/completions, no chat template).

| Benchmark | Qwen3.6 base Q6_K (bartowski) | Omnimerge-v2 (Qwen3.5 base) | Omnimerge-v4-MLP (this) | Δ vs base | Δ vs v2 |

|---|---|---|---|---|---|

| HumanEval pass@1 (164q) | 84.76% | 79.27% | 83.54% (137/164) | −1.22 pp | +4.27 pp |

| MBPP pass@1 (500q) — corrected\* | 57.60% | 74.60% | 73.00% (365/500) | +15.40 pp | −1.60 pp |

| GPQA Diamond pass@1 (flex) — full greedy§ | not measured | 69.19% (full 198q) | 78.28% (155/198) | — | +9.09 pp |

\* MBPP scores are post-<think>-stripping (lm_eval's raw scorer SyntaxErrors on literal < in exec(prompt+completion+tests)). See the v4 model card for the per-model recovery breakdown.

§ Canonical full-198q greedy GPQA result measured 2026-05-22 on pod 37268930 (Vast.ai 3090) with the patched eval chain (lm-eval 0.4.11 + max_length=32768 override + the api_models.py:545 UnboundLocalError patch + aiohttp lifecycle workaround). Sampler: do_sample=False, temperature=0.0, max_gen_toks=8192. Wall time 4 h 55 min. Companion strict-match (rigid Answer: X template) is 7.58 % — the model emits CoT verbosely rather than the strict template, so flex is the real quality signal. Earlier card revisions reported an ≈ 84.75 % partial result (177/198 sampled at T=0.6, budget=16384); that number is superseded by this canonical greedy measurement on the full bench — the 6.5 pp difference is driven by the methodology change (sampler / budget / completeness), not by a model change.

Sampled cohort (recommended, T=0.6) — not comparable to the greedy table above

These five benches were measured 2026-08-22/23 under a different sampler from the

greedy head-to-head table. They are reported as their own cohort and must never be pooled

with, or differenced against, the greedy rows.

Basis: Q6_K · llama.cpp b9700 · backend llama · sampler profile qwen3.6-35B-A3B,

preset recommendedtemperature 0.6, top_p 0.95, top_k 20, do_sample true.

| Benchmark | n | Score | metric / filter |

|---|---|---|---|

| GPQA Diamond | 198 | 78.79% | exact_match / flexible-extract ⚠¹ |

| HumanEval (thinking) | 164 | 98.17% | pass@1 / extract_chat |

| IFEval | 100 | 95.00% | prompt_level_strict_acc |

| LiveCodeBench v6 | 77 | 81.82% | pass_at_1 ⚠² |

| MultiPL-E | 300 | 87.67% | pass_at_1 ⚠³ |

⚠¹ Truncation-taxed. 5/198 completions (2.53%) stopped at the cap (8191 tok,

answer_allowance; max_gen_toks=16384, thinking_token_budget=8192). 1 of the 5 still

scored. Not comparable to a cell run at a different budget.

⚠² Truncation-taxed. 3/77 completions (3.9%) stopped at the cap (32767 tok;

max_gen_toks=32768, thinking_token_budget=12288). 0 of the capped rows scored.

⚠³ MultiPL-E reports 300 samples (3 languages × 100), not 100. Scored post-bug-604

(chat_to_body extraction fix, 2026-08-20); this run finished 2026-08-22, so it is a

post-fix cell.

Why a second table rather than more rows: the greedy figures above are the cross-cohort

comparison anchor. HumanEval illustrates the gap — 83.54% greedy on raw /v1/completions

vs 98.17% here, which differs by both sampler and bench construction

(humaneval_full_think uses a thinking scaffold and extract_chat). Those are two

different measurements, not two estimates of one number.

Available Quantizations

All 27 files (F16 + 26 imatrix-quantized tiers, ~417 GB total) are uploaded and ready. imatrix.dat (used for every quant) is in the repo root for audit and reproduction.

| Quantization | File size | Use case |

|---|---|---|

| F16 (full precision) | 50.11 GB | Conversion source / lossless reference |

| Q8_0 | 26.63 GB | Highest fidelity, large |

| Q6_K_L | 21.14 GB | Q6_K with embed/output at Q8_0 |

| Q6_K | 20.57 GB | Recommended high tier — eval methodology used this |

| Q5_K_L | 18.64 GB | Q5_K_M with embed/output at Q8_0 |

| Q5_K_M | 17.91 GB | Strong fidelity, balanced |

| Q5_K_S | 17.40 GB | Slightly smaller K-mix |

| Q4_K_L | 16.29 GB | Q4_K_M with embed/output at Q8_0 |

| Q4_1 | 15.91 GB | Legacy 4-bit, dense |

| Q4_K_M | 15.41 GB | Recommended balanced tier for most users |

| IQ4_NL | 14.72 GB | Importance-aware 4-bit non-linear |

| Q4_K_S | 14.52 GB | K-mix small variant |

| Q4_0 | 14.41 GB | Legacy 4-bit |

| IQ4_XS | 14.05 GB | IQ4 extra-small |

| Q3_K_XL | 13.42 GB | Q3_K_L with embed/output at Q8_0 |

| Q3_K_L | 13.36 GB | 3-bit K-mix large |

| Q3_K_M | 12.39 GB | 3-bit K-mix medium |

| IQ3_M | 11.72 GB | Importance-aware 3-bit medium |

| Q3_K_S | 11.24 GB | 3-bit K-mix small |

| IQ3_XS | 11.15 GB | IQ3 extra-small |

| Q2_K_L | 11.13 GB | Q2_K with embed/output at Q8_0 |

| IQ3_XXS | 10.42 GB | IQ3 extra-extra-small |

| Q2_K | 9.98 GB | 2-bit K-mix |

| IQ2_M | 9.32 GB | Importance-aware 2-bit medium |

| IQ2_S | 8.72 GB | IQ2 small |

| IQ2_XS | 8.47 GB | IQ2 extra-small |

| IQ2_XXS | 7.85 GB | IQ2 extra-extra-small (smallest) |

How to Use

With llama.cpp:

# Recommended args for reasoning-tag-emitting models:
llama-server \
    -m Qwen3.6-27B-Omnimerge-v4-Q4_K_M.gguf \
    -c 32768 -ngl 99 -t 12 --no-warmup \
    --reasoning-budget 8192

See Reasoning budget and thinking stop phrase below for the budget, the wrap-up phrase that stops reasoning

leaking into the answer, and why --reasoning-format plays no part in it.

The published evals additionally pin --reasoning-format deepseek so that

lm_eval sees only the answer; it is not needed for ordinary serving.

Swap Q4_K_M for any tier from the table above. Q6_K matches the methodology used in our published evals; Q4_K_M is the typical "balanced" choice for most users.

For multimodal (vision) inference: the mmproj projector is in bartowski/Qwen_Qwen3.6-27B-GGUF and works with this model unchanged (vision tower is preserved verbatim from the base).

With ollama: use a Modelfile pointing to one of the GGUFs above, or HF direct load.

Reasoning budget and thinking stop phrase (llama.cpp)

Qwen 3.6 reasons at length by design, and on a hard prompt it can consume the

whole context window before it answers. llama.cpp can bound the thinking block

with a sampler, and — the part that actually matters — tell the model why the

block is being closed.

Needs llama.cpp b8508 or newer for the flags, b10091 or newer for the

per-request overrides.

Serve with a bounded thinking block

llama-server -m Qwen3.6-27B-Omnimerge-v4-Q4_K_M.gguf -c 32768 -ngl 99 \
    --jinja \
    --reasoning-budget 8192 \
    --reasoning-budget-message $'\n\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n' \
    --temp 0.6 --top-k 20 --top-p 0.95

| flag | meaning |

|---|---|

| --reasoning-budget N | -1 unrestricted (default), 0 close the block immediately, N > 0 cap it at N tokens |

| --reasoning-budget-message | text written into the block just before the closing tag is forced |

| --jinja | required — the delimiters come from the chat template (<think></think>). Without it llama.cpp has no tags to count and the budget silently does nothing |

Both flags also read from the environment: LLAMA_ARG_THINK_BUDGET and

LLAMA_ARG_THINK_BUDGET_MESSAGE.

--reasoning-format is not part of this. It only decides how the thinking

is handed back — message.reasoning_content versus left inline in

message.content — and never whether the budget is enforced: the delimiters the

sampler counts are set by the chat template regardless, so the cap binds under

auto, deepseek and none alike. The default auto already extracts

reasoning and is behaviourally identical to deepseek (they differ only in

name; the sole branch in the parser is != none). Leave it at the default so

the model's own tool-call and channel handling stays in play, and pin

deepseek only when a harness needs the thinking kept out of content.

--reasoning-budget on its own forces the closing tag the moment the budget

runs out, wherever the model happens to be. When that lands mid-thought the

model frequently does not register that it was interrupted: it carries on

reasoning, now inside the visible answer. The stop phrase is what prevents

that — it gives the model a reason to be finishing.

Two wordings that work

# "qwen" — the string Qwen's own service uses, from their docs
--reasoning-budget-message $'\n\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n'

# "voice" — shorter, in the model's own reasoning voice
--reasoning-budget-message $'\n\nOK, I have enough to answer now.\n'

Wording is model-specific: Qwen note that the ability to act on such a message

"is not explicitly trained but emerges naturally", so it is worth trying both

on your own workload. Leading and trailing newlines matter — they keep the

phrase off whatever half-finished line the cut landed on.

What it measures out to

Measured on the Qwen3.6-35B-A3B base this model is pruned from. Three hard

questions, temperature 0.6, fixed seed, answer characters with wall time in

brackets. Every run answered all three correctly, and thinking length is

unchanged by the message in every row:

| budget | no message | qwen | voice |

|---|---|---|---|

| 2048 | 1907 (69 s) | 1838 (42 s) | 1615 (41 s) |

| 4096 | 18015 (170 s) | 2642 (78 s) | 1441 (104 s) |

| 8192 | 3642 (158 s) | 1848 (129 s) | 2023 (175 s)|

The 4096 row is the failure this exists for: the cap lands mid-thought and the

reasoning simply continues in the answer, ten times longer and 2.2x the wall

time, for the same three correct answers. Both phrases remove it.

Per request, instead of per server

The server accepts both as request fields, overriding the command line:

{
  "messages": [ ... ],
  "thinking_budget_tokens": 8192,
  "reasoning_budget_message": "\n\nOK, I have enough to answer now.\n"
}

On the raw /completion endpoint the delimiters are not inferred, so they have

to be supplied with the budget:

{
  "prompt": "...",
  "reasoning_budget_tokens": 8192,
  "reasoning_budget_start_tag": "<think>",
  "reasoning_budget_end_tag": "</think>",
  "reasoning_budget_message": "\n\nOK, I have enough to answer now.\n"
}

On b10091 the message field must be present on /completion requests even

when empty: llama.cpp builds the sequence it forces from message + end_tag

inside that field's handler, so omitting it leaves the budget with nothing to

force — the sampler logs as though the cap fired while the thinking block stays

open.

Rules of thumb

  • Keep -c several times larger than the budget. A budget equal to the context

lets the thinking phase fill the window on its own.

  • A quarter of the context is a sensible starting point: 8192 at -c 32768.
  • Qwen recommend keeping a thinking budget above 1024 tokens; below that

the cap tends to land before the model has committed to an approach.

  • The budget is per thinking block, not per response — the sampler re-arms

when it sees a new opening tag, so a multi-turn agent gets a fresh window each

time.

imatrix.dat

The imatrix.dat (~14 MB) used to generate every quant in this repo is uploaded alongside the GGUFs at the repo root. Reproducible, auditable.

Reproducing

See scripts/ on the source v4 model repo:

  • dare_ties_merge.py — main merger (auto-detects Qwen3.6 base via output_gate_type and applies MLP-skip)
  • v4_mlp_passthrough.py — post-process: rebuild merged dir with MLP layers from base
  • quantize_gguf.py — the script that built this repo

For dense (non-Gemma-4-MoE) models, pass --exclude CD-Q6_K,CD-Q5_K_M,CD-Q4_K_M,CD-Q3_K_M,CD-Q2_K to skip ContribDynamic tiers (those require Gemma 4 expert-contribution maps).

License

Apache-2.0 (inherited from Qwen/Qwen3.6-27B and the fine-tune sources).

Acknowledgements

Tool-calling benchmark — tool-eval-bench hardmode (88 scenarios, 176 pts)

> Measured on the MTP build of this quant, not on this repo's file. Both repos

> ship a Qwen3.6-27B-Omnimerge-v4-Q4_K_M.gguf, and they are different artifacts:

> 16,547,399,232 B here vs 16,810,713,696 B in

> -MTP-GGUF

> — a ~263 MB difference, which is the 15 mtp.* tensors. The cohort ran with

> nextn=YES spec=mtp, and the served file was confirmed by byte-exact size to be

> the MTP one.

>

> The score is reported here because **speculative decoding is distribution-

> preserving**: the MTP head changes draft-acceptance rate and decode speed, not

> what the model computes. This build is therefore expected to score the same

> within the ±2.7 seed noise. That is a reasoned expectation, not a measurement

> — this exact file has not been run.

v4 scores 146.2 ±2.7 of 176, joint third of ten, tied to the decimal with

Ornith-1.5-35B and 2.2 pts above its own Qwen3.6-27B base (144.0). Its successor

Omnimerge-v6 scores 156.4,

and the two CIs do not overlap — on this benchmark v6 supersedes v4 outright.

Category profile (mean over 5 seeds): perfect on Tool Selection, Parameter Precision,

Localization, Creative Composition and Structured Output (12/12). Hard Mode 31.4/38

(82.6%). Weakest at Autonomous Planning 4.0/6 (66.7%) and Safety & Boundaries

18.4/26 (70.8%).

Safety caveat, stated plainly: 16 safety-critical failures across five seeds —

TC-31 (Ambiguity Resolution), TC-34 (Prompt Injection Resistance) and TC-60 (Cross-Turn

Sleeper Injection) fail on every seed. TC-60 is a cohort-wide weakness (every model

here fails it 5/5 except v6), but TC-31 and TC-34 are not: v6 passes both on all seeds.

If your deployment exposes the model to untrusted tool output, prefer v6.

v4 has no cell affected by the TC-62 scorer crash described below, so its score is

not inflated or deflated by it.

!Tool-calling benchmark

Full cohort

| model | quant | Total Points (mean, 5 seeds) | 95% CI | safety-critical (5 seeds) |

|---|---|---|---|---|

| Qwen3.8-27B-Omnimerge-v6 | Q4_K_M | 156.4 ±3.5 | [152.0, 160.8] | 3 |

| Qwen3.8-27B (base) | UD-Q4_K_M | 150.8 ±2.5 | [147.7, 153.9] | 9 |

| Ornith-1.5-35B | IQ4_XS | 146.2 ±2.6 | [143.0, 149.4] | 10 |

| Qwen3.6-27B-Omnimerge-v4 | Q4_K_M | 146.2 ±2.7 | [142.9, 149.5] | 16 |

| Qwen3.6-27B (base) | Q4_K_M | 144.0 ±3.4 | [139.8, 148.2] | 14 |

| Qwen3.6-35B-A3B (base) | IQ4_XS | 141.6 ±2.4 | [138.6, 144.6] | 15 |

| Qwen3.6-27B-A3B-CoderX | Q4_K_M | 137.4 ±4.9 | [131.3, 143.5] | 17 |

| Ornith-1.5-27B-A3B-Coder | IQ4_XS | 136.8 ±4.8 | [130.9, 142.7] | 12 |

| Ornith-1.5-27B-A3B-CoderX | IQ4_XS | 134.0 ±2.5 * | [130.8, 137.2] | 14 |

| Qwen3.6-27B-A3B-Coder | Q4_K_M | 123.2 ±2.3 | [120.4, 126.0] | 15 |

* one seed (s42) is graded on 174 pts, not 176 — see that model's card.

<details>

<summary><b>Basis — read before comparing these numbers to anything</b></summary>

  • Scorer: tool-eval-bench v2.6.0 (the pip/uv-installed package, verified via

tool_eval_bench.__file__, not a git checkout). An earlier note in the runner claimed

cf54b4b (v2.6.0-45); that is wrong and has been corrected — no cell ever ran it.

All 50 cells ran the same v2.6.0, so the cohort is internally consistent.

  • v2.6.0 carries a known scorer crash on TC-62. email_calls[-1] raises IndexError

when a model sent no valid CFO email; the orchestrator catches it and returns

FAIL / 0 points while keeping the scenario in the denominator. It hits 11 of 38

scored cells, 2 pts each, and it is **not neutral — it concentrates on the weakest

models**. Later harness commits credit that behaviour instead, so a fixed scorer would

raise affected scores, unevenly.

  • 5 paired seeds [42–46], 64k context, context-pressure 0.25, max 8 turns, 120 s timeout,

thinking enabled, sampler temp 0.6 / top-p 0.95 / top-k 20 (not greedy).

  • Served on llama.cpp b1788384120-c588c4f47 with MTP speculative decoding enabled

(nextn=YES spec=mtp), one model per GPU, sequential.

  • Quant tiers are not uniform across the cohort (Q4_K_M for the Omnimerge/A3B rows,

IQ4_XS for Ornith and 35B-A3B, UD-Q4_K_M for the Qwen3.8 base). Cross-row gaps

therefore carry a quantisation component and are not purely architectural.

  • Do not pool these with the r/LocalLLaMA published tool-eval-bench figures: those

were run at 256k context and are a different basis despite the shared scorer version.

</details>

Run ManniX-ITA/Qwen3.6-27B-Omnimerge-v4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models