final-bench/darwin-35b-a3b-opus-q8-gguf 00003 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
final-bench/darwin-35b-a3b-opus-q8-gguf overview
Q8_0 GGUF of Darwin-35B-A3B-Opus | ~37GB (3 shards) | GPQA Diamond 90.0% | Near-lossless quality | MoE 35B (3B active) | 201 Languages | 262K Context | Apache 2.0 ---
Downloads
2,781
Likes
17
Pipeline
text-generation
Library
gguf
Visibility
Public
Access
Open
Repository Files & Downloads
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"license": "apache-2.0",
"base_model": [
"FINAL-Bench/Darwin-35B-A3B-Opus"
],
"tags": [
"llama-cpp",
"gguf",
"quantized",
"Q8_0",
"merge",
"evolutionary-merge",
"darwin",
"darwin-v5",
"reasoning",
"qwen3.5",
"qwen",
"moe",
"mixture-of-experts",
"claude-opus",
"distillation",
"multilingual",
"gpqa",
"open-source",
"apache-2.0",
"layer-wise-merge",
"coding-agent",
"tool-calling",
"long-context",
"262k-context"
],
"language": [
"en",
"zh",
"ko",
"ja",
"de",
"fr",
"es",
"ru",
"ar",
"multilingual"
],
"pipeline_tag": "text-generation",
"library_name": "gguf",
"quantized_by": "VIDRAFT",
"model-index": [
{
"name": "Darwin-35B-A3B-Opus-Q8_0-GGUF",
"results": [
{
"task": {
"type": "text-generation",
"name": "Graduate-Level Reasoning"
},
"dataset": {
"type": "Idavidrein/gpqa",
"name": "GPQA Diamond",
"config": "gpqa_diamond",
"split": "train"
},
"metrics": [
{
"type": "accuracy",
"value": 90,
"name": "Accuracy",
"verified": false
}
]
},
{
"task": {
"type": "text-generation",
"name": "Multilingual Knowledge"
},
"dataset": {
"type": "openai/MMMLU",
"name": "MMMLU"
},
"metrics": [
{
"type": "accuracy",
"value": 85,
"name": "Accuracy",
"verified": false
}
]
}
]
}
],
"frontmatter": {
"license": "apache-2.0",
"base_model": [
"FINAL-Bench/Darwin-35B-A3B-Opus"
],
"tags": [
"llama-cpp",
"gguf",
"quantized",
"Q8_0",
"merge",
"evolutionary-merge",
"darwin",
"darwin-v5",
"reasoning",
"qwen3.5",
"qwen",
"moe",
"mixture-of-experts",
"claude-opus",
"distillation",
"multilingual",
"gpqa",
"open-source",
"apache-2.0",
"layer-wise-merge",
"coding-agent",
"tool-calling",
"long-context",
"262k-context"
],
"language": [
"en",
"zh",
"ko",
"ja",
"de",
"fr",
"es",
"ru",
"ar",
"multilingual"
],
"pipeline_tag": "text-generation",
"library_name": "gguf",
"quantized_by": [
"name: Darwin-35B-A3B-Opus-Q8_0-GGUF",
"task:",
"type: accuracy",
"task:",
"type: accuracy"
]
},
"hero_image_url": "https://img.shields.io/badge/๐งฌ_Gen1-Darwin--4B--Opus-blue?style=for-the-badge",
"summary": "> Q8_0 GGUF of Darwin-35B-A3B-Opus | ~37GB (3 shards) | GPQA Diamond 90.0% | Near-lossless quality | MoE 35B (3B active) | 201 Languages | 262K Context | Apache 2.0 ---",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nlicense: apache-2.0\nbase_model:\n - FINAL-Bench/Darwin-35B-A3B-Opus\ntags:\n - llama-cpp\n - gguf\n - quantized\n - Q8_0\n - merge\n - evolutionary-merge\n - darwin\n - darwin-v5\n - reasoning\n - qwen3.5\n - qwen\n - moe\n - mixture-of-experts\n - claude-opus\n - distillation\n - multilingual\n - gpqa\n - open-source\n - apache-2.0\n - layer-wise-merge\n - coding-agent\n - tool-calling\n - long-context\n - 262k-context\nlanguage:\n - en\n - zh\n - ko\n - ja\n - de\n - fr\n - es\n - ru\n - ar\n - multilingual\npipeline_tag: text-generation\nlibrary_name: gguf\nquantized_by: VIDRAFT\nmodel-index:\n - name: Darwin-35B-A3B-Opus-Q8_0-GGUF\n results:\n - task:\n type: text-generation\n name: Graduate-Level Reasoning\n dataset:\n type: Idavidrein/gpqa\n name: GPQA Diamond\n config: gpqa_diamond\n split: train\n metrics:\n - type: accuracy\n value: 90.0\n name: Accuracy\n verified: false\n - task:\n type: text-generation\n name: Multilingual Knowledge\n dataset:\n type: openai/MMMLU\n name: MMMLU\n metrics:\n - type: accuracy\n value: 85.0\n name: Accuracy\n verified: false\n---\n\n# Darwin-35B-A3B-Opus-Q8_0-GGUF\n\n<p align=\"center\">\n <a href=\"https://huggingface.co/FINAL-Bench/Darwin-4B-Opus\"><img src=\"https://img.shields.io/badge/๐งฌ_Gen1-Darwin--4B--Opus-blue?style=for-the-badge\" alt=\"Gen1\"></a>\n <a href=\"https://huggingface.co/FINAL-Bench/Darwin-4B-David\"><img src=\"https://img.shields.io/badge/๐งฌ_Gen2-Darwin--4B--David-blue?style=for-the-badge\" alt=\"Gen2\"></a>\n <a href=\"https://huggingface.co/FINAL-Bench/Darwin-4B-Genesis\"><img src=\"https://img.shields.io/badge/โญ_Gen3-Darwin--4B--Genesis-gold?style=for-the-badge\" alt=\"Gen3\"></a>\n</p>\n\n<p align=\"center\">\n <a href=\"https://huggingface.co/FINAL-Bench/Darwin-9B-Opus\"><img src=\"https://img.shields.io/badge/๐งฌ_Model-Darwin--9B--Opus-blue?style=for-the-badge\" alt=\"9B\"></a>\n <a href=\"https://huggingface.co/spaces/FINAL-Bench/Darwin-9B-Opus\"><img src=\"https://img.shields.io/badge/๐_Space-9B_Demo-purple?style=for-the-badge\" alt=\"9B Space\"></a>\n <a href=\"https://huggingface.co/FINAL-Bench/Darwin-31B-Opus\"><img src=\"https://img.shields.io/badge/๐งฌ_Model-Darwin--31B--Opus-blue?style=for-the-badge\" alt=\"31B\"></a>\n <a href=\"https://huggingface.co/spaces/FINAL-Bench/Darwin-31B-Opus\"><img src=\"https://img.shields.io/badge/๐_Space-31B_Demo-purple?style=for-the-badge\" alt=\"31B Space\"></a>\n</p>\n\n<p align=\"center\">\n <a href=\"https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus\"><img src=\"https://img.shields.io/badge/๐งฌ_Model-Darwin--35B--A3B--Opus-blue?style=for-the-badge\" alt=\"35B\"></a>\n <a href=\"https://huggingface.co/spaces/FINAL-Bench/Darwin-35B-A3B-Opus\"><img src=\"https://img.shields.io/badge/๐_Space-35B_Demo-purple?style=for-the-badge\" alt=\"35B Space\"></a>\n <a href=\"https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF\"><img src=\"https://img.shields.io/badge/๐ฆ_GGUF-Q8--Official-yellow?style=for-the-badge\" alt=\"Q8 GGUF\"></a>\n <a href=\"https://huggingface.co/bartowski/FINAL-Bench_Darwin-35B-A3B-Opus-GGUF\"><img src=\"https://img.shields.io/badge/๐ฆ_GGUF-bartowski-yellow?style=for-the-badge\" alt=\"bartowski GGUF\"></a>\n</p>\n\n<p align=\"center\">\n <a href=\"https://huggingface.co/spaces/FINAL-Bench/Leaderboard\"><img src=\"https://img.shields.io/badge/๐_FINAL_Bench-Leaderboard-green?style=for-the-badge\" alt=\"FINAL Bench\"></a>\n <a href=\"https://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard\"><img src=\"https://img.shields.io/badge/๐_ALL_Bench-Leaderboard-orange?style=for-the-badge\" alt=\"ALL Bench\"></a>\n</p>\n\n> Q8_0 GGUF of Darwin-35B-A3B-Opus | ~37GB (3 shards) | GPQA Diamond 90.0% | Near-lossless quality | MoE 35B (3B active) | 201 Languages | 262K Context | Apache 2.0\n\n---\n\n## About This Quantization\n\nQ8_0 GGUF of [`FINAL-Bench/Darwin-35B-A3B-Opus`](https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus).\n\n| | Original (BF16) | This Model (Q8_0 GGUF) |\n|---|---|---|\n| Format | SafeTensors | GGUF |\n| Size | 65.5 GB | ~37 GB (3 shards) |\n| Quality | Baseline | Near-lossless (~99.9% of BF16) |\n| VRAM Required | 65+ GB | ~37 GB |\n| Runs on | H100, A100 80GB | A100 40GB, Mac 64GB, 2x RTX 4090 |\n| Framework | Transformers, vLLM, SGLang | llama.cpp, Ollama, LM Studio |\n\n---\n\n## Files\n\n| File | Size | Description |\n|---|---|---|\n| `darwin-35b-a3b-opus-q8_0-00001-of-00003.gguf` | ~13.6 GB | Shard 1 of 3 |\n| `darwin-35b-a3b-opus-q8_0-00002-of-00003.gguf` | ~12.5 GB | Shard 2 of 3 |\n| `darwin-35b-a3b-opus-q8_0-00003-of-00003.gguf` | ~10.7 GB | Shard 3 of 3 |\n| Total | ~36.8 GB | All 3 shards required |\n\n> Download all 3 shard files. llama.cpp and Ollama will automatically load them together.\n\n---\n\n## Hardware Requirements\n\n| Setup | Memory | Status |\n|---|---|---|\n| NVIDIA A100 40GB | 40 GB VRAM | Fits |\n| NVIDIA A100 80GB | 80 GB VRAM | Comfortable |\n| NVIDIA H100 93GB | 93 GB VRAM | Comfortable |\n| 2x RTX 4090 (24GB each) | 48 GB VRAM | With tensor parallel |\n| Mac Studio M2/M3 Ultra 64GB | 64 GB Unified | Fits |\n| Mac M3 Max 48GB | 48 GB Unified | Fits |\n| Single RTX 4090 24GB | 24 GB VRAM | Insufficient (use Q4_K_M) |\n\n> As a MoE model, only 3B parameters are active per token. Inference is fast despite the 37GB model size.\n\n---\n\n## Usage\n\n### llama.cpp (CLI)\n\n```bash\nllama-cli \\\n --hf-repo FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF \\\n --hf-file darwin-35b-a3b-opus-q8_0-00001-of-00003.gguf \\\n -p \"The meaning to life and the universe is\" \\\n -n 512 -ngl 99\n```\n\n### llama.cpp (Server)\n\n```bash\nllama-server \\\n --hf-repo FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF \\\n --hf-file darwin-35b-a3b-opus-q8_0-00001-of-00003.gguf \\\n -c 32768 -ngl 99\n```\n\n### Ollama\n\n```bash\necho 'FROM ./darwin-35b-a3b-opus-q8_0-00001-of-00003.gguf' > Modelfile\nollama create darwin-opus -f Modelfile\nollama run darwin-opus\n```\n\n### LM Studio\n\n1. Download all 3 `.gguf` shard files\n2. Place them in the same folder\n3. Open LM Studio, load the first shard\n4. LM Studio auto-detects and loads all shards\n\n### MoE Expert Offload (Limited VRAM)\n\n```bash\nllama-cli \\\n --hf-repo FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF \\\n --hf-file darwin-35b-a3b-opus-q8_0-00001-of-00003.gguf \\\n -ot \".ffn_.*_exps.=CPU\" \\\n -ngl 99 -c 32768\n```\n\n---\n\n## Benchmark Results (Original Model)\n\nQ8_0 preserves near-identical performance to BF16.\n\nGPQA Diamond (198 Questions, Graduate-Level Reasoning)\n\n| Model | Accuracy |\n|---|---|\n| Darwin-35B-A3B-Opus | 90.0% |\n| Mother (Jackrong Claude 4.6 Opus Distilled) | 85.0% |\n| Father (Qwen3.5-35B-A3B Official) | 84.2% |\n\nMMMLU (Multilingual Knowledge, 29 Languages)\n\n| Model | Accuracy |\n|---|---|\n| Darwin-35B-A3B-Opus | 85.0% |\n| Father (Qwen3.5-35B-A3B Official) | 85.2% |\n\n---\n\n## How Darwin Was Created\n\nDarwin-35B-A3B-Opus was created using Darwin V5, a diagnostic-guided evolutionary merge engine built on [mergekit](https://github.com/arcee-ai/mergekit).\n\nBoth parent models share the identical Qwen3.5-35B-A3B architecture. The Mother is a LoRA SFT on the same base โ not a different architecture.\n\nDarwin V5 adds three phases over standard mergekit evolve:\n1. Pre-merge parent profiling (40 layers x 256 experts: activation frequency, routing entropy, probe cosine distance)\n2. Evolution with diagnostic-informed initial population and constrained search space\n3. Post-merge child validation (layer-by-layer comparison against both parents)\n\nKey diagnostic finding: Mother had 50-65% dead experts (activation < 5%) from text-only LoRA SFT. Darwin compensated by reducing Mother density and using Father's living experts to fill inactive slots.\n\nMerge configuration:\n```yaml\n# Method: DARE-TIES via mergekit\nL0-L37: t=0.5988 (Mother 60%) โ router from Mother\nL38: t=0.9000 (Mother 90%) โ reasoning core (peak probe cosine distance)\nL39: t=0.5336 (Father 47%) โ router from Father (output routing)\n```\n\nFor full technical details, diagnostics, and health check results, see the [original model card](https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus).\n\n---\n\n## Other Quantizations\n\n| Quantization | Size | Quality | Use Case |\n|---|---|---|---|\n| Q8_0 (this) | ~37 GB | Near-lossless | Maximum quality |\n| Q4_K_M (coming soon) | ~20 GB | Good | RTX 4090, Mac 32GB |\n\n---\n\n## Model Specifications\n\n| | |\n|---|---|\n| Base Model | [FINAL-Bench/Darwin-35B-A3B-Opus](https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus) |\n| Architecture | Qwen3.5 MoE (Gated DeltaNet + MoE) |\n| Total Parameters | 35B |\n| Active Parameters | 3B per forward pass |\n| Experts | 256 (8 routed + 1 shared active) |\n| Context Length | 262,144 native |\n| Languages | 201 |\n| Quantization | Q8_0 (8-bit integer) |\n| GGUF Shards | 3 files |\n| License | Apache 2.0 |\n| Quantized by | VIDRAFT via llama.cpp |\n\n---\n\n## Acknowledgements\n\n- Korean Government โ GPU Support Program research grant\n- [Qwen Team](https://huggingface.co/Qwen) โ Qwen3.5-35B-A3B base architecture\n- [Jackrong](https://huggingface.co/Jackrong) โ Claude 4.6 Opus Reasoning Distilled model\n- [mergekit](https://github.com/arcee-ai/mergekit) โ Merge backend infrastructure\n- [llama.cpp](https://github.com/ggml-org/llama.cpp) โ GGUF conversion and quantization\n\n---\n\n## Citation\n\n```bibtex\n@misc{vidraft_darwin_35b_opus_gguf,\n title = {Darwin-35B-A3B-Opus-Q8_0-GGUF},\n author = {VIDRAFT},\n year = {2026},\n publisher = {Hugging Face},\n howpublished = {\\url{https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF}}\n}\n```",
"related_quantizations": []
},
"tags": [
"gguf",
"qwen3_5_moe",
"llama-cpp",
"quantized",
"Q8_0",
"merge",
"evolutionary-merge",
"darwin",
"darwin-v5",
"reasoning",
"qwen3.5",
"qwen",
"moe",
"mixture-of-experts",
"claude-opus",
"distillation",
"multilingual",
"gpqa",
"open-source",
"apache-2.0",
"layer-wise-merge",
"coding-agent",
"tool-calling",
"long-context",
"262k-context",
"text-generation",
"en",
"zh",
"ko",
"ja",
"de",
"fr",
"es",
"ru",
"ar",
"base_model:FINAL-Bench/Darwin-35B-A3B-Opus",
"base_model:quantized:FINAL-Bench/Darwin-35B-A3B-Opus",
"license:apache-2.0",
"model-index",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 17,
"downloads": 2781,
"gated": false,
"private": false,
"last_modified": "2026-04-11T03:10:21.000Z",
"created_at": "2026-04-02T06:10:29.000Z",
"pipeline_tag": "text-generation",
"library_name": "gguf"
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69ce0855423f8be15fedf553",
"id": "FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF",
"modelId": "FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF",
"sha": "dac09294d0235e9b50a18ec110da72714e9a70af",
"createdAt": "2026-04-02T06:10:29.000Z",
"lastModified": "2026-04-11T03:10:21.000Z",
"author": "FINAL-Bench",
"downloads": 2781,
"likes": 17,
"gated": false,
"private": false,
"pipeline_tag": "text-generation",
"library_name": "gguf",
"siblings_count": 18
}