GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

AtomicChat/Ling-3.0-flash-Fin-GGUF overview

How to Run Ling 3.0 Flash Fin Locally <p style="margin top: 0; margin bottom: 0;" <em Built from InclusionAI's original weights with our own importance matrix.…

ggufatomic-chatlingfinancefinancial-researchagentstool-uselong-contextmixture-of-expertsimatrixquantizedllama.cpptext-generationbase_model:inclusionAI/Ling-3.0-flash-Finbase_model:quantized:inclusionAI/Ling-3.0-flash-Finlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~4.08 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
2
Pipeline
text-generation

Repository Files & Downloads

13 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
AD-Q6_K/Ling-3.0-flash-Fin-AD-Q6_K-00001-of-00003.ggufGGUFQ6_K41.19 GBDownload
AD-Q6_K/Ling-3.0-flash-Fin-AD-Q6_K-00002-of-00003.ggufGGUFQ6_K41.35 GBDownload
AD-Q6_K/Ling-3.0-flash-Fin-AD-Q6_K-00003-of-00003.ggufGGUFQ6_K18.53 GBDownload
AD-Q8_0/Ling-3.0-flash-Fin-AD-Q8_0-00001-of-00004.ggufGGUFQ8_041.91 GBDownload
AD-Q8_0/Ling-3.0-flash-Fin-AD-Q8_0-00002-of-00004.ggufGGUFQ8_041.58 GBDownload
AD-Q8_0/Ling-3.0-flash-Fin-AD-Q8_0-00003-of-00004.ggufGGUFQ8_041.45 GBDownload
AD-Q8_0/Ling-3.0-flash-Fin-AD-Q8_0-00004-of-00004.ggufGGUFQ8_04.08 GBDownload
BF16/Ling-3.0-flash-Fin-BF16-00001-of-00006.ggufGGUFBF1641.68 GBDownload
BF16/Ling-3.0-flash-Fin-BF16-00002-of-00006.ggufGGUFBF1640.12 GBDownload
BF16/Ling-3.0-flash-Fin-BF16-00003-of-00006.ggufGGUFBF1640.86 GBDownload
BF16/Ling-3.0-flash-Fin-BF16-00004-of-00006.ggufGGUFBF1640.12 GBDownload
BF16/Ling-3.0-flash-Fin-BF16-00005-of-00006.ggufGGUFBF1640.12 GBDownload
BF16/Ling-3.0-flash-Fin-BF16-00006-of-00006.ggufGGUFBF1634.67 GBDownload

Model Details

Model IDAtomicChat/Ling-3.0-flash-Fin-GGUF
AuthorAtomicChat
Pipelinetext-generation
Licensemit
Base modelinclusionAI/Ling-3.0-flash-Fin
Last modified2026-09-03T20:09:36.000Z

Model README

---

license: mit

license_link: https://huggingface.co/inclusionAI/Ling-3.0-flash-Fin/blob/main/LICENSE

base_model:

  • inclusionAI/Ling-3.0-flash-Fin

base_model_relation: quantized

quantized_by: AtomicChat

pipeline_tag: text-generation

library_name: gguf

tags:

  • atomic-chat
  • ling
  • finance
  • financial-research
  • agents
  • tool-use
  • long-context
  • mixture-of-experts
  • gguf
  • imatrix
  • quantized
  • llama.cpp

---

How to Run Ling 3.0 Flash Fin Locally

<p style="margin-top: 0; margin-bottom: 0;">

<em>Built from InclusionAI's original weights with our own importance matrix. The <a href="https://huggingface.co/datasets/AtomicChat/calib-corpora">calibration corpora</a> behind our builds are public.</em>

</p>

<div style="display: flex; gap: 8px; align-items: center; margin-top: 10px; margin-bottom: 10px;">

<a href="https://atomic.chat/?utm_source=huggingface&utm_medium=referral&utm_campaign=hf_ling_3_0_flash_fin&utm_content=btn_atomic"><img src="https://huggingface.co/AtomicChat/Ling-3.0-flash-Fin-GGUF/resolve/main/btn_atomic.png" width="162" alt="Atomic Chat"></a>

<a href="https://discord.gg/8wGSsvmg4V"><img src="https://huggingface.co/AtomicChat/Ling-3.0-flash-Fin-GGUF/resolve/main/btn_discord.png" width="119" alt="Discord"></a>

<a href="https://github.com/AtomicBot-ai/Atomic-Chat"><img src="https://huggingface.co/AtomicChat/Ling-3.0-flash-Fin-GGUF/resolve/main/btn_github.png" width="115" alt="GitHub"></a>

</div>

<ul style="margin: 0 0 12px 0;">

<li>Ling 3.0 Flash Fin is InclusionAI's finance-enhanced Ling model for research, valuation, spreadsheets, and long-horizon agent workflows.</li>

<li>These GGUFs are self-quantized from InclusionAI's original BF16 weights with our own importance matrix.</li>

<li>The repo currently includes BF16, AD-Q8_0, and AD-Q6_K builds; lower-bit quants are still uploading.</li>

</ul>

<hr style="margin: 0 0 16px 0;">

Ling 3.0 Flash Fin, self-quantized to GGUF by Atomic Chat. It has 124B total parameters, activates 5.1B per token, and supports a 256K context window. The checkpoint extends Ling 3.0 Flash with continued training on high-quality financial data.

Highlights

  • End-to-end financial research: connects retrieval, evidence review, calculation, modeling, and report preparation in one workflow.
  • Source-grounded search: prioritizes authoritative sources and traceable answers. InclusionAI publishes FinFIRST for transparent evaluation.
  • Multi-document reasoning: reconciles periods, definitions, assumptions, and conflicting figures across filings, earnings materials, and research.
  • Valuation and spreadsheet workflows: understands formulas, estimate updates, cross-sheet dependencies, balance checks, and scenario analysis.
  • Research-ready output: separates facts, analysis, judgments, and charts into material that can be reviewed and edited.

> [!NOTE]

> These GGUFs are self-quantized from the original weights, not a repack.

> [!IMPORTANT]

> Thinking mode is enabled by default. InclusionAI recommends temperature=1.0, top_p=0.95, and top_k=20.

Model overview

| Property | Value |

|---|---|

| Base model | inclusionAI/Ling-3.0-flash-Fin |

| Type | Finance-enhanced mixture-of-experts language model |

| Total / active parameters | 124B total / 5.1B active |

| Context length | 256K tokens |

| Focus | Financial research, source review, valuation, spreadsheets, and agent workflows |

| This repo | GGUF builds made directly from the original BF16 checkpoint |

Pick a file

| Build | Download size | Notes |

|---|---:|---|

| AD-Q6_K | 108.5 GB | Recommended current download. Near-lossless and 30 GB smaller than Q8_0. |

| AD-Q8_0 | 138.5 GB | Reference-quality quant for machines with enough memory. |

| BF16 | 255.1 GB | Full-precision GGUF reference. Not intended for most local systems. |

The download must fit alongside the context cache and runtime overhead. Leave several gigabytes of headroom beyond the file size.

Get started

Run Ling 3.0 Flash Fin locally with:

  • Atomic Chat: open the app, search AtomicChat/Ling-3.0-flash-Fin-GGUF, pick a build, and select Use this model.
  • llama.cpp: download the AD-Q6_K folder and open its first shard with llama-server.
  • LM Studio / Jan: search the repo ID and download the build that fits your machine.

Download the recommended build:

hf download AtomicChat/Ling-3.0-flash-Fin-GGUF \
  --include "AD-Q6_K/*" \
  --local-dir Ling-3.0-flash-Fin-GGUF

Run it:

llama-server \
  -m Ling-3.0-flash-Fin-GGUF/AD-Q6_K/Ling-3.0-flash-Fin-AD-Q6_K-00001-of-00003.gguf \
  --jinja -ngl 99 -c 32768

The chat template is embedded in the GGUF. Keep --jinja enabled so thinking mode and tool-call formatting are applied correctly.

Best practices

| Parameter | Value |

|---|---|

| temperature | 1.0 |

| top_p | 0.95 |

| top_k | 20 |

Allocate enough output length for research and agent tasks. Financial conclusions, valuation assumptions, and investment decisions still require professional review.

What the model is built for

The producer evaluates the model on FinFIRST, FinSearchComp Verified, FinCRAFT, Finance Agent, APEX-Agents, SpreadsheetBench, and tau3-Banking. These benchmarks cover source-grounded retrieval, investment research, long-horizon execution, valuation modeling, spreadsheet operations, and banking workflows.

See the official model card for the producer's benchmark results, deployment guidance, and limitations.

How these were made

  1. Start from inclusionAI/Ling-3.0-flash-Fin, the original BF16 checkpoint.
  2. Convert the checkpoint directly to GGUF.
  3. Build a per-tensor importance matrix over the public Atomic Chat calibration corpora.
  4. Quantize the shipped builds and validate them against the BF16 reference.

The raw evaluation logs currently available for the uploaded builds are included in the logs/ directory of this repository.

License

Released by InclusionAI under the MIT License. Quantized by Atomic Chat.

Run AtomicChat/Ling-3.0-flash-Fin-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models