unsloth/lfm2-350m-gguf overview
LFM2 is a new generation of hybrid models developed by Liquid AI, specifically designed for edge AI and on-device deployment. It sets a new standard in terms of quality, speed, and memory efficiency. We're releasing the weights of three post-trained checkpoints with 350M, 700M, and 1.2B parameters. They provide the following key features to create AI-powered edge applications: Fast training & inference – LFM2 achieves 3x faster training compared to its previous generation. It also benefits from 2x faster decode and prefill speed on CPU compared to Qwen3. Best performance – LFM2 outperforms similarly-sized models across multiple benchmark categories, including knowledge, mathematics, instruction following, and multilingual capabilities. New architecture – LFM2 is a new hybrid Liquid model with multiplicative gates and short convolutions. Flexible deployment – LFM2 runs efficiently on CPU, GPU, and NPU hardware for flexible deployment on smartphones, laptops, or vehicles. Find more information about LFM2 in our blog post.
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| LFM2-350M-F16.gguf | GGUF | F16 | 678.52 MB | Download |
| LFM2-350M-Q2_K.gguf | GGUF | Q2_K | 153.15 MB | Download |
| LFM2-350M-Q2_K_L.gguf | GGUF | Q2_K_L | 153.15 MB | Download |
| LFM2-350M-Q3_K_M.gguf | GGUF | Q3_K_M | 184.20 MB | Download |
| LFM2-350M-Q3_K_S.gguf | GGUF | Q3_K_S | 172.76 MB | Download |
| LFM2-350M-Q4_0.gguf | GGUF | — | 209.15 MB | Download |
| LFM2-350M-Q4_1.gguf | GGUF | — | 226.27 MB | Download |
| LFM2-350M-Q4_K_M.gguf | GGUF | Q4_K_M | 218.69 MB | Download |
| LFM2-350M-Q4_K_S.gguf | GGUF | Q4_K_S | 210.52 MB | Download |
| LFM2-350M-Q5_K_M.gguf | GGUF | Q5_K_M | 248.31 MB | Download |
| LFM2-350M-Q5_K_S.gguf | GGUF | Q5_K_S | 243.40 MB | Download |
| LFM2-350M-Q6_K.gguf | GGUF | Q6_K | 279.79 MB | Download |
| LFM2-350M-Q8_0.gguf | GGUF | — | 361.65 MB | Download |
| LFM2-350M-UD-Q2_K_XL.gguf | GGUF | Q2_K_XL | 153.15 MB | Download |
| LFM2-350M-UD-Q3_K_XL.gguf | GGUF | Q3_K_XL | 184.20 MB | Download |
| LFM2-350M-UD-Q4_K_XL.gguf | GGUF | Q4_K_XL | 218.69 MB | Download |
| LFM2-350M-UD-Q5_K_XL.gguf | GGUF | Q5_K_XL | 248.31 MB | Download |
| LFM2-350M-UD-Q6_K_XL.gguf | GGUF | Q6_K_XL | 295.29 MB | Download |
| LFM2-350M-UD-Q8_K_XL.gguf | GGUF | Q8_K_XL | 421.65 MB | Download |
Related Quantizations
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"base_model": [
"LiquidAI/LFM2-350M"
],
"library_name": "transformers",
"license": "other",
"license_name": "lfm1.0",
"license_link": "LICENSE",
"language": [
"en",
"ar",
"zh",
"fr",
"de",
"ja",
"ko",
"es"
],
"pipeline_tag": "text-generation",
"tags": [
"liquid",
"unsloth",
"lfm2",
"edge"
],
"frontmatter": {
"base_model": [
"LiquidAI/LFM2-350M"
],
"library_name": "transformers",
"license": "other",
"license_name": "lfm1.0",
"license_link": "LICENSE",
"language": [
"en",
"ar",
"zh",
"fr",
"de",
"ja",
"ko",
"es"
],
"pipeline_tag": "text-generation",
"tags": [
"liquid",
"unsloth",
"lfm2",
"edge"
]
},
"hero_image_url": "https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png",
"summary": "LFM2 is a new generation of hybrid models developed by Liquid AI, specifically designed for edge AI and on-device deployment. It sets a new standard in terms of quality, speed, and memory efficiency. We're releasing the weights of three post-trained checkpoints with 350M, 700M, and 1.2B parameters. They provide the following key features to create AI-powered edge applications: * **Fast training & inference** – LFM2 achieves 3x faster training compared to its previous generation. It also benefits from 2x faster decode and prefill speed on CPU compared to Qwen3. * **Best performance** – LFM2 outperforms similarly-sized models across multiple benchmark categories, including knowledge, mathematics, instruction following, and multilingual capabilities. * **New architecture** – LFM2 is a new hybrid Liquid model with multiplicative gates and short convolutions. * **Flexible deployment** – LFM2 runs efficiently on CPU, GPU, and NPU hardware for flexible deployment on smartphones, laptops, or vehicles. Find more information about LFM2 in our blog post.",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nbase_model:\n- LiquidAI/LFM2-350M\nlibrary_name: transformers\nlicense: other\nlicense_name: lfm1.0\nlicense_link: LICENSE\nlanguage:\n- en\n- ar\n- zh\n- fr\n- de\n- ja\n- ko\n- es\npipeline_tag: text-generation\ntags:\n- liquid\n- unsloth\n- lfm2\n- edge\n---\n> [!NOTE]\n> Includes our **chat template fixes**! <br> For `llama.cpp`, use `--jinja`\n>\n\n<div>\n<p style=\"margin-top: 0;margin-bottom: 0;\">\n <em><a href=\"https://docs.unsloth.ai/basics/unsloth-dynamic-v2.0-gguf\">Unsloth Dynamic 2.0</a> achieves superior accuracy & outperforms other leading quants.</em>\n </p>\n <div style=\"display: flex; gap: 5px; align-items: center; \">\n <a href=\"https://github.com/unslothai/unsloth/\">\n <img src=\"https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png\" width=\"133\">\n </a>\n <a href=\"https://discord.gg/unsloth\">\n <img src=\"https://github.com/unslothai/unsloth/raw/main/images/Discord%20button.png\" width=\"173\">\n </a>\n <a href=\"https://docs.unsloth.ai/\">\n <img src=\"https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/documentation%20green%20button.png\" width=\"143\">\n </a>\n </div>\n</div>\n\n\n<center>\n<div style=\"text-align: center;\">\n <img \n src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/7_6D7rWrLxp2hb6OHSV1p.png\" \n alt=\"Liquid AI\"\n style=\"width: 100%; max-width: 66%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;\"\n />\n</div>\n\n<a href=\"https://playground.liquid.ai/chat\">\n<svg width=\"114.8\" height=\"20\" viewBox=\"0 0 1300 200\" xmlns=\"http://www.w3.org/2000/svg\" role=\"img\" aria-label=\"Liquid Playground\" style=\"margin-bottom: 1em;\">\n <title>Liquid: Playground</title>\n <g>\n <rect fill=\"#fff\" width=\"600\" height=\"200\"></rect>\n <rect fill=\"url(#x)\" x=\"600\" width=\"700\" height=\"200\"></rect>\n </g>\n <g transform=\"translate(20, 30) scale(0.4, 0.4)\">\n <path d=\"M172.314 129.313L172.219 129.367L206.125 188.18C210.671 195.154 213.324 203.457 213.324 212.382C213.324 220.834 210.956 228.739 206.839 235.479L275.924 213.178L167.853 33.6L141.827 76.9614L172.314 129.313Z\" fill=\"black\"/>\n <path d=\"M114.217 302.4L168.492 257.003C168.447 257.003 168.397 257.003 168.352 257.003C143.515 257.003 123.385 237.027 123.385 212.387C123.385 203.487 126.023 195.204 130.55 188.24L162.621 132.503L135.966 86.7327L60.0762 213.183L114.127 302.4H114.217Z\" fill=\"black\"/>\n <path d=\"M191.435 250.681C191.435 250.681 191.43 250.681 191.425 250.686L129.71 302.4H221.294L267.71 226.593L191.435 250.686V250.681Z\" fill=\"black\"/>\n </g>\n <g aria-hidden=\"true\" fill=\"#fff\" text-anchor=\"start\" font-family=\"Verdana,DejaVu Sans,sans-serif\" font-size=\"110\">\n <text x=\"200\" y=\"148\" textLength=\"329\" fill=\"#000\" opacity=\"0.1\">Liquid</text>\n <text x=\"190\" y=\"138\" textLength=\"329\" fill=\"#000\">Liquid</text>\n <text x=\"655\" y=\"148\" textLength=\"619\" fill=\"#000\" opacity=\"0.1\">Playground</text>\n <text x=\"645\" y=\"138\" textLength=\"619\">Playground</text>\n </g>\n \n <linearGradient id=\"x\" x1=\"0%\" y1=\"0%\" x2=\"100%\" y2=\"0%\">\n <stop offset=\"0%\" style=\"stop-color:#000000\"></stop>\n <stop offset=\"100%\" style=\"stop-color:#000000\"></stop>\n </linearGradient>\n</svg>\n</a>\n</center>\n\n# LFM2-350M\n\nLFM2 is a new generation of hybrid models developed by [Liquid AI](https://www.liquid.ai/), specifically designed for edge AI and on-device deployment. It sets a new standard in terms of quality, speed, and memory efficiency. \n\nWe're releasing the weights of three post-trained checkpoints with 350M, 700M, and 1.2B parameters. They provide the following key features to create AI-powered edge applications:\n\n* **Fast training & inference** – LFM2 achieves 3x faster training compared to its previous generation. It also benefits from 2x faster decode and prefill speed on CPU compared to Qwen3.\n* **Best performance** – LFM2 outperforms similarly-sized models across multiple benchmark categories, including knowledge, mathematics, instruction following, and multilingual capabilities.\n* **New architecture** – LFM2 is a new hybrid Liquid model with multiplicative gates and short convolutions.\n* **Flexible deployment** – LFM2 runs efficiently on CPU, GPU, and NPU hardware for flexible deployment on smartphones, laptops, or vehicles.\n\nFind more information about LFM2 in our [blog post](https://www.liquid.ai/blog/liquid-foundation-models-v2-our-second-series-of-generative-ai-models).\n\n## 📄 Model details\n\nDue to their small size, **we recommend fine-tuning LFM2 models on narrow use cases** to maximize performance. \nThey are particularly suited for agentic tasks, data extraction, RAG, creative writing, and multi-turn conversations. \nHowever, we do not recommend using them for tasks that are knowledge-intensive or require programming skills.\n\n| Property | Value |\n| ------------------- | ----------------------------- |\n| **Parameters** | 354,483,968 |\n| **Layers** | 16 (10 conv + 6 attn) |\n| **Context length** | 32,768 tokens |\n| **Vocabulary size** | 65,536 |\n| **Precision** | bfloat16 |\n| **Training budget** | 10 trillion tokens |\n| **License** | LFM Open License v1.0 |\n\n**Supported languages**: English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.\n\n**Generation parameters**: We recommend the following parameters:\n* `temperature=0.3`\n* `min_p=0.15`\n* `repetition_penalty=1.05`\n\n**Chat template**: LFM2 uses a ChatML-like chat template as follows:\n\n```\n<|startoftext|><|im_start|>system\nYou are a helpful assistant trained by Liquid AI.<|im_end|>\n<|im_start|>user\nWhat is C. elegans?<|im_end|>\n<|im_start|>assistant\nIt's a tiny nematode that lives in temperate soil environments.<|im_end|>\n```\n\nYou can apply it using the dedicated [`.apply_chat_template()`](https://huggingface.co/docs/transformers/en/chat_templating#applychattemplate) function from Hugging Face transformers.\n\n**Tool use**: It consists of four main steps:\n1. **Function definition**: LFM2 takes JSON function definitions as input (JSON objects between `<|tool_list_start|>` and `<|tool_list_end|>` special tokens), usually in the system prompt\n2. **Function call**: LFM2 writes Pythonic function calls (a Python list between `<|tool_call_start|>` and `<|tool_call_end|>` special tokens), as the assistant answer.\n3. **Function execution**: The function call is executed and the result is returned (string between `<|tool_response_start|>` and `<|tool_response_end|>` special tokens), as a \"tool\" role.\n4. **Final answer**: LFM2 interprets the outcome of the function call to address the original user prompt in plain text.\n\nHere is a simple example of a conversation using tool use:\n\n```\n<|startoftext|><|im_start|>system\nList of tools: <|tool_list_start|>[{\"name\": \"get_candidate_status\", \"description\": \"Retrieves the current status of a candidate in the recruitment process\", \"parameters\": {\"type\": \"object\", \"properties\": {\"candidate_id\": {\"type\": \"string\", \"description\": \"Unique identifier for the candidate\"}}, \"required\": [\"candidate_id\"]}}]<|tool_list_end|><|im_end|>\n<|im_start|>user\nWhat is the current status of candidate ID 12345?<|im_end|>\n<|im_start|>assistant\n<|tool_call_start|>[get_candidate_status(candidate_id=\"12345\")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>\n<|im_start|>tool\n<|tool_response_start|>{\"candidate_id\": \"12345\", \"status\": \"Interview Scheduled\", \"position\": \"Clinical Research Associate\", \"date\": \"2023-11-20\"}<|tool_response_end|><|im_end|>\n<|im_start|>assistant\nThe candidate with ID 12345 is currently in the \"Interview Scheduled\" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>\n```\n\n**Architecture**: Hybrid model with multiplicative gates and short convolutions: 10 double-gated short-range LIV convolution blocks and 6 grouped query attention (GQA) blocks.\n\n**Pre-training mixture**: Approximately 75% English, 20% multilingual, and 5% code data sourced from the web and licensed materials.\n\n**Training approach**:\n* Knowledge distillation using [LFM1-7B](https://www.liquid.ai/blog/introducing-lfm-7b-setting-new-standards-for-efficient-language-models) as teacher model\n* Very large-scale SFT on 50% downstream tasks, 50% general domains\n* Custom DPO with length normalization and semi-online datasets\n* Iterative model merging\n\n## 🏃 How to run LFM2\n\nTo run LFM2, you need to install Hugging Face [`transformers`](https://github.com/huggingface/transformers) from source (v4.54.0.dev0).\nYou can update or install it with the following command: `pip install \"transformers @ git+https://github.com/huggingface/transformers.git@main\"`.\n\nHere is an example of how to generate an answer with transformers in Python:\n\n```python\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\n\n# Load model and tokenizer\nmodel_id = \"LiquidAI/LFM2-350M\"\nmodel = AutoModelForCausalLM.from_pretrained(\n model_id,\n device_map=\"auto\",\n torch_dtype=\"bfloat16\",\n trust_remote_code=True,\n# attn_implementation=\"flash_attention_2\" <- uncomment on compatible GPU\n)\ntokenizer = AutoTokenizer.from_pretrained(model_id)\n\n# Generate answer\nprompt = \"What is C. elegans?\"\ninput_ids = tokenizer.apply_chat_template(\n [{\"role\": \"user\", \"content\": prompt}],\n add_generation_prompt=True,\n return_tensors=\"pt\",\n tokenize=True,\n).to(model.device)\n\noutput = model.generate(\n input_ids,\n do_sample=True,\n temperature=0.3,\n min_p=0.15,\n repetition_penalty=1.05,\n max_new_tokens=512,\n)\n\nprint(tokenizer.decode(output[0], skip_special_tokens=False))\n\n# <|startoftext|><|im_start|>user\n# What is C. elegans?<|im_end|>\n# <|im_start|>assistant\n# C. elegans, also known as Caenorhabditis elegans, is a small, free-living\n# nematode worm (roundworm) that belongs to the phylum Nematoda.\n```\n\nYou can directly run and test the model with this [Colab notebook](https://colab.research.google.com/drive/1_q3jQ6LtyiuPzFZv7Vw8xSfPU5FwkKZY?usp=sharing).\n\n## 🔧 How to fine-tune LFM2\n\nWe recommend fine-tuning LFM2 models on your use cases to maximize performance.\n\n| Notebook | Description | Link |\n|-------|------|------|\n| SFT + LoRA | Supervised Fine-Tuning (SFT) notebook with a LoRA adapter in TRL. | <a href=\"https://colab.research.google.com/drive/1j5Hk_SyBb2soUsuhU0eIEA9GwLNRnElF?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"120\" alt=\"Colab link\"></a> |\n| DPO | Preference alignment with Direct Preference Optimization (DPO) in TRL. | <a href=\"https://colab.research.google.com/drive/1MQdsPxFHeZweGsNx4RH7Ia8lG8PiGE1t?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"120\" alt=\"Colab link\"></a> |\n\n## 📈 Performance\n\nLFM2 outperforms similar-sized models across different evaluation categories.\n\n### 1. Automated benchmarks\n\n\n\n| Model | MMLU | GPQA | IFEval | IFBench | GSM8K | MGSM | MMMLU |\n|-------|------|------|--------|---------|-------|------|-------|\n| LFM2-350M | 43.43 | 27.46 | 65.12 | 16.41 | 30.1 | 29.52 | 37.99 |\n| LFM2-700M | 49.9 | 28.48 | 72.23 | 20.56 | 46.4 | 45.36 | 43.28 |\n| LFM2-1.2B | *55.23* | **31.47** | **74.89** | *20.7* | *58.3* | *55.04* | **46.73** |\n| Qwen3-0.6B | 44.93 | 22.14 | 64.24 | 19.75 | 36.47 | 41.28 | 30.84 |\n| Qwen3-1.7B | **59.11** | 27.72 | *73.98* | **21.27** | 51.4 | **66.56** | *46.51* |\n| Llama-3.2-1B-Instruct | 46.6 | *28.84* | 52.39 | 16.86 | 35.71 | 29.12 | 38.15 |\n| gemma-3-1b-it | 40.08 | 21.07 | 62.9 | 17.72 | **59.59** | 43.6 | 34.43 |\n\n### 2. LLM-as-a-Judge\n\n\n\n\n### 3. Inference\n\n#### Throughput comparison on CPU in ExecuTorch\n\n\n\n#### Throughput comparison on CPU in Llama.cpp\n\n\n\n## 📬 Contact\n\nIf you are interested in custom solutions with edge deployment, please contact [our sales team](https://www.liquid.ai/contact).\n",
"related_quantizations": [
{
"id": "Felladrin/gguf-sharded-Q4_K_S-LFM2-350M",
"author": "Felladrin",
"downloads": null,
"likes": null,
"pipeline_tag": "",
"library_name": "",
"tags": []
}
]
},
"tags": [
"transformers",
"gguf",
"liquid",
"unsloth",
"lfm2",
"edge",
"text-generation",
"en",
"ar",
"zh",
"fr",
"de",
"ja",
"ko",
"es",
"base_model:LiquidAI/LFM2-350M",
"base_model:quantized:LiquidAI/LFM2-350M",
"license:other",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 17,
"downloads": 1092,
"gated": false,
"private": false,
"last_modified": "2025-07-14T12:19:36.000Z",
"created_at": "2025-07-11T19:57:27.000Z",
"pipeline_tag": "text-generation",
"library_name": "transformers"
}
Source payload excerpt (from Hugging Face API)
{
"_id": "68716ca74d0e077c4d0a6f46",
"id": "unsloth/LFM2-350M-GGUF",
"modelId": "unsloth/LFM2-350M-GGUF",
"sha": "d70a1c657baf5aff4a2258488fa8c5884b595bb6",
"createdAt": "2025-07-11T19:57:27.000Z",
"lastModified": "2025-07-14T12:19:36.000Z",
"author": "unsloth",
"downloads": 1092,
"likes": 17,
"gated": false,
"private": false,
"pipeline_tag": "text-generation",
"library_name": "transformers",
"siblings_count": 21
}