GraySoft
Projects Models About FAQ Contact Download guIDE →
Model Intelligence Sheet

unsloth/lfm2.5-1.2b-instruct-gguf overview

LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning. !image Find more information about LFM2.5 in our blog post.

transformersggufliquidunslothlfm2.5edgetext-generationenarzhfrdejakoesarxiv:2511.23404base_model:LiquidAI/LFM2.5-1.2B-Instructbase_model:quantized:LiquidAI/LFM2.5-1.2B-Instructlicense:otherendpoints_compatibleregion:usconversational
unsloth/lfm2.5-1.2b-instruct-gguf visual
Downloads
4,975
Likes
38
Pipeline
text-generation
Library
transformers
Visibility
Public
Access
Open

Repository Files & Downloads

19 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
LFM2.5-1.2B-Instruct-BF16.gguf GGUF BF16 2.18 GB Download
LFM2.5-1.2B-Instruct-Q2_K.gguf GGUF Q2_K 461.01 MB Download
LFM2.5-1.2B-Instruct-Q2_K_L.gguf GGUF Q2_K_L 461.01 MB Download
LFM2.5-1.2B-Instruct-Q3_K_M.gguf GGUF Q3_K_M 572.54 MB Download
LFM2.5-1.2B-Instruct-Q3_K_S.gguf GGUF Q3_K_S 532.30 MB Download
LFM2.5-1.2B-Instruct-Q4_0.gguf GGUF 663.52 MB Download
LFM2.5-1.2B-Instruct-Q4_1.gguf GGUF 725.27 MB Download
LFM2.5-1.2B-Instruct-Q4_K_M.gguf GGUF Q4_K_M 697.04 MB Download
LFM2.5-1.2B-Instruct-Q4_K_S.gguf GGUF Q4_K_S 668.02 MB Download
LFM2.5-1.2B-Instruct-Q5_K_M.gguf GGUF Q5_K_M 804.29 MB Download
LFM2.5-1.2B-Instruct-Q5_K_S.gguf GGUF Q5_K_S 787.02 MB Download
LFM2.5-1.2B-Instruct-Q6_K.gguf GGUF Q6_K 918.24 MB Download
LFM2.5-1.2B-Instruct-Q8_0.gguf GGUF 1.16 GB Download
LFM2.5-1.2B-Instruct-UD-Q2_K_XL.gguf GGUF Q2_K_XL 461.01 MB Download
LFM2.5-1.2B-Instruct-UD-Q3_K_XL.gguf GGUF Q3_K_XL 572.54 MB Download
LFM2.5-1.2B-Instruct-UD-Q4_K_XL.gguf GGUF Q4_K_XL 697.04 MB Download
LFM2.5-1.2B-Instruct-UD-Q5_K_XL.gguf GGUF Q5_K_XL 804.29 MB Download
LFM2.5-1.2B-Instruct-UD-Q6_K_XL.gguf GGUF Q6_K_XL 949.24 MB Download
LFM2.5-1.2B-Instruct-UD-Q8_K_XL.gguf GGUF Q8_K_XL 1.28 GB Download

Model Details Live

Model Slug
unsloth/lfm2.5-1.2b-instruct-gguf
Author
unsloth
Pipeline Task
text-generation
Library
transformers
Created
2026-01-06
Last Modified
2026-01-06
Gated
No
Private
No
HF SHA
bf1ebe055f24ddd24f3622d933a63b42606773f3
License
other
Language
en, ar, zh, fr, de, ja, ko, es
Base Model
LiquidAI/LFM2.5-1.2B-Instruct

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "library_name": "transformers",
    "license": "other",
    "license_name": "lfm1.0",
    "license_link": "LICENSE",
    "language": [
      "en",
      "ar",
      "zh",
      "fr",
      "de",
      "ja",
      "ko",
      "es"
    ],
    "pipeline_tag": "text-generation",
    "tags": [
      "liquid",
      "unsloth",
      "lfm2.5",
      "edge"
    ],
    "base_model": [
      "LiquidAI/LFM2.5-1.2B-Instruct"
    ],
    "frontmatter": {
      "library_name": "transformers",
      "license": "other",
      "license_name": "lfm1.0",
      "license_link": "LICENSE",
      "language": [
        "en",
        "ar",
        "zh",
        "fr",
        "de",
        "ja",
        "ko",
        "es"
      ],
      "pipeline_tag": "text-generation",
      "tags": [
        "liquid",
        "unsloth",
        "lfm2.5",
        "edge"
      ],
      "base_model": [
        "LiquidAI/LFM2.5-1.2B-Instruct"
      ]
    },
    "hero_image_url": "https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png",
    "summary": "LFM2.5 is a new family of hybrid models designed for **on-device deployment**. It builds on the LFM2 architecture with extended pre-training and reinforcement learning. !image Find more information about LFM2.5 in our blog post.",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlibrary_name: transformers\nlicense: other\nlicense_name: lfm1.0\nlicense_link: LICENSE\nlanguage:\n- en\n- ar\n- zh\n- fr\n- de\n- ja\n- ko\n- es\npipeline_tag: text-generation\ntags:\n- liquid\n- unsloth\n- lfm2.5\n- edge\nbase_model:\n- LiquidAI/LFM2.5-1.2B-Instruct\n---\n> [!NOTE]\n>  Includes Unsloth **chat template fixes**! <br> For `llama.cpp`, use `--jinja`\n>\n\n<div>\n<p style=\"margin-top: 0;margin-bottom: 0;\">\n    <em><a href=\"https://docs.unsloth.ai/basics/unsloth-dynamic-v2.0-gguf\">Unsloth Dynamic 2.0</a> achieves superior accuracy & outperforms other leading quants.</em>\n  </p>\n  <div style=\"display: flex; gap: 5px; align-items: center; \">\n    <a href=\"https://github.com/unslothai/unsloth/\">\n      <img src=\"https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png\" width=\"133\">\n    </a>\n    <a href=\"https://discord.gg/unsloth\">\n      <img src=\"https://github.com/unslothai/unsloth/raw/main/images/Discord%20button.png\" width=\"173\">\n    </a>\n    <a href=\"https://docs.unsloth.ai/\">\n      <img src=\"https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/documentation%20green%20button.png\" width=\"143\">\n    </a>\n  </div>\n</div>\n\n\n<div align=\"center\">\n  <img \n    src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png\" \n    alt=\"Liquid AI\" \n    style=\"width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;\"\n  />\n  <div style=\"display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;\">\n    <a href=\"https://playground.liquid.ai/\"><strong>Try LFM</strong></a> • \n    <a href=\"https://docs.liquid.ai/lfm\"><strong>Documentation</strong></a> • \n    <a href=\"https://leap.liquid.ai/\"><strong>LEAP</strong></a>\n  </div>\n</div>\n\n# LFM2.5-1.2B-Instruct\n\nLFM2.5 is a new family of hybrid models designed for **on-device deployment**. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.\n\n- **Best-in-class performance**: A 1.2B model rivaling much larger models, bringing high-quality AI to your pocket.\n- **Fast edge inference**: 239 tok/s decode on AMD CPU, 82 tok/s on mobile NPU. Runs under 1GB of memory with day-one support for llama.cpp, MLX, and vLLM.\n- **Scaled training**: Extended pre-training from 10T to 28T tokens and large-scale multi-stage reinforcement learning.\n\n![image](https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/dxnYF2fuLpulismtFSGFi.png)\n\nFind more information about LFM2.5 in our [blog post](https://www.liquid.ai/blog/introducing-lfm2-5-the-next-generation-of-on-device-ai).\n\n## 🗒️ Model Details\n\n| Model | Parameters | Description |\n|-------|------------|-------------|\n| [LFM2.5-1.2B-Base](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Base) | 1.2B | Pre-trained base model for fine-tuning |\n| [**LFM2.5-1.2B-Instruct**](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct) | 1.2B | General-purpose instruction-tuned model |\n| [LFM2.5-1.2B-JP](https://huggingface.co/LiquidAI/LFM2.5-1.2B-JP) | 1.2B | Japanese-optimized chat model |\n| [LFM2.5-VL-1.6B](https://huggingface.co/LiquidAI/LFM2.5-VL-1.6B) | 1.6B | Vision-language model with fast inference |\n| [LFM2.5-Audio-1.5B](https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B) | 1.5B | Audio-language model for speech and text I/O |\n\nLFM2.5-1.2B-Instruct is a general-purpose text-only model with the following features:\n\n- **Number of parameters**: 1.17B\n- **Number of layers**: 16 (10 double-gated LIV convolution blocks + 6 GQA blocks)\n- **Training budget**: 28T tokens\n- **Context length**: 32,768 tokens\n- **Vocabulary size**: 65,536\n- **Languages**: English, Arabic, Chinese, French, German, Japanese, Korean, Spanish\n- **Generation parameters**:\n  - `temperature: 0.1`\n  - `top_k: 50`\n  - `top_p: 0.1`\n  - `repetition_penalty: 1.05`\n\n| Model | Description |\n|-------|-------------|\n| [**LFM2.5-1.2B-Instruct**](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct) | Original model checkpoint in native format. Best for fine-tuning or inference with Transformers and vLLM. |\n| [LFM2.5-1.2B-Instruct-GGUF](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF) | Quantized format for llama.cpp and compatible tools. Optimized for CPU inference and local deployment with reduced memory usage. |\n| [LFM2.5-1.2B-Instruct-ONNX](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-ONNX) | ONNX Runtime format for cross-platform deployment. Enables hardware-accelerated inference across diverse environments (cloud, edge, mobile). |\n\nWe recommend using it for agentic tasks, data extraction, and RAG. It is not recommended for knowledge-intensive tasks and programming.\n\n### Chat Template\n\nLFM2.5 uses a ChatML-like format. See the [Chat Template documentation](https://docs.liquid.ai/lfm/key-concepts/chat-template) for details. Example:\n\n```\n<|startoftext|><|im_start|>system\nYou are a helpful assistant trained by Liquid AI.<|im_end|>\n<|im_start|>user\nWhat is C. elegans?<|im_end|>\n<|im_start|>assistant\n```\n\nYou can use [`tokenizer.apply_chat_template()`](https://huggingface.co/docs/transformers/en/chat_templating#using-applychattemplate) to format your messages automatically.\n\n### Tool Use\n\nLFM2.5 supports function calling as follows:\n\n1. **Function definition**: We recommend providing the list of tools as a JSON object in the system prompt. You can also use the [`tokenizer.apply_chat_template()`](https://huggingface.co/docs/transformers/en/chat_extras#passing-tools) function with tools.\n2. **Function call**: By default, LFM2.5 writes Pythonic function calls (a Python list between `<|tool_call_start|>` and `<|tool_call_end|>` special tokens), as the assistant answer. You can override this behavior by asking the model to output JSON function calls in the system prompt.\n3. **Function execution**: The function call is executed, and the result is returned as a \"tool\" role.\n4. **Final answer**: LFM2 interprets the outcome of the function call to address the original user prompt in plain text.\n\nSee the [Tool Use documentation](https://docs.liquid.ai/lfm/key-concepts/tool-use) for the full guide. Example:\n\n```\n<|startoftext|><|im_start|>system\nList of tools: [{\"name\": \"get_candidate_status\", \"description\": \"Retrieves the current status of a candidate in the recruitment process\", \"parameters\": {\"type\": \"object\", \"properties\": {\"candidate_id\": {\"type\": \"string\", \"description\": \"Unique identifier for the candidate\"}}, \"required\": [\"candidate_id\"]}}]<|im_end|>\n<|im_start|>user\nWhat is the current status of candidate ID 12345?<|im_end|>\n<|im_start|>assistant\n<|tool_call_start|>[get_candidate_status(candidate_id=\"12345\")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>\n<|im_start|>tool\n[{\"candidate_id\": \"12345\", \"status\": \"Interview Scheduled\", \"position\": \"Clinical Research Associate\", \"date\": \"2023-11-20\"}]<|im_end|>\n<|im_start|>assistant\nThe candidate with ID 12345 is currently in the \"Interview Scheduled\" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>\n```\n\n## 🏃 Inference\n\nLFM2.5 is supported by many inference frameworks. See the [Inference documentation](https://docs.liquid.ai/lfm/inference/transformers) for the full list.\n\n| Name | Description | Docs | Notebook |\n|------|-------------|------|:--------:|\n| [Transformers](https://github.com/huggingface/transformers) | Simple inference with direct access to model internals. | <a href=\"https://docs.liquid.ai/lfm/inference/transformers\">Link</a> | <a href=\"https://colab.research.google.com/drive/1_q3jQ6LtyiuPzFZv7Vw8xSfPU5FwkKZY?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"110\" alt=\"Colab link\"></a> |\n| [vLLM](https://github.com/vllm-project/vllm) | High-throughput production deployments with GPU. | <a href=\"https://docs.liquid.ai/lfm/inference/vllm\">Link</a> | <a href=\"https://colab.research.google.com/drive/1VfyscuHP8A3we_YpnzuabYJzr5ju0Mit?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"110\" alt=\"Colab link\"></a> |\n| [llama.cpp](https://github.com/ggml-org/llama.cpp) | Cross-platform inference with CPU offloading. | <a href=\"https://docs.liquid.ai/lfm/inference/llama-cpp\">Link</a> | <a href=\"https://colab.research.google.com/drive/1ohLl3w47OQZA4ELo46i5E4Z6oGWBAyo8?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"110\" alt=\"Colab link\"></a> |\n| [MLX](https://github.com/ml-explore/mlx) | Apple's machine learning framework optimized for Apple Silicon. | <a href=\"https://docs.liquid.ai/lfm/inference/mlx\">Link</a> | — |\n| [LM Studio](https://lmstudio.ai/) | Desktop application for running LLMs locally. | <a href=\"https://docs.liquid.ai/lfm/inference/lm-studio\">Link</a> | — |\n\nHere's a quick start example with Transformers:\n\n```python\nfrom transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer\n\nmodel_id = \"LiquidAI/LFM2.5-1.2B-Instruct\"\nmodel = AutoModelForCausalLM.from_pretrained(\n    model_id,\n    device_map=\"auto\",\n    dtype=\"bfloat16\",\n#   attn_implementation=\"flash_attention_2\" <- uncomment on compatible GPU\n)\ntokenizer = AutoTokenizer.from_pretrained(model_id)\nstreamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)\n\nprompt = \"What is C. elegans?\"\n\ninput_ids = tokenizer.apply_chat_template(\n    [{\"role\": \"user\", \"content\": prompt}],\n    add_generation_prompt=True,\n    return_tensors=\"pt\",\n    tokenize=True,\n).to(model.device)\n\noutput = model.generate(\n    input_ids,\n    do_sample=True,\n    temperature=0.1,\n    top_k=50,\n    top_p=0.1,\n    repetition_penalty=1.05,\n    max_new_tokens=512,\n    streamer=streamer,\n)\n```\n\n## 🔧 Fine-Tuning\n\nWe recommend fine-tuning LFM2.5 for your specific use case to achieve the best results.\n\n| Name | Description | Docs | Notebook |\n|------|-------------|------|----------|\n| SFT ([Unsloth](https://github.com/unslothai/unsloth)) | Supervised Fine-Tuning with LoRA using Unsloth. | <a href=\"https://docs.liquid.ai/lfm/fine-tuning/unsloth\">Link</a> | <a href=\"https://colab.research.google.com/drive/1HROdGaPFt1tATniBcos11-doVaH7kOI3?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"110\" alt=\"Colab link\"></a> |\n| SFT ([TRL](https://github.com/huggingface/trl)) | Supervised Fine-Tuning with LoRA using TRL. | <a href=\"https://docs.liquid.ai/lfm/fine-tuning/trl\">Link</a> | <a href=\"https://colab.research.google.com/drive/1j5Hk_SyBb2soUsuhU0eIEA9GwLNRnElF?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"110\" alt=\"Colab link\"></a> |\n| DPO ([TRL](https://github.com/huggingface/trl)) | Direct Preference Optimization with LoRA using TRL. | <a href=\"https://docs.liquid.ai/lfm/fine-tuning/trl\">Link</a> | <a href=\"https://colab.research.google.com/drive/1MQdsPxFHeZweGsNx4RH7Ia8lG8PiGE1t?usp=sharing\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHa_b_LXysEu2E.png\" width=\"110\" alt=\"Colab link\"></a> |\n\n## 📊 Performance\n\n### Benchmarks\n\nWe compared LFM2.5-1.2B-Instruct with relevant sub-2B models on a diverse suite of benchmarks.\n\n| Model | GPQA | MMLU-Pro | IFEval | IFBench | Multi-IF | AIME25 | BFCLv3 |\n|-------|------|----------|--------|---------|----------|--------|--------|\n| **LFM2.5-1.2B-Instruct** | 38.89 | 44.35 | 86.23 | 47.33 | 60.98 | 14.00 | 49.12 |\n| Qwen3-1.7B | 34.85 | 42.91 | 73.68 | 21.33 | 56.48 | 9.33 | 46.30 |\n| Granite 4.0-1B | 24.24 | 33.53 | 79.61 | 21.00 | 43.65 | 3.33 | 52.43 |\n| Llama 3.2 1B Instruct | 16.57 | 20.80 | 52.37 | 15.93 | 30.16 | 0.33 | 21.44 |\n| Gemma 3 1B IT | 24.24 | 14.04 | 63.25 | 20.47 | 44.31 | 1.00 | 16.64 |\n\nGPQA, MMLU-Pro, IFBench, and AIME25 follow [ArtificialAnalysis's methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking). For IFEval and Multi-IF, we report the average score across strict and loose prompt and instruction accuracies. For BFCLv3, we report the final weighted average score with a custom Liquid handler to support our tool use template. \n\n### Inference speed\n\nLFM2.5-1.2B-Instruct offers extremely fast inference speed on CPUs with a low memory profile compared to similar-sized models.\n\n![image](https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/dbbI-15p9re2ROhAkqnZm.png)\n\nIn addition, we are partnering with AMD, Qualcomm, and Nexa AI to bring the LFM2.5 family to NPUs. These optimized models are available through our partners, enabling highly efficient on-device inference.\n\n| Device                                               | Inference | Framework        | Model                | Prefill (tok/s) | Decode (tok/s) | Memory (GB) |\n| ---------------------------------------------------- | --------- | ---------------- | -------------------- | --------------- | -------------- | ----------- |\n| Qualcomm Snapdragon® X Elite                         | NPU       | NexaML           | LFM2.5-1.2B-instruct | 2591            | 63             | 0.9GB       |\n| Qualcomm Snapdragon® Gen4 (ROG Phone9 Pro)           | NPU       | NexaML           | LFM2.5-1.2B-instruct | 4391            | 82             | 0.9GB       |\n| Qualcomm Snapdragon® Gen4 (Samsung Galaxy S25 Ultra) | CPU       | llama.cpp (Q4_0) | LFM2.5-1.2B-instruct | 335             | 70             | 719MB       |\n| Qualcomm Snapdragon® Gen4 (Samsung Galaxy S25 Ultra) | CPU       | llama.cpp (Q4_0) | Qwen3-1.7B           | 181             | 40             | 1306MB      |\n\nThese capabilities unlock new deployment scenarios across various devices, including vehicles, mobile devices, laptops, IoT devices, and embedded systems.\n\n## Contact\n\nFor enterprise solutions and edge deployment, contact [sales@liquid.ai](mailto:sales@liquid.ai).\n\n## Citation\n\n```bibtex\n@article{liquidai2025lfm2,\n  title={LFM2 Technical Report},\n  author={Liquid AI},\n  journal={arXiv preprint arXiv:2511.23404},\n  year={2025}\n}\n```",
    "related_quantizations": []
  },
  "tags": [
    "transformers",
    "gguf",
    "liquid",
    "unsloth",
    "lfm2.5",
    "edge",
    "text-generation",
    "en",
    "ar",
    "zh",
    "fr",
    "de",
    "ja",
    "ko",
    "es",
    "arxiv:2511.23404",
    "base_model:LiquidAI/LFM2.5-1.2B-Instruct",
    "base_model:quantized:LiquidAI/LFM2.5-1.2B-Instruct",
    "license:other",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 38,
  "downloads": 4975,
  "gated": false,
  "private": false,
  "last_modified": "2026-01-06T08:28:52.000Z",
  "created_at": "2026-01-06T05:18:51.000Z",
  "pipeline_tag": "text-generation",
  "library_name": "transformers"
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "695c9b3b5478be093b12b3c2",
  "id": "unsloth/LFM2.5-1.2B-Instruct-GGUF",
  "modelId": "unsloth/LFM2.5-1.2B-Instruct-GGUF",
  "sha": "bf1ebe055f24ddd24f3622d933a63b42606773f3",
  "createdAt": "2026-01-06T05:18:51.000Z",
  "lastModified": "2026-01-06T08:28:52.000Z",
  "author": "unsloth",
  "downloads": 4975,
  "likes": 38,
  "gated": false,
  "private": false,
  "pipeline_tag": "text-generation",
  "library_name": "transformers",
  "siblings_count": 22
}