Agntro author hub
license: apache 2.0 base model: allenai/OLMoE 1B 7B 0924 library name: llama.cpp tags: gguf, moe, quantization, llama.cpp, tq2 0, quaternary Update 2026 07 : now GPU native. The TQ2 0 kernel is merged for CUDA / Metal / AMD / Vulkan — OLMoE runs GPU native ~384 t/s decode on an …
Models
2
Downloads
10
Run models locally with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.