Ferrum 0.8.4 — a single-binary Rust LLM runtime for Metal and CUDA

· Users Rust Lang · Sept. 2, 2026, 3:23 a.m.
Summary
Ferrum 0.8.4 is an MIT-licensed local LLM inference runtime written in Rust, designed for easy installation and serving of local LLMs using a single binary without the need for Python or complex libraries. It supports Apple Silicon with Metal acceleration and NVIDIA GPUs with CUDA, along with an API compatible with OpenAI.
AUTHOR