Ferrum 0.8.4 — a single-binary Rust LLM runtime for Metal and CUDA

· Users Rust Lang · Sept. 2, 2026, 3:23 a.m.
Summary
Ferrum 0.8.4 is an MIT-licensed local LLM inference runtime written in Rust, designed for easy installation and serving of local LLMs using a single binary without the need for Python or complex libraries. It supports Apple Silicon with Metal acceleration and NVIDIA GPUs with CUDA, along with an API compatible with OpenAI.
AUTHOR
Sponsored
Zulip logo Zulip
Organized team chat for people who take work seriously. Topic-based threading keeps conversations focused.
Try Zulip
Become a sponsor →