Ferrum 0.8.4 is an MIT-licensed local LLM inference runtime written in Rust, designed for easy installation and serving of local LLMs using a single binary without the need for Python or complex libraries. It supports Apple Silicon with Metal acceleration and NVIDIA GPUs with CUDA, along with an API compatible with OpenAI.