Modernising Grab’s model serving platform with NVIDIA Triton Inference Server

178 · Grab · Oct. 21, 2025, 4:13 a.m.
Summary
This blog post describes the modernization of Grab's machine learning model serving platform, Catwalk, using NVIDIA Triton Inference Server. The transition aims to improve performance, reduce latency, and manage costs efficiently by consolidating multiple inference engines into one. Key takeaways include significant enhancements in model serving capabilities, cost reductions, and the importance of a seamless migration for internal users. Overall, the implementation of Triton has greatly benefited the performance and stability of machine learning models at Grab, leading to a successful rollout with minimal impact on end users.