How DigitalOcean’s Agentic Inference Cloud powered by NVIDIA GPUs Achieved 67% Lower Inference Costs for Workato

240 · DigitalOcean · March 3, 2026, 8:40 a.m.
Summary
This blog post details how Workato improved LLM inference efficiency by partnering with DigitalOcean to leverage NVIDIA's Dynamo framework. By implementing KV-aware routing, they achieved significant performance boosts in throughput and latency while reducing costs. The architecture focused on optimizing coordination between GPU resources, ultimately leading to a more effective inference stack that balances computational load, minimizes redundant processing, and retains cache efficiency. Key results include a 67% higher tokens per second per GPU and substantial reductions in inference costs.