Batch inference on OpenShift AI with Ray Data, vLLM, and CodeFlare

· Red Hat · Aug. 7, 2025, 7:08 a.m.
Summary
This blog post addresses how to efficiently perform large-scale batch inference using the CodeFlare SDK with Ray Data and vLLM on OpenShift AI. It highlights the distinction between online and offline inference, and offers a step-by-step guide for setting up a remote batch inference job, detailing configuration, model sourcing, performance tuning, and job submission to a Ray cluster. It emphasizes that data scientists can leverage these tools without deep infrastructure knowledge.
AUTHOR
Sponsored
Zulip logo Zulip
Organized team chat for people who take work seriously. Topic-based threading keeps conversations focused.
Try Zulip
Become a sponsor →
BLOG POST FEATURED ON

Add this plugin to your blog