Workload Example
An example of online model serving implemented with Ray Serve.
In this module, you’ll deploy a Hugging Face sentiment model as a scalable online inference service using Ray Serve with a FastAPI HTTP endpoint. You’ll learn how to run and scale the deployment with replicas, send test client requests, and properly shut down the Serve app and Ray cluster.