Professional Data EngineerOperationalizing machine learning modelsMedium
A data science team has developed a new image classification model. They want to deploy this model to serve online predictions with low latency. The model is packaged as a custom container image. They anticipate highly variable traffic patterns, with potential spikes, and need the deployment to scale automatically while minimizing costs during idle periods. Which Google Cloud deployment option is best suited for these requirements?
- ADeploy to a Compute Engine instance with a fixed VM size.
- BDeploy to Cloud Functions with GPU acceleration.
- CDeploy to Vertex AI Endpoints with autoscaling enabled.
- DDeploy to a Kubernetes Engine cluster with manual scaling.
Show answer & explanationAnswer & explanation
Correct answer: C. Deploy to Vertex AI Endpoints with autoscaling enabled.
Vertex AI Endpoints provides a fully managed service for deploying models, including support for custom containers, and offers robust autoscaling capabilities to handle variable traffic and scale down to zero during idle times, optimizing cost and performance.
Why the other options are wrong
- A. Fixed Compute Engine instances do not scale automatically and would be inefficient for variable traffic, leading to over-provisioning or under-provisioning.
- B. Cloud Functions are generally not designed for long-running, resource-intensive ML models, and while they can scale, they are not optimized for GPU-accelerated custom containers in the same way as Vertex AI Endpoints.
- D. Kubernetes Engine offers flexibility but requires more operational overhead for managing the cluster and configuring autoscaling compared to Vertex AI Endpoints.
Vertex AI Endpoints
A managed service within Vertex AI for deploying machine learning models to serve online predictions with high availability, low latency, and automatic scaling.
- Supports various model formats and custom containers.
- Provides managed autoscaling for variable traffic.
- Offers A/B testing and traffic splitting for model updates.
Memory trick: Vertex AI Endpoints: Your model's managed, auto-scaling launchpad for predictions.