Google Associate Cloud EngineerDeploying and implementing a cloud solutionHard
A client needs to deploy a custom machine learning model for real-time inference. The model is containerized and needs to be highly available, scalable, and deployed as a managed endpoint that can be accessed via a REST API. They also need integrated monitoring and logging for the deployed model. Which Google Cloud service should they use?
- AGoogle Kubernetes Engine (GKE)
- BVertex AI Prediction
- CCompute Engine
- DCloud Functions
Show answer & explanationAnswer & explanation
Correct answer: B. Vertex AI Prediction
Vertex AI Prediction is specifically designed for deploying machine learning models as managed endpoints for real-time inference, offering high availability, scalability, container support, and integrated monitoring/logging.
Why the other options are wrong
- A. GKE can host ML models but requires managing Kubernetes clusters and configuring scaling, monitoring, and API endpoints manually, unlike the fully managed Vertex AI Prediction service.
- C. Compute Engine would require manual setup and management of the inference server, scaling, and monitoring, which is not a 'managed endpoint'.
- D. Cloud Functions are for event-driven, short-lived tasks and are not optimized for continuous, high-performance real-time ML inference with complex models.
Vertex AI Prediction
A managed service within Vertex AI that enables deploying machine learning models as scalable, highly available endpoints for real-time online predictions.
- Supports custom containerized models.
- Provides automatic scaling and load balancing.
- Offers integrated monitoring, logging, and explainability features.
Memory trick: Vertex AI for managed ML, GKE for custom control, Functions for simple tasks.