Google Associate Cloud EngineerDeploying and implementing a cloud solutionHard
A client needs to deploy a custom machine learning model for real-time inference. The model is packaged as a Docker container and requires GPU acceleration for optimal performance. They also need to monitor the model's performance and manage different model versions. Which Google Cloud service should be used?
- AGoogle Kubernetes Engine (GKE) with GPU nodes
- BCloud Run with GPU acceleration enabled
- CCompute Engine instances with GPUs and custom deployment scripts
- DVertex AI Prediction
Show answer & explanationAnswer & explanation
Correct answer: D. Vertex AI Prediction
Vertex AI Prediction is a fully managed service designed for deploying and serving machine learning models, including those packaged as Docker containers and requiring GPU acceleration. It provides built-in features for model monitoring, version management, and auto-scaling, significantly reducing operational overhead compared to self-managing infrastructure.
Why the other options are wrong
- A. GKE can host containers with GPUs, but Vertex AI Prediction offers a higher-level, ML-specific managed service with integrated features like monitoring and versioning.
- B. Cloud Run currently does not support GPU acceleration for inference.
- C. Compute Engine requires significant manual effort for deployment, scaling, monitoring, and version management of ML models with GPUs.
Vertex AI Prediction
A fully managed service within Google Cloud's Vertex AI platform for deploying and serving machine learning models for online (real-time) predictions.
- Supports custom containers for model deployment.
- Can leverage GPU acceleration for inference.
- Provides model monitoring and version management.
- Offers auto-scaling based on traffic patterns.
- Integrates with other Vertex AI services for MLOps.
Memory trick: Vertex AI Prediction: It's the 'peak' for putting your machine learning models to work, smart and fast.