Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium
A company is deploying a new application that uses a custom machine learning model. The model requires GPU acceleration for inference and needs to be deployed as a highly available and scalable service. They prefer to manage the underlying infrastructure as little as possible. Which Google Cloud compute service is the most appropriate choice?
- AGoogle Kubernetes Engine (GKE) with GPU nodes
- BCloud Run with GPU-enabled containers
- CCompute Engine with custom GPU VMs
- DCloud AI Platform Prediction
Show answer & explanationAnswer & explanation
Correct answer: D. Cloud AI Platform Prediction
Cloud AI Platform Prediction is a fully managed service specifically designed for deploying machine learning models, including those requiring GPUs. It handles scaling, availability, and infrastructure management, minimizing operational overhead.
Why the other options are wrong
- A. GKE with GPU nodes offers scalability and container orchestration but still involves managing a Kubernetes cluster and its nodes, not fully minimizing infrastructure management for ML inference.
- B. Cloud Run currently does not directly support GPU acceleration for inference, making it unsuitable for this specific requirement.
- C. Compute Engine provides flexibility but requires significant management of VMs, GPUs, and scaling, which goes against 'manage infrastructure as little as possible'.
Google Cloud AI Platform Prediction
A fully managed service for deploying machine learning models into production at scale, handling infrastructure, scaling, and model serving.
- Supports models from various ML frameworks.
- Manages infrastructure, scaling, and high availability.
- Can leverage GPUs for accelerated inference.
Memory trick: AI Platform is the 'AI' for 'platform' deployment.