Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium

A company is deploying a new application that uses a custom machine learning model. The model requires GPU acceleration for inference and needs to be deployed as a highly available and scalable service. They prefer to manage the underlying infrastructure as little as possible. Which Google Cloud compute service is the most appropriate choice?

  1. AGoogle Kubernetes Engine (GKE) with GPU nodes
  2. BCloud Run with GPU-enabled containers
  3. CCompute Engine with custom GPU VMs
  4. DCloud AI Platform Prediction
Show answer & explanation

Correct answer: D. Cloud AI Platform Prediction

Cloud AI Platform Prediction is a fully managed service specifically designed for deploying machine learning models, including those requiring GPUs. It handles scaling, availability, and infrastructure management, minimizing operational overhead.

Why the other options are wrong

  • A. GKE with GPU nodes offers scalability and container orchestration but still involves managing a Kubernetes cluster and its nodes, not fully minimizing infrastructure management for ML inference.
  • B. Cloud Run currently does not directly support GPU acceleration for inference, making it unsuitable for this specific requirement.
  • C. Compute Engine provides flexibility but requires significant management of VMs, GPUs, and scaling, which goes against 'manage infrastructure as little as possible'.

Google Cloud AI Platform Prediction

A fully managed service for deploying machine learning models into production at scale, handling infrastructure, scaling, and model serving.

  • Supports models from various ML frameworks.
  • Manages infrastructure, scaling, and high availability.
  • Can leverage GPUs for accelerated inference.

Memory trick: AI Platform is the 'AI' for 'platform' deployment.

More Deploying and implementing a cloud solution questions