Professional Data EngineerBuilding and operationalizing data processing systemsHard
A data engineer is tasked with migrating an on-premises data processing pipeline to Google Cloud. The existing pipeline uses custom scripts written in Python and Bash to extract data from various databases, transform it, and load it into a data warehouse. The team wants to keep their existing code largely intact but needs a scalable, serverless environment to run these scripts without managing virtual machines. They also need to ensure that the environment can handle fluctuating workloads and integrate with other Google Cloud services. Which Google Cloud service should they choose to run their existing scripts?
- ACloud Run
- BCompute Engine
- CCloud Dataflow
- DCloud Dataproc
Show answer & explanationAnswer & explanation
Correct answer: A. Cloud Run
Cloud Run is a fully managed, serverless platform for containerized applications. It allows you to deploy and run stateless containers that can execute custom Python and Bash scripts, scale automatically with demand, and integrate easily with other Google Cloud services without managing VMs. This aligns perfectly with the requirement to lift-and-shift existing scripts into a serverless environment.
Why the other options are wrong
- B. Compute Engine requires managing VMs, which contradicts the 'without managing virtual machines' requirement.
- C. Cloud Dataflow is for running Apache Beam pipelines, requiring rewriting custom scripts into Beam, not running existing Python/Bash scripts directly.
- D. Cloud Dataproc is a managed Apache Spark/Hadoop service, suitable for big data frameworks, but not for running arbitrary custom Python/Bash scripts in a serverless manner.
Cloud Run
A fully managed serverless platform for running containerized applications, supporting any language or library.
- Scales automatically from zero to millions of requests
- Pay-per-use billing
- Ideal for stateless services and custom scripts
Memory trick: Run your scripts in a container, serverless and free.