Professional Data Engineer practice questions

262 free questions with answers and explanations.

Practice test
  1. 1.A financial institution needs to process daily batch files containing millions of customer transactions. These files arrive at a specific time each day and must be processed within a strict 2-hour window. The processing involves complex transformations and aggregations, which can be computationally intensive. The company wants to optimize the Dataflow job for both performance (completing within 2 hours) and cost, ensuring that resources are perfectly matched to the workload without over-provisioning. Which Dataflow Prime feature is specifically designed to address this challenge?Building and operationalizing data processing systems
  2. 2.A global ride-sharing company collects real-time GPS coordinates from millions of active vehicles. This data needs to be ingested into Google Cloud for immediate processing by a Dataflow streaming pipeline. The ingestion system must handle extreme, unpredictable spikes in message volume, scale automatically without manual intervention, and guarantee high durability to prevent data loss. Which Google Cloud service is best suited to ingest this data reliably and scalably?Building and operationalizing data processing systems
  3. 3.A global online gaming company needs to store petabytes of user gameplay data, including session logs, in-game events, and player statistics. This data is characterized by extremely high write throughput (millions of writes per second), low-latency reads for real-time leaderboards and player profiles, and requires horizontal scalability to handle unpredictable traffic spikes. Data consistency is eventually consistent for most use cases, but strong consistency is preferred where possible. Which Google Cloud database service is best suited for this workload?Building and operationalizing data processing systems
  4. 4.A financial institution is implementing a new data warehousing solution on Google Cloud. They store highly sensitive customer financial data in BigQuery. Due to regulatory requirements, they must ensure that data at rest is encrypted with customer-managed encryption keys (CMEK), and these keys must be automatically rotated annually. Which Google Cloud service should they use to manage and automatically rotate these encryption keys for BigQuery?Managing and securing data
  5. 5.A gaming company is experiencing intermittent spikes in latency and error rates in their game analytics pipeline, which uses Cloud Pub/Sub for ingestion and Dataflow for real-time processing. They need a comprehensive solution to quickly detect, diagnose, and alert on these issues across their entire data processing stack. Which combination of Google Cloud services would provide the best solution?Building and operationalizing data processing systems
  6. 6.A global manufacturing company uses BigQuery for its operational analytics. They have several datasets containing sensitive production data, including intellectual property (IP) related to product designs. They need to ensure that only specific engineering teams can view the 'design_specifications' column in the `production_designs` table, while other teams can still access other columns in the same table for general reporting. How should they implement this granular access control?Managing and securing data
  7. 7.A financial services company is building a data lake on Google Cloud. They need to store raw, semi-structured transaction data (e.g., JSON, CSV) from various sources, including on-premises databases and third-party APIs. The data volume is expected to grow to petabytes, and it will be accessed by different analytical tools (BigQuery, Spark) for ad-hoc queries and machine learning models. The solution must be highly durable, globally accessible, and cost-effective for long-term storage, with flexible schema support. Which Google Cloud storage service is most appropriate for their data lake's raw data storage layer?Building and operationalizing data processing systems
  8. 8.A data engineering team is developing a complex batch data pipeline that involves data ingestion from various sources (Cloud Storage, external APIs), transformation using custom Python scripts, and loading into BigQuery. The pipeline has multiple dependencies between stages, conditional logic based on data availability, and needs to be scheduled daily. The team wants a fully managed service that allows them to define, schedule, and monitor these workflows using a familiar open-source tool. Which Google Cloud service meets these requirements?Building and operationalizing data processing systems
  9. 9.A data team is developing a new streaming pipeline to process real-time sensor data from IoT devices. The pipeline uses Dataflow to perform aggregations and transformations before storing the results in BigQuery. They notice that during periods of high data volume, the Dataflow job's CPU utilization is consistently high, and the processing latency increases significantly. They have already increased the number of workers, but the problem persists. Upon inspecting the Dataflow UI, they see that a specific `DoFn` responsible for a complex, CPU-intensive calculation is the bottleneck. What is the most effective approach to optimize this bottleneck in the Dataflow pipeline?Building and operationalizing data processing systems
  10. 10.A data engineering team is building a new real-time analytics pipeline. During testing, they observe that some data events arrive out of order or with significant delays, leading to inaccurate windowed aggregations. They need to ensure that aggregations correctly account for all data within a specific time window, regardless of when events actually arrive. Which Apache Beam concept is crucial for handling this problem in Dataflow?Building and operationalizing data processing systems
  11. 11.A data analytics team is migrating an on-premises data warehouse to Google Cloud. They have petabytes of historical data stored in various formats (CSV, JSON, Parquet) that need to be loaded into BigQuery for analysis. The team wants to perform initial data cleaning and transformation before loading, but they also need to support incremental daily loads. They prefer a serverless approach that minimizes operational overhead. Which Google Cloud service combination would be most suitable for this migration and ongoing data loading?Building and operationalizing data processing systems
  12. 12.A data team is developing a new Dataflow pipeline that processes gigabytes of data hourly. After initial deployment, they notice that the pipeline is consistently running behind schedule, and the Dataflow monitoring interface shows high CPU utilization and low data freshness. The pipeline uses complex custom transformations written in Python. Which action should they take to improve the pipeline's performance?Building and operationalizing data processing systems
  13. 13.A data team is building a complex data pipeline that involves ingesting data from various sources (databases, APIs, files), performing transformations, and loading into BigQuery. The pipeline has dependencies between different stages, requires scheduling, error handling, and robust monitoring. They need to orchestrate this multi-step workflow in a managed, scalable, and fault-tolerant manner. Which Google Cloud service is the most appropriate for orchestrating this pipeline?Building and operationalizing data processing systems
  14. 14.A data scientist needs to train a machine learning model using a large dataset stored in BigQuery. The training process requires performing complex statistical aggregations and feature engineering steps that are best expressed using SQL, but the dataset is too large to fit into memory on a single machine. The data scientist wants to leverage the scalability of BigQuery for these data preparation steps before exporting the final features for model training in a separate environment. Which BigQuery feature is most suitable for this scenario?Building and operationalizing data processing systems
  15. 15.A data engineering team is building a real-time recommendation engine. They need to store user interaction data (e.g., clicks, views, purchases) and serve personalized recommendations with extremely low latency (sub-10ms) to millions of concurrent users. The data model is relatively simple, consisting of key-value pairs and wide-column structures. High write throughput and read throughput are critical. Which Google Cloud service is best suited for this operational database requirement?Building and operationalizing data processing systems
  16. 16.A financial institution is building a data warehouse in BigQuery. They need to ensure that data consumed by downstream analytical applications is trustworthy and that its origin and transformations can be tracked. Specifically, they need to visualize the path of data from its ingestion sources through various BigQuery transformations to its final destination tables. Which Google Cloud service provides this capability?Managing and securing data
  17. 17.A data engineering team is developing a new streaming pipeline to process real-time sensor data. The pipeline runs on Dataflow, and the processed data is eventually stored in BigQuery. The team needs to ensure that the Dataflow workers can securely communicate with BigQuery and other Google Cloud services without exposing credentials directly in the code or to the internet. They also want to restrict access based on the principle of least privilege. Which authentication and networking approach should they implement?Building and operationalizing data processing systems
  18. 18.A global e-commerce company needs to store petabytes of historical transaction data for auditing and compliance purposes. This data is accessed very infrequently, perhaps once or twice a year, but must be retained for at least 10 years. The company prioritizes minimizing storage costs while ensuring data durability and availability when needed. Which Google Cloud Storage class should be used for this data?Building and operationalizing data processing systems
  19. 19.A data engineer is implementing a custom data processing job on Compute Engine. This job needs to read data from a Cloud Storage bucket, process it, and then write the results to a BigQuery table. The engineer wants to grant the Compute Engine instance the necessary permissions securely and follow the principle of least privilege. Which Google Cloud identity should be associated with the Compute Engine instance to achieve this?Building and operationalizing data processing systems
  20. 20.A data engineering team is building a complex data pipeline that involves ingesting data from various sources, performing multiple transformation steps (filtering, aggregation, joining), and loading the results into BigQuery. The pipeline has dependencies between stages, conditional execution logic, and requires robust error handling and retry mechanisms. The team needs a fully managed service to define, schedule, and monitor these workflows, leveraging Python for custom logic. Which Google Cloud service is best suited for orchestrating this data pipeline?Building and operationalizing data processing systems
  21. 21.A data analytics team is developing a real-time anomaly detection system using Dataflow. The system processes sensor readings from IoT devices, identifies unusual patterns over rolling 5-minute windows, and publishes alerts. During testing, they observe that the system occasionally generates duplicate alerts for the same anomaly, especially after worker restarts or re-processing. The team needs to ensure that each anomaly is processed and alerted exactly once, even in the face of infrastructure failures or re-processing. Which Dataflow feature should they leverage to guarantee this behavior?Building and operationalizing data processing systems
  22. 22.A media streaming service is migrating its on-premises user activity database to Google Cloud. The database contains millions of user profiles and preferences, requiring low-latency reads and writes for real-time personalization features. The data schema is highly flexible and subject to frequent changes as new personalization features are introduced. Which Google Cloud database service is the most appropriate choice for this migration?Building and operationalizing data processing systems
  23. 23.A large media company manages petabytes of video assets that are accessed frequently for the first 30 days after upload, then rarely (once a quarter or less) for the next 5 years, and finally almost never (once a year or less) for long-term archival. They need a cost-effective storage solution on Google Cloud that automatically transitions data between storage classes based on its age and access patterns, minimizing storage costs while ensuring availability when needed. Which Cloud Storage feature should they implement?Building and operationalizing data processing systems
  24. 24.A logistics company uses a BigQuery data warehouse for analyzing shipment data. They notice that queries filtering on `delivery_date` and `warehouse_id` columns are consistently slow and expensive, even though these columns are frequently used together in WHERE clauses. The table contains billions of rows and is partitioned by `shipment_date`. Which BigQuery feature should they implement to improve query performance and reduce costs for these specific queries?Building and operationalizing data processing systems
  25. 25.A gaming company uses Cloud Spanner for its global leaderboard, which stores player IDs and high scores. To comply with privacy regulations, they need to ensure that player IDs are pseudonymized when used for analytics, but can be reversed to their original form for customer support purposes. The pseudonymization process must be consistent across all systems and securely managed. Which Google Cloud service should be used to manage this reversible pseudonymization?Managing and securing data
  26. 26.A global ride-sharing company collects real-time GPS coordinates from millions of active vehicles. This data needs to be ingested immediately, with very low latency, and then processed by a stream analytics engine to detect anomalies and optimize routes. The data volume can fluctuate significantly, peaking during rush hours. The company requires a messaging service that can handle high throughput, ensure message durability, and scale automatically without manual intervention. Which Google Cloud service is most suitable for ingesting this data?Building and operationalizing data processing systems
  27. 27.A financial services company is migrating its on-premises transactional database to Cloud Spanner. The database contains highly sensitive client financial records, and the company has a strict regulatory requirement that all encryption keys for this data must be stored and managed in a FIPS 140-2 Level 3 certified hardware security module (HSM) that is physically located within their own data centers. They want to ensure that Cloud Spanner uses these keys for encrypting the data at rest. Which solution should the data engineering team choose?Managing and securing data
  28. 28.A global gaming company uses Cloud Spanner for its global leaderboard, which stores player IDs and high scores. To comply with privacy regulations, they need to ensure that actual player IDs are not stored directly in the leaderboard table but can be reversibly linked back to real player identities when necessary (e.g., for customer support or fraud investigation). The mapping between the pseudonymized ID and the real ID must be highly secure and centrally managed. Which approach should the data engineering team implement?Managing and securing data
  29. 29.A global e-commerce company needs to store petabytes of historical transaction data for audit and compliance purposes. This data is accessed very rarely, perhaps once a year, but must be retained for at least 10 years. Cost optimization for storage is a primary concern. Which Cloud Storage class should they choose?Building and operationalizing data processing systems
  30. 30.A financial services company needs to ingest real-time transaction data from thousands of point-of-sale (POS) terminals across various retail locations. The data stream is high-volume, requires low-latency ingestion, and must be reliably delivered for immediate fraud detection analysis. Which Google Cloud service is the most appropriate choice for ingesting this data?Building and operationalizing data processing systems
  31. 31.A data team is developing a new streaming pipeline to process IoT sensor data using Dataflow. They need to monitor the pipeline's performance, identify bottlenecks, and troubleshoot errors in real-time. Specifically, they want to track metrics like data freshness, system latency, and element count, and also view logs for individual pipeline steps. Which Google Cloud services should they primarily use for monitoring and logging this Dataflow pipeline?Building and operationalizing data processing systems
  32. 32.A data engineering team is troubleshooting a Dataflow job that processes real-time sensor data. Users are reporting that some data points appear to be missing from the final aggregated results in BigQuery, particularly during periods of high data volume. The Dataflow job itself reports no errors and appears to be running, but the output is incomplete. The team suspects that some data elements might be getting dropped before they even reach the Dataflow pipeline or are being lost in transit. Which initial step should the team take to investigate potential data loss at the ingestion layer?Building and operationalizing data processing systems
  33. 33.A logistics company uses BigQuery to analyze shipment data. They have a `shipments` table with billions of rows, including columns for `shipment_id`, `origin_city`, `destination_city`, `shipment_date`, and `weight_kg`. Analysts frequently query data filtered by `shipment_date` and `destination_city`, and often aggregate results by `origin_city`. They want to optimize query performance and reduce costs for these common queries. Which BigQuery table optimization strategy should they implement?Building and operationalizing data processing systems
  34. 34.A data team is developing a new streaming pipeline using Google Cloud Dataflow to process sensor data from factory equipment. They realize that some sensor readings might arrive several minutes after their actual event time due to network delays or device buffering. It's critical that these late events are still included in the correct hourly aggregates, but results for an hour should not be delayed indefinitely. Which Dataflow feature should they configure to address this scenario effectively?Building and operationalizing data processing systems
  35. 35.A data engineering team is building a complex data pipeline that involves ingesting data from various sources (CSV, JSON, databases), applying multiple transformations (filtering, joining, aggregation), and loading the results into BigQuery. The pipeline has dependencies between tasks, needs to be scheduled daily, and requires robust error handling and monitoring. They also want to reuse common transformation logic across different datasets. Which Google Cloud service should they use for orchestration?Building and operationalizing data processing systems
  36. 36.A financial institution needs to process daily batch files containing millions of customer transactions. Each file is approximately 50 GB and arrives in Cloud Storage around midnight. The processing involves complex transformations, data validation, and aggregation before loading the results into BigQuery. The entire process must complete within a 4-hour window, and the solution needs to be cost-effective, paying only for the resources consumed during processing. Which Dataflow runner configuration should they prioritize to meet these requirements?Building and operationalizing data processing systems
  37. 37.A data engineering team is developing a new real-time fraud detection system. The system needs to maintain a running count of transactions per user over the last 5 minutes to identify suspicious activity. This count must be continuously updated as new transactions arrive. Which Dataflow feature is essential for implementing this type of logic?Building and operationalizing data processing systems
  38. 38.A global ride-sharing company collects real-time GPS coordinates from millions of active vehicles. This data needs to be ingested, processed by a Dataflow streaming pipeline, and then written to a time-series database. The ingestion service must handle extremely high throughput (millions of messages per second) with low latency, automatically scale to accommodate fluctuating traffic, and provide at-least-once delivery guarantees. Which Google Cloud service should be used for ingesting this real-time data?Building and operationalizing data processing systems
  39. 39.A healthcare provider is migrating its legacy patient record system to Google Cloud. The system requires a globally distributed database that offers strong consistency, high availability, and horizontal scalability to handle potential growth. The data model is relational, and strict ACID properties are essential for transactional integrity. Which Google Cloud database service is the most appropriate choice for these requirements?Building and operationalizing data processing systems
  40. 40.A data engineering team is developing a new streaming pipeline to process real-time sensor data from IoT devices using Dataflow. The pipeline needs to perform complex aggregations and transformations, and then write the results to a BigQuery table. The Dataflow job is running in a private network, and BigQuery is also configured with Private Google Access. To ensure secure and private communication between the Dataflow workers and BigQuery, what authentication and networking setup is required?Building and operationalizing data processing systems
  41. 41.A data team developed a new Dataflow pipeline that processes gigabytes of data hourly. After deploying to production, they observe that the pipeline frequently stalls and reports 'Out of memory' errors on certain worker nodes, particularly during peak load. The pipeline code has been reviewed and optimized for memory efficiency as much as possible. Which Dataflow feature should they adjust to address these memory issues and ensure stable operation?Building and operationalizing data processing systems
  42. 42.A global e-commerce company experiences peak traffic during flash sales, which can cause their traditional batch data processing jobs to fall behind. They need to analyze customer behavior in near real-time during these events to dynamically adjust recommendations and promotions. The solution must handle sudden, massive spikes in data volume without manual intervention. Which data processing pattern should they implement?Building and operationalizing data processing systems
  43. 43.A data analytics team is migrating an on-premises Hadoop cluster to Google Cloud. They need to store petabytes of raw, unstructured, and semi-structured data (logs, images, JSON files) for long-term retention and future analysis. The data will be accessed by various analytics tools, including BigQuery, Dataflow, and custom machine learning models. They require a scalable, durable, and cost-effective storage solution that can handle diverse data formats without imposing a rigid schema. Which Google Cloud service is the most appropriate for this data lake requirement?Building and operationalizing data processing systems
  44. 44.A data engineer is designing a new data ingestion pipeline for a global IoT platform. Millions of devices send telemetry data every second to a regional Cloud Pub/Sub topic. The data needs to be processed by a Dataflow streaming job and then stored in BigQuery. The engineer is concerned about potential message loss or duplicates during ingestion and processing, especially under high load or network interruptions. They need to ensure that each unique telemetry event is processed exactly once in BigQuery. How can this be achieved in the Dataflow pipeline?Building and operationalizing data processing systems
  45. 45.A data engineering team is building a data pipeline that ingests sensitive customer data from various sources into Cloud Storage. Due to strict regulatory requirements, all data at rest in Cloud Storage must be encrypted using customer-provided encryption keys (CSEK), where the keys are managed and supplied by the customer, not Google. Which encryption option should they configure for their Cloud Storage buckets and objects?Managing and securing data
  46. 46.A data engineering team is building a new real-time fraud detection system. The system needs to process financial transactions as they occur, maintain a running count of transactions per user within a 5-minute window, and identify users exceeding a certain transaction threshold. The processing must guarantee that all transactions are processed exactly once, even in the event of worker failures or restarts. Which Dataflow feature is crucial for maintaining the running count accurately and ensuring exactly-once processing?Building and operationalizing data processing systems
  47. 47.A global logistics company uses BigQuery for analyzing shipment data. They have a dataset containing sensitive customer information that is classified as 'Confidential'. All access to this dataset must be logged, and any attempts by unauthorized users to access it should trigger an immediate alert to the security operations center (SOC). How should you implement this data governance policy?Managing and securing data
  48. 48.A global ride-sharing company collects real-time GPS coordinates from millions of active vehicles. This data needs to be ingested into a scalable, low-latency messaging system that can handle extremely high throughput and fan-out to multiple downstream processing services (e.g., for real-time mapping, analytics, and driver incentives). Which Google Cloud service is best suited for this ingestion requirement?Building and operationalizing data processing systems
  49. 49.A global IoT company collects sensor data from millions of devices worldwide. This data is ingested into Cloud Pub/Sub and then processed by a Dataflow streaming pipeline. The data volume can fluctuate significantly throughout the day, with peak loads reaching 5x the average. The Dataflow pipeline needs to dynamically adjust its worker resources to handle these fluctuations without manual intervention, ensuring consistent processing latency and minimizing operational costs. Which Dataflow feature is crucial for achieving this dynamic resource allocation?Building and operationalizing data processing systems
  50. 50.A healthcare provider is building a new system to store anonymized patient records. They need a highly available, globally consistent, transactional database that can scale horizontally to handle millions of records and thousands of concurrent read/write operations. Data integrity and strong consistency are paramount. Which Google Cloud database service should they choose?Building and operationalizing data processing systems