Professional Data Engineer practice questions

262 free questions with answers and explanations.

Practice test
  1. 51.A data team is developing a new streaming pipeline using Google Cloud Dataflow to process financial transactions. The pipeline needs to calculate the sum of transaction amounts for each user within a 1-minute fixed window. However, transactions can arrive out of order, with some events arriving several minutes late. The team needs to ensure that all late-arriving data for a given window is eventually included in the window's calculation, even if it arrives significantly after the window's end. Which Dataflow concept should they use to handle this requirement efficiently?Building and operationalizing data processing systems
  2. 52.A data engineer is tasked with migrating an on-premises data processing pipeline to Google Cloud. The existing pipeline uses custom scripts written in Python and Bash to extract data from various databases, transform it, and load it into a data warehouse. The team wants to keep their existing code largely intact but needs a scalable, serverless environment to run these scripts without managing virtual machines. They also need to ensure that the environment can handle fluctuating workloads and integrate with other Google Cloud services. Which Google Cloud service should they choose to run their existing scripts?Building and operationalizing data processing systems
  3. 53.A global media company needs to process millions of video frames per second for real-time content moderation. Each frame needs to be analyzed by a machine learning model, and the processing must be highly scalable and fault-tolerant. The processing logic is complex and requires custom code. Which Google Cloud service is best suited for this task?Building and operationalizing data processing systems
  4. 54.A media company needs to store petabytes of video and audio content for archival purposes. This data is rarely accessed after initial ingestion but must be available within minutes when requested. Cost optimization for storage is a primary concern. Which Google Cloud storage class is most suitable for this requirement?Building and operationalizing data processing systems
  5. 55.A data engineering team is building a real-time analytics pipeline using Cloud Dataflow. The pipeline ingests sensor data from IoT devices, performs aggregations over 5-minute windows, and then writes the results to BigQuery. During testing, they observe that some sensor readings arrive several minutes late due to network fluctuations. These late readings are crucial for accurate analytics and must be included in the correct 5-minute window, even if they arrive after the window's normal processing time. Which Dataflow feature should they configure to handle these late-arriving elements?Building and operationalizing data processing systems
  6. 56.A research institution is building a data lake on Google Cloud using Cloud Storage. They frequently ingest new datasets from various external sources, and data scientists need to quickly discover, understand, and access these datasets. The institution requires a centralized metadata management solution that automatically catalogs data, supports custom metadata, and helps with data governance. Which Google Cloud service should they integrate?Managing and securing data
  7. 57.A data analytics team is migrating an on-premises Hadoop cluster to Google Cloud. They have petabytes of diverse, unstructured, and semi-structured data (logs, images, videos, CSV files) that need to be stored cost-effectively, made accessible for various analytical tools (Dataproc, BigQuery, Dataflow), and serve as the single source of truth. Which Google Cloud service is the most appropriate foundational component for building this data lake?Building and operationalizing data processing systems
  8. 58.A research institution is building a data lake on Google Cloud using Cloud Storage. They have numerous research projects, each generating large volumes of diverse data (structured, semi-structured, unstructured). To ensure proper data governance, they need a centralized way to discover, understand, and manage metadata for all their data assets, regardless of where they reside within the data lake. This includes tracking schema, lineage, and business descriptions. Which Google Cloud service is best suited for this requirement?Managing and securing data
  9. 59.A financial institution is migrating its on-premises data archive to Google Cloud Storage. The archive contains sensitive customer transaction records that must be encrypted at rest. The institution has a strict compliance requirement to manage its own encryption keys entirely outside of Google Cloud's infrastructure, while still leveraging Google Cloud Storage for data durability and availability. Which encryption option should the data engineering team recommend?Managing and securing data
  10. 60.A global online gaming company needs to store petabytes of user gameplay data, including session logs, in-game actions, and chat messages. This data will be frequently accessed for real-time analytics dashboards, machine learning model training, and historical trend analysis. The company requires a highly scalable, low-latency storage solution that can handle millions of reads and writes per second. Which Google Cloud storage service is best suited for this requirement?Building and operationalizing data processing systems
  11. 61.A data engineering team is developing a new streaming pipeline to process IoT sensor data from millions of devices. They observe that the Dataflow job's CPU utilization is consistently high (above 80%), even with autoscaling enabled and increasing worker count. This is leading to increased processing latency and backlogs. The pipeline involves several `GroupByKey` and `Combine` operations. What is the most effective strategy to reduce CPU utilization and improve performance in this scenario?Building and operationalizing data processing systems
  12. 62.A data engineering team is developing a new streaming pipeline that processes real-time sensor data from industrial machinery. The data needs to be enriched with historical maintenance records stored in a Cloud SQL database. Due to data privacy regulations, the Cloud SQL database should only be accessible by specific service accounts with the least necessary privileges. Which Google Cloud security feature should be used to securely connect the Dataflow job to Cloud SQL?Building and operationalizing data processing systems
  13. 63.A global e-commerce company uses BigQuery for its analytics platform. They have several datasets containing customer personal identifiable information (PII) that must be masked for analysts who do not have explicit permission to view raw PII. The company needs to implement a solution that dynamically masks sensitive columns without creating multiple copies of the data. How should you implement this data privacy requirement?Managing and securing data
  14. 64.A retail company is analyzing customer purchase data stored in BigQuery. The `transactions` table contains billions of rows, with columns like `transaction_id`, `customer_id`, `transaction_timestamp`, `product_id`, and `amount`. Analysts frequently query this table to find transactions for specific `customer_id` values within a given date range. Queries are becoming slow and expensive. The team needs to optimize query performance and reduce costs for these common analytical patterns. Which BigQuery table optimization strategy should they implement?Building and operationalizing data processing systems
  15. 65.A global e-commerce company experiences peak traffic during flash sales, which can cause their data ingestion and processing pipelines to fall behind. They use Cloud Pub/Sub for ingestion and Dataflow for real-time processing. During off-peak hours, traffic is significantly lower. The company needs to ensure that the Dataflow pipeline can automatically scale its resources up and down to match the fluctuating ingestion rate from Pub/Sub, processing all data immediately while optimizing costs. Which Dataflow feature is essential for this requirement?Building and operationalizing data processing systems
  16. 66.A pharmaceutical company stores sensitive clinical trial data in BigQuery. The data must be accessible only by specific analysts within the R&D department, and their access should be restricted to aggregated, depersonalized views of the data. Furthermore, due to an upcoming audit, the company needs to demonstrate that access to this sensitive data is strictly controlled and auditable. How should you design the access control and auditing strategy?Managing and securing data
  17. 67.A global e-commerce company uses BigQuery for its analytics platform. They have sensitive customer data, including payment card information, which must be protected according to PCI DSS. While the data is stored securely, analysts need to query this data for aggregate reporting. The company wants to ensure that raw payment card numbers are never directly visible to analysts, even if they have access to the tables, but other fields in the same table should be accessible. How should the company implement this data access control?Managing and securing data
  18. 68.A data engineering team is building a data pipeline that ingests data from various sources into a Cloud Storage data lake. They need to ensure that all data at rest in the data lake is encrypted with keys that are centrally managed and controlled by the security team, and that the keys are protected by a FIPS 140-2 Level 3 validated hardware security module (HSM). Which Google Cloud service should be used to manage these encryption keys?Managing and securing data
  19. 69.A global media company uses BigQuery for analyzing user engagement data. They want to ensure that all queries accessing sensitive user demographic information (e.g., age, gender) are logged, including the user who ran the query, the query text, and the specific columns accessed. This information is crucial for their internal auditing and compliance requirements. How should they configure BigQuery to meet this logging requirement?Managing and securing data
  20. 70.A healthcare analytics startup is using BigQuery to store de-identified patient data. They need to ensure that data is stored in the EU region to comply with GDPR, and that all data transfers between Google Cloud services within their project also remain within the EU region. How can they enforce this data residency requirement?Managing and securing data
  21. 71.A data engineering team is troubleshooting a Dataflow job that processes real-time sensor data. Users are reporting that the dashboards powered by this pipeline are showing stale data, often delayed by several minutes. Upon investigation, the Dataflow monitoring interface shows that the 'System Latency' metric is consistently high, while 'Data freshness' is low. The pipeline's 'Element count' is steady, and 'CPU utilization' on workers is around 60%. What is the most likely cause of the stale data and high system latency?Building and operationalizing data processing systems
  22. 72.A large e-commerce company is experiencing inconsistent performance in their data processing pipelines, leading to missed SLAs for critical reports. These pipelines are built using Dataflow and integrate with various Google Cloud services. The data engineering team needs a centralized solution to monitor pipeline health, identify bottlenecks, and troubleshoot issues across different stages of the data processing workflow. What is the most appropriate Google Cloud service to implement for comprehensive monitoring and logging?Building and operationalizing data processing systems
  23. 73.A data engineering team is building a pipeline that processes sensor data from IoT devices. The data arrives continuously and needs to be aggregated into 5-minute windows before being stored in BigQuery. Due to network fluctuations, some data points might arrive a few minutes late. The team needs to ensure that all data points, including late ones, are included in the correct 5-minute window without indefinitely delaying the pipeline. Which Apache Beam concept should they use to achieve this?Building and operationalizing data processing systems
  24. 74.A financial institution processes sensitive customer transaction data daily. The data is generated on-premises and needs to be ingested into Google Cloud for processing and analysis in BigQuery. Due to strict security and compliance requirements, the data cannot traverse the public internet, and a direct, private, and highly available connection is mandated. Which Google Cloud networking solution should be used to securely ingest this data?Building and operationalizing data processing systems
  25. 75.A retail company uses BigQuery for analyzing customer purchasing behavior. They have a dataset containing sensitive customer PII that needs to be retained for 7 years for auditing purposes, but older, less frequently accessed data should be moved to a cheaper storage class after 3 years. Additionally, specific tables containing temporary analytical results should be automatically deleted after 30 days to save costs. Which BigQuery feature or combination of features should they use to manage this data lifecycle effectively?Managing and securing data
  26. 76.A data team developed a new Dataflow pipeline that processes gigabytes of data hourly. After initial deployment, they observe that the pipeline consistently falls behind, leading to increasing backlogs in Pub/Sub. The Dataflow job metrics show high CPU utilization and low memory utilization on the workers. What is the most effective action to improve the pipeline's throughput and catch up with the incoming data?Building and operationalizing data processing systems
  27. 77.A telecommunications company uses Pub/Sub for ingesting real-time call detail records (CDRs). Due to regulatory compliance, these records must be retained for exactly 30 days and then automatically purged. They also need to ensure that if a subscriber is temporarily unavailable, messages are not lost and can be redelivered later. How should you configure the Pub/Sub topic and subscription to meet these requirements?Managing and securing data
  28. 78.A financial services company is migrating its on-premises transactional database to Google Cloud. The database contains critical customer account information that requires high availability, strong consistency, and robust disaster recovery capabilities across multiple regions. Due to strict regulatory requirements, all data at rest must be encrypted with customer-managed encryption keys (CMEK). Which Google Cloud database service and encryption setup should they choose?Managing and securing data
  29. 79.A data team is optimizing the cost of their BigQuery data warehouse. They have several tables that store historical logs, where data older than 90 days is very rarely accessed, but must be retained for compliance purposes. Queries on this older data are acceptable to run with slightly higher latency. The team wants to reduce storage costs for this older, infrequently accessed data without moving it out of BigQuery. Which BigQuery feature should they use?Building and operationalizing data processing systems
  30. 80.A global logistics company uses BigQuery for analyzing shipment data. They have a dataset containing sensitive customer addresses and shipment contents. To ensure data privacy and compliance, they need to restrict access to this dataset to only authorized personnel, and all access attempts, including failed ones, must be logged for auditing purposes. Additionally, they need to track who accessed what data, when, and from where. Which combination of Google Cloud services should they use to meet these requirements?Managing and securing data
  31. 81.A data team is developing a new streaming analytics pipeline using Dataflow. They observe that some events arrive significantly out of order, sometimes minutes or hours late, due to network issues or upstream system delays. These late events are critical for accurate historical aggregation and should not be discarded. The team needs to configure their Dataflow pipeline to correctly process these late events without unduly delaying the processing of on-time events. Which Dataflow concept should they configure?Building and operationalizing data processing systems
  32. 82.A media company is building a data pipeline to process user engagement metrics from their streaming platform. The pipeline uses Dataflow to aggregate events (e.g., 'play', 'pause', 'seek') into 1-minute fixed windows. Due to network conditions, some events might arrive slightly late, but they are still valuable if processed within 30 seconds of their window's end. Events arriving more than 30 seconds late should be discarded. How should the Dataflow pipeline be configured to handle these late events?Building and operationalizing data processing systems
  33. 83.A payment processing company uses Cloud SQL for PostgreSQL to store transactional data, including customer payment details. Due to strict industry regulations (e.g., PCI DSS), all data at rest must be encrypted using customer-managed keys. The company wants to leverage Cloud KMS for key management to centralize key lifecycle and access control. How should the data engineering team configure Cloud SQL to meet this requirement?Managing and securing data
  34. 84.A data analytics team is designing a new batch processing pipeline to analyze historical customer purchase data. The data, currently stored in CSV files in Cloud Storage, needs to be loaded into BigQuery for complex analytical queries. The team wants to ensure data quality and perform transformations like data type conversions and column renaming before loading, without writing extensive custom code. Which Google Cloud service should they use to achieve this efficiently?Building and operationalizing data processing systems
  35. 85.A global e-commerce company uses BigQuery for its analytical data warehouse. They have a daily batch job that loads millions of new customer orders into a BigQuery table. This job needs to start immediately after the previous day's order processing is complete and must complete within a specific time window to ensure dashboards are updated before business hours. The company wants to implement robust scheduling and dependency management for this job, along with visibility into its execution status. Which Google Cloud service is best suited for orchestrating this daily batch load?Building and operationalizing data processing systems
  36. 86.A global e-commerce company uses BigQuery for its analytical data warehouse. They have a daily batch job that processes customer order data, performs complex transformations, and loads the results into another BigQuery table. This job is part of a larger pipeline that includes data ingestion from multiple sources, data quality checks, and report generation. The company wants to ensure reliable, scheduled execution of this entire pipeline with dependencies between tasks. Which Google Cloud service would be most appropriate for orchestrating this workflow?Building and operationalizing data processing systems
  37. 87.A healthcare provider is storing patient medical records in Cloud Storage. Due to HIPAA regulations, they must ensure that all data is encrypted at rest using keys managed by the organization, and that these keys can be revoked when necessary. The data engineering team needs to implement a solution that provides the highest level of control over the encryption keys. Which encryption method should they choose for Cloud Storage?Managing and securing data
  38. 88.A large e-commerce company uses BigQuery as its primary data warehouse. They have a daily sales table containing billions of rows, with columns such as `order_id`, `customer_id`, `order_timestamp`, `product_id`, and `order_total`. Analysts frequently query this table, often filtering by `order_timestamp` (for specific dates or date ranges) and `customer_id`. The table is currently partitioned by `order_timestamp`. To further optimize these common queries for both performance and cost, what additional BigQuery feature should be applied?Building and operationalizing data processing systems
  39. 89.A financial institution is migrating its on-premises data warehouse to Google Cloud. They have stringent compliance requirements that mandate all sensitive data at rest must be encrypted with customer-managed encryption keys (CMEK), and the keys themselves must be generated and stored within FIPS 140-2 Level 3 certified hardware security modules (HSMs). Which Google Cloud Key Management Service (KMS) key type should they provision to meet these requirements for their data stored in services like BigQuery and Cloud Storage?Managing and securing data
  40. 90.A global IoT company collects sensor data from millions of devices worldwide. This data is ingested into Cloud Pub/Sub and then processed by a Dataflow streaming job. The data volume can fluctuate significantly throughout the day, with peak loads reaching ten times the average. The company needs to ensure the Dataflow job can handle these spikes without manual intervention, maintaining consistent processing latency. Which Dataflow feature is crucial for this requirement?Building and operationalizing data processing systems
  41. 91.A data engineer is implementing a custom data processing job on Compute Engine. This job needs to access sensitive customer data stored in Cloud Storage. To ensure secure access, the engineer wants to grant the Compute Engine VM only the necessary permissions and avoid storing credentials directly on the VM. Which Google Cloud security feature should be configured for the Compute Engine instance?Building and operationalizing data processing systems
  42. 92.A financial institution is building a data warehouse in BigQuery. They need to ensure that data lineage is tracked for all transformations, from ingestion to final aggregated reports, to meet strict auditing and compliance requirements (e.g., BCBS 239). The lineage information must include schema changes, transformation logic, and responsible users. Which Google Cloud service should you integrate to achieve this comprehensive data lineage tracking?Managing and securing data
  43. 93.A data analytics team is analyzing a large dataset in BigQuery. They have a table with over a billion rows representing customer interactions, with columns like `customer_id`, `event_timestamp`, and `event_type`. Queries frequently filter by `event_timestamp` within specific date ranges and then aggregate results by `customer_id`. The team notices that these queries are often slow and scan a large amount of data, leading to high costs. What BigQuery optimization strategy should they implement to improve query performance and reduce costs for these specific queries?Building and operationalizing data processing systems
  44. 94.A global logistics company uses BigQuery for analyzing shipment data. They have a dataset containing sensitive customer addresses and shipment contents. To comply with GDPR, they need to ensure that access to this dataset is strictly controlled and all access attempts, successful or not, are logged with details about the user and the specific columns or rows accessed. Furthermore, they need to be able to trace who accessed what data for audit purposes. How should they implement this comprehensive data governance and auditing solution?Managing and securing data
  45. 95.A logistics company uses a BigQuery data warehouse for analyzing shipment data. They notice that queries on their `shipments` table, which contains billions of rows, are becoming increasingly slow, especially when filtering by `destination_country` and `shipment_date`. The table is partitioned by `shipment_date`. What BigQuery feature should they implement to improve query performance specifically for filters on `destination_country`?Building and operationalizing data processing systems
  46. 96.A global e-commerce company uses BigQuery for its analytics platform. They have several datasets containing sensitive customer information (e.g., email addresses, phone numbers) that must be protected. Data analysts need to perform queries on these datasets for marketing campaigns, but they should only see masked versions of the sensitive data while still being able to join tables and perform aggregate analysis. Which BigQuery feature should be used to achieve this without creating separate copies of the data?Managing and securing data
  47. 97.A financial institution processes sensitive customer transaction data daily. The data is generated in various on-premises systems and needs to be ingested into Google Cloud for analysis in BigQuery. Due to strict regulatory compliance requirements, the data must be encrypted at rest and in transit, and access must be tightly controlled and auditable. Furthermore, the ingestion process must be highly available and resilient to network disruptions. Which Google Cloud services and practices should the data engineer prioritize for a secure and reliable ingestion strategy?Building and operationalizing data processing systems
  48. 98.A healthcare analytics startup is using BigQuery to store de-identified patient data. To comply with regional data residency requirements, they must ensure that all data for European patients is stored only in data centers located within the European Union, while data for North American patients must be stored only in North American data centers. The startup needs an automated and enforceable mechanism to prevent users from accidentally or intentionally creating datasets in non-compliant regions. Which Google Cloud feature should they implement?Managing and securing data
  49. 99.A data engineer is designing a real-time analytics pipeline using Cloud Pub/Sub and Dataflow. The pipeline needs to process millions of messages per second, and the incoming data stream has highly variable throughput. To ensure message delivery and prevent data loss, the Pub/Sub topic configuration needs to be optimized for high volume and bursty traffic. Which Pub/Sub feature should be enabled to handle this scenario effectively?Building and operationalizing data processing systems
  50. 100.A data team is building a new real-time fraud detection system. The system needs to process millions of transactions per second, perform complex aggregations across multiple transactions from the same user within a short time window, and maintain state information (e.g., total spend in the last 5 minutes) for each user. The results must be available with low latency for immediate decision-making. Which Google Cloud service is best suited for this scenario?Building and operationalizing data processing systems