Professional Data Engineer practice questions

262 free questions with answers and explanations.

Practice test
  1. 201.A startup is building a new mobile application that generates user activity logs. They expect a rapidly growing user base and unpredictable spikes in traffic. The logs need to be ingested in real-time, stored, and then later processed for analytics. The team wants a solution that can automatically scale to handle varying loads, is fully managed to minimize operational burden, and offers a flexible schema to accommodate future changes in log formats. Which Google Cloud service should they use for real-time ingestion?Designing data processing systems
  2. 202.A global ride-sharing company processes millions of GPS coordinates per second from active vehicles. This real-time location data is critical for dynamic pricing, driver dispatch, and route optimization. The system requires ultra-low latency reads and writes (milliseconds) and must be highly scalable to accommodate fluctuating demand. The data schema is relatively simple (key-value pairs with timestamps). Which Google Cloud service is the most appropriate for storing and serving this operational data?Designing data processing systems
  3. 203.A global online retailer is building a new real-time fraud detection system. The system needs to ingest millions of events per second, perform complex aggregations and pattern matching, and detect fraudulent activities with sub-second latency. Data must be durable and available for historical analysis. Which Google Cloud services should be combined to meet these requirements efficiently?Designing data processing systems
  4. 204.A financial services company needs to build a highly available and durable data ingestion pipeline for real-time stock market data. The data arrives as a continuous stream of JSON messages, and it's critical that no data is lost even during temporary downstream processing failures. Which Google Cloud service should be used for the ingestion layer to ensure message persistence and reliable delivery?Designing data processing systems
  5. 205.A data analytics team uses BigQuery for its data warehouse. They have several tables containing historical data that are accessed infrequently (e.g., once a quarter for compliance reports) but must be retained for 10 years. These tables consume a significant amount of storage, leading to high costs. The team wants to optimize storage costs while maintaining the ability to query this data when needed. Which BigQuery feature should they leverage?Designing data processing systems
  6. 206.A healthcare provider is building a data lake for patient records, which contain highly sensitive Protected Health Information (PHI). They need to implement strong data governance to ensure data quality, enforce access policies, track data lineage, and classify sensitive data. Which Google Cloud service is specifically designed to address these data governance challenges in a unified manner?Designing data processing systems
  7. 207.A healthcare provider is migrating its on-premises electronic health records (EHR) system to Google Cloud. The data contains highly sensitive patient information and must comply with HIPAA regulations. The data lake will store raw, unstructured, and semi-structured data from various sources. The design requires robust access controls, encryption at rest and in transit, and auditing capabilities to track all data access. Which Google Cloud storage service is the most appropriate foundational layer for this data lake, and what key security features should be configured?Designing data processing systems
  8. 208.A pharmaceutical company is building a data pipeline to process genomic sequencing data. This data is extremely large (petabytes per study) and requires complex, computationally intensive transformations (e.g., alignment, variant calling) that can take days to complete. The processing jobs are not time-sensitive but must be fault-tolerant and scale elastically to handle varying workloads and data volumes across different research projects. Which Google Cloud service is most suitable for orchestrating and executing these batch processing jobs?Designing data processing systems
  9. 209.A global e-commerce company needs to store customer order history, which can grow to petabytes over time. This data is primarily accessed for historical reporting and infrequent ad-hoc analysis, but not for real-time transactions. The company wants to optimize storage costs while maintaining high availability. Which Google Cloud storage solution is most suitable for this use case?Designing data processing systems
  10. 210.A global ride-sharing company processes millions of sensor data points per second from vehicles, including GPS coordinates, accelerometer readings, and engine diagnostics. This data is critical for real-time fleet management, predictive maintenance, and driver safety analytics. The company requires a database that can handle extremely high write throughput, efficient time-series queries, and scale globally to petabytes of data with minimal latency. Which Google Cloud database is best suited for this demanding workload?Designing data processing systems
  11. 211.A startup is building a new mobile application that generates user activity logs. They expect millions of users, leading to a high volume of events (hundreds of thousands per second) that need to be captured reliably and immediately. The events will then be processed by downstream analytics systems. The solution must be fully managed, highly scalable, and provide at-least-once delivery guarantees. Which Google Cloud service is the most appropriate for ingesting these high-volume event streams?Designing data processing systems
  12. 212.A large retail company needs to migrate its on-premises operational analytics database to Google Cloud. This database handles millions of transactions per second, requires extremely low-latency reads and writes (single-digit milliseconds), and stores data with a simple key-value structure but needs to support complex aggregations and time-series analysis. The data grows to petabytes. Which Google Cloud service is the most appropriate for this workload?Designing data processing systems
  13. 213.A data science team is developing a recommendation engine that needs to perform real-time feature lookups for user personalization. The engine requires serving features (e.g., user preferences, interaction history) with single-digit millisecond latency to support live inference. The data volume for these features can grow substantially, and the system needs to be highly available and scalable. Which Google Cloud database is the most suitable for this real-time feature serving requirement?Designing data processing systems
  14. 214.A global financial institution needs to build a highly secure and compliant data processing system on Google Cloud. The system will handle sensitive customer financial transactions, requiring strict data residency, encryption, and access control. Due to regulatory requirements, all data processing activities must be auditable, and data must be immediately available for analysis. The solution must also be cost-effective for petabyte-scale storage. Which combination of Google Cloud services should you recommend to meet these requirements?Designing data processing systems
  15. 215.A global ride-sharing company processes millions of GPS coordinates per second from active vehicles. This data is critical for real-time tracking, dispatch, and estimated time of arrival (ETA) calculations, requiring extremely low-latency reads and writes (milliseconds). The data schema is simple: device ID, timestamp, latitude, longitude. Which Google Cloud database is best suited for storing and serving this high-velocity, low-latency operational time-series data?Designing data processing systems
  16. 216.A data analytics company has developed a proprietary machine learning model for fraud detection. The model's training and prediction pipelines involve multiple steps: data extraction from various sources, data cleaning and transformation using Spark, model training on GPUs, and model deployment to an endpoint. These steps run on a schedule, have interdependencies, and require error handling and retries. Which Google Cloud service should be used to orchestrate these complex, multi-step workflows?Designing data processing systems
  17. 217.A large e-commerce company needs to store customer order history, which can grow to petabytes of data. This data is primarily used for historical analysis, trend reporting, and machine learning model training. Queries are often complex, involving joins across many tables and aggregations over vast date ranges. The solution must be highly scalable, performant for analytical queries, and cost-effective for petabyte-scale storage. Which Google Cloud service is the most appropriate for this data warehousing requirement?Designing data processing systems
  18. 218.A global e-commerce company needs to process millions of transactions per second for real-time fraud detection. The system must be highly available, scalable, and capable of ingesting data from various sources (web, mobile, IoT devices). After ingestion, the data needs to be processed with minimal latency to identify fraudulent activities. Which Google Cloud service is best suited for ingesting these high-volume, real-time event streams?Designing data processing systems
  19. 219.A financial services company needs to process daily batch jobs that involve complex transformations and aggregations on terabytes of historical transaction data. The jobs have strict Service Level Agreements (SLAs) for completion time and require a fully managed, serverless solution to minimize operational overhead. The solution must also be cost-effective, paying only for the resources consumed during job execution. Which Google Cloud service is most suitable for this scenario?Designing data processing systems
  20. 220.A global e-commerce company needs to store product catalog data, which is frequently updated and accessed with low latency by customer-facing applications worldwide. The data model is semi-structured and requires flexible schema management. The solution must provide strong consistency and high availability across multiple regions. Which Google Cloud service should the company use?Designing data processing systems
  21. 221.A data analytics team needs to process daily batches of customer order data, which can range from gigabytes to several terabytes. The processing involves complex transformations, aggregations, and joins before loading the data into BigQuery for reporting. The team requires a serverless solution that automatically scales based on data volume, provides fault tolerance, and minimizes operational overhead. Which Google Cloud service is the most appropriate for this batch processing workload?Designing data processing systems
  22. 222.A financial institution processes millions of transactions daily. Due to regulatory requirements, all transaction data must be retained for 7 years for auditing and compliance, but only the most recent 90 days of data are actively queried for operational reporting. The remaining historical data is rarely accessed but must be available if needed. The company wants to minimize storage costs while ensuring data availability and compliance. Which BigQuery feature should be primarily used to optimize costs for this scenario?Designing data processing systems
  23. 223.A data engineering team is building a new data lake on Google Cloud to store vast amounts of raw, unstructured, and semi-structured data from various sources. The data lake must support flexible schema evolution, integrate with various analytics tools, and provide cost-effective long-term storage. Which Google Cloud service is the most appropriate foundational component for this data lake?Designing data processing systems
  24. 224.A data analytics team needs to build a robust data pipeline to ingest real-time sensor data from thousands of IoT devices. The data needs to be processed with minimal latency, enriched with contextual information, and then stored in a data warehouse for further analysis. The system must be highly scalable to handle fluctuating data volumes and resilient to device failures or network outages. Which set of Google Cloud services should be used to design this pipeline?Designing data processing systems
  25. 225.A global ride-sharing company processes millions of GPS coordinates per second from active vehicles. This time-series data needs to be ingested continuously, stored efficiently for at least 30 days for immediate operational analysis (e.g., real-time traffic, dispatch optimization), and then archived for long-term historical analysis spanning several years. The system must support high write throughput and low-latency reads for recent data. Which combination of Google Cloud services should you recommend?Designing data processing systems
  26. 226.A media company needs to process millions of video transcoding jobs daily. Each job involves downloading a source video, applying various transformations, and uploading the processed video to Cloud Storage. The workload is highly spiky, with peak demands occurring during prime-time hours. The company requires a solution that is fully managed, scales automatically, and only charges for the actual compute used. Which Google Cloud product is best suited for this use case?Designing data processing systems
  27. 227.A global gaming company collects telemetry data from millions of active users. This data, which includes game events, player progress, and in-game purchases, needs to be ingested in real-time, processed to detect anomalies and personalize user experiences, and then stored for historical analysis. The system must be highly available, fault-tolerant, and scale dynamically to handle sudden spikes in user activity during new game releases. Which Google Cloud services should be used to design this hybrid real-time and batch data pipeline?Designing data processing systems
  28. 228.A data engineering team needs to implement a data governance strategy for a new data lake on Google Cloud. The strategy requires centralizing metadata management, ensuring data quality, and providing a unified view of all data assets across various Google Cloud services and on-premises sources. Which Google Cloud service is designed to address these data governance requirements?Designing data processing systems
  29. 229.A financial institution is building a data lake on Google Cloud using Cloud Storage and BigQuery. They need to implement strict data retention policies for different types of data (e.g., transaction logs for 7 years, audit trails for 10 years, temporary processing files for 30 days). You need to automate the enforcement of these retention policies in Cloud Storage to ensure compliance and optimize costs. Which Cloud Storage feature should you use?Ensuring solution quality
  30. 230.A pharmaceutical company is processing clinical trial data, which is highly sensitive and subject to strict regulatory audits. They need to ensure that all data transformations performed by their Dataflow pipelines are fully traceable and immutable. Any change to the data schema or transformation logic must be recorded and auditable. You need to design a solution that provides this level of traceability and immutability for Dataflow pipelines. Which approach is best?Ensuring solution quality
  31. 231.A fast-growing startup is building a new real-time analytics platform on Google Cloud. They anticipate rapid increases in data volume and query complexity, and need a BigQuery solution that can handle petabytes of data and thousands of concurrent users without performance degradation. The solution must provide consistent, high performance and low latency, even under extreme load, to support interactive dashboards and ad-hoc analysis. Which BigQuery feature should they prioritize to ensure scalable and high-performance analytics for their growing needs?Ensuring solution quality
  32. 232.A global ride-sharing company is building a new data processing pipeline to analyze driver and rider location data. Due to the massive scale (trillions of records) and the need for extremely low-latency queries (milliseconds) for real-time decision-making (e.g., dynamic pricing, driver matching), a traditional relational database or standard data warehouse is insufficient. The data is primarily time-series, with new data constantly appended and historical data frequently accessed. You need to select a Google Cloud database service that can handle this scale and performance requirement. Which service is most appropriate?Ensuring solution quality
  33. 233.A multinational retail company is building a new customer analytics platform on Google Cloud. They store customer personal identifiable information (PII) in BigQuery and need to comply with GDPR and CCPA regulations. This requires ensuring that customer data can be securely deleted upon request (right to be forgotten) and that data access is restricted based on user roles and data sensitivity. Which BigQuery features should they combine to meet the 'right to be forgotten' and granular access control requirements?Ensuring solution quality
  34. 234.A global logistics company uses BigQuery for its operational analytics, processing billions of rows daily. They notice that certain complex queries, particularly those involving large joins and aggregations on historical data, sometimes run for several minutes, impacting dashboard refresh times. The queries are critical and cannot be simplified. You need to optimize the performance of these specific, complex queries while minimizing cost impact. What BigQuery feature should you leverage?Ensuring solution quality
  35. 235.A media company is developing a new recommendation engine that uses machine learning models. The data scientists frequently experiment with new features and model architectures, requiring rapid iteration and testing. They need a data processing environment that allows them to quickly provision and de-provision clusters, run various open-source data tools (e.g., Spark, Hadoop, Presto), and scale resources up or down on demand without managing underlying infrastructure. You need to recommend a Google Cloud service for this flexible and agile data processing. Which service is most suitable?Ensuring solution quality
  36. 236.A marketing analytics team uses BigQuery for ad-hoc analysis and reporting. They frequently run queries on large datasets, leading to unpredictable monthly costs due to BigQuery's on-demand pricing model. The team has a consistent budget for data analytics and prefers predictable spending. You need to recommend a BigQuery pricing model that provides cost predictability and potentially better performance for their workload. Which pricing model should they choose?Ensuring solution quality
  37. 237.A global media company uses Google Cloud for its data analytics platform. They have multiple BigQuery datasets across different regions, each containing sensitive user data. To comply with GDPR, CCPA, and other regional regulations, they need to ensure that data access is strictly controlled based on the user's role and their need-to-know, regardless of the dataset's physical location. Specifically, they want to restrict access to certain columns (e.g., email addresses) for specific user groups, while allowing them to view other columns in the same table. Which BigQuery security feature should be implemented to achieve this granular control?Ensuring solution quality
  38. 238.A global media streaming service processes billions of user interaction events (clicks, views, searches) in real-time. They need to monitor their streaming data pipelines for data quality issues, such as missing events, malformed data, or sudden drops in throughput, and be alerted immediately. The solution must provide real-time visibility into data health and trigger notifications to the data engineering team. Which combination of Google Cloud services is best suited to implement real-time data quality monitoring and alerting for this streaming pipeline?Ensuring solution quality
  39. 239.A large manufacturing company wants to centralize its operational data from various factory sensors and machinery into a Google Cloud data lake. The data arrives in varying formats and structures, and they need to define and enforce a consistent schema for this data as it lands in Cloud Storage, before being processed by downstream analytics tools. The solution must allow for schema evolution over time without breaking existing pipelines. Which open-source file format, commonly used in the Hadoop ecosystem and supported by Google Cloud, is best suited for this requirement?Ensuring solution quality
  40. 240.A global e-commerce company uses BigQuery for its data warehouse. They experience peak loads during seasonal sales events, which can cause query performance degradation and potential service disruptions if not managed properly. The data engineering team needs to design a solution that automatically scales BigQuery resources to handle these fluctuating workloads efficiently, ensuring consistent query performance without manual intervention. Which BigQuery pricing model and feature combination should they prioritize to meet these requirements?Ensuring solution quality
  41. 241.A global marketing analytics team uses BigQuery for ad-hoc analysis and reporting. They frequently run complex queries that scan petabytes of data, leading to high unpredictable costs under the on-demand pricing model. The team needs to optimize costs while ensuring sufficient compute capacity for their analytical workloads, which have fluctuating peak demands. They want to avoid large, fixed monthly commitments if possible. Which BigQuery pricing model and feature combination offers the best balance of cost predictability, flexibility, and performance for this scenario?Ensuring solution quality
  42. 242.A financial institution is implementing a data lake on Google Cloud. They need to ensure that all data at rest is encrypted, and they must have complete control over the encryption keys due to strict regulatory compliance requirements. The solution should also integrate seamlessly with Google Cloud services like Cloud Storage and BigQuery. Which Google Cloud service should they use to manage their encryption keys?Ensuring solution quality
  43. 243.A media streaming service processes billions of user interaction events (clicks, views, searches) daily using Pub/Sub and Dataflow. The service needs to detect anomalies in real time, such as sudden spikes in error rates or unusual user behavior patterns, to trigger alerts for the operations team. You need to implement a robust monitoring and alerting solution that integrates with these services. Which Google Cloud service combination is most appropriate for this task?Ensuring solution quality
  44. 244.A large enterprise is building a new data platform on Google Cloud that integrates data from various on-premises and cloud sources. They need a centralized mechanism to enforce consistent data quality rules, manage metadata, and discover data assets across their entire data landscape. The solution should also facilitate data sharing and governance for different business units. Which Google Cloud service is designed to provide a unified data management platform for data governance, discovery, and quality across an organization's disparate data assets?Ensuring solution quality
  45. 245.A gaming company is building a real-time analytics platform to track player engagement and in-game purchases. The platform needs to ingest millions of events per second, process them with very low latency (sub-second), and store them for immediate querying. Data durability and ordering are critical to ensure accurate analytics. You need to select the most appropriate Google Cloud service for event ingestion that meets these requirements. Which service should you choose?Ensuring solution quality
  46. 246.A financial services company is migrating its on-premises data warehouse to BigQuery. They have strict requirements for data reliability, including the ability to recover data to a specific point in time in case of accidental deletions or corruptions, without relying on manual backups. The data in BigQuery is updated frequently. Which BigQuery feature directly supports this requirement?Ensuring solution quality
  47. 247.A media company processes large volumes of user interaction data (clicks, views, searches) using a Dataflow streaming pipeline. They need to ensure high availability and fault tolerance for this pipeline, as any data loss or significant downtime directly impacts real-time analytics and user experience. The solution must automatically recover from failures without manual intervention and maintain data integrity. Which Dataflow feature is crucial for achieving these reliability and fault tolerance goals?Ensuring solution quality
  48. 248.A global media company uses Dataproc clusters for large-scale data processing and machine learning model training. They need to minimize the cost of their Dataproc clusters, especially for batch jobs that are not time-sensitive, while ensuring that critical, time-sensitive jobs still get the necessary resources. The solution must intelligently utilize cheaper, pre-emptible resources without sacrificing the stability of long-running, critical workloads. Which Dataproc feature should they leverage to achieve cost optimization for non-time-sensitive jobs?Ensuring solution quality
  49. 249.A retail company uses Dataflow for real-time processing of customer order data. During peak shopping seasons, the data volume can increase tenfold within minutes. The current Dataflow pipeline occasionally experiences delays and backlogs during these spikes, leading to stale analytics. You need to ensure the Dataflow pipeline can handle these sudden, large increases in data volume without performance degradation. Which Dataflow feature is most relevant for addressing this challenge?Ensuring solution quality
  50. 250.A multinational financial services company is migrating its on-premises data warehouses to BigQuery. They have strict data residency requirements, meaning certain sensitive financial data must remain within specific geographic regions (e.g., EU, US). They need to design their BigQuery environment to ensure data is stored and processed exclusively within the designated regions, preventing any accidental cross-region data movement. Which BigQuery dataset configuration is essential to enforce these data residency requirements?Ensuring solution quality