Professional Data Engineer practice questions

262 free questions with answers and explanations.

Practice test
  1. 151.A manufacturing company uses a machine learning model to predict equipment failures. The model's predictions are used to schedule preventative maintenance. The company has observed that while the model's technical metrics (e.g., F1-score, accuracy) remain high, the actual number of unexpected equipment downtimes has increased. This discrepancy suggests that the model is no longer effectively serving its business purpose. What type of monitoring is needed to address this gap?Operationalizing machine learning models
  2. 152.A research institution is experimenting with a new deep learning model for medical image analysis. They are trying out various neural network architectures, optimization algorithms, and hyperparameter configurations to achieve the best diagnostic accuracy. They need a systematic way to track all these experiments, compare their results (metrics, loss curves), and manage the associated model artifacts and training code to ensure reproducibility. Which Google Cloud service is designed for this purpose?Operationalizing machine learning models
  3. 153.A data scientist is experimenting with different model architectures and hyperparameter configurations for a new natural language processing (NLP) model. Each experiment generates various metrics (e.g., accuracy, loss, F1-score) and model artifacts. The team needs a systematic way to track these experiments, compare results, and manage the lineage of models to ensure reproducibility and facilitate selection of the best-performing model. Which Google Cloud service should they use?Operationalizing machine learning models
  4. 154.A social media company uses an ML model to filter out inappropriate content. The model is deployed on a Vertex AI Endpoint. To ensure responsible AI practices, they need to understand which parts of the input content (e.g., specific words, phrases, or image regions) are most influential in the model's decision to flag content as inappropriate. This understanding is crucial for auditing decisions and improving model transparency. Which Vertex AI feature can provide this insight for their deployed model?Operationalizing machine learning models
  5. 155.A healthcare provider is developing a machine learning model to predict patient readmission risk. The training data contains sensitive patient health information (PHI). To comply with HIPAA regulations, they must ensure that all data used for model training and evaluation is de-identified before being processed by the ML pipeline. Which Google Cloud service should be integrated into the data ingestion pipeline to automatically detect and transform PHI?Operationalizing machine learning models
  6. 156.A large e-commerce company needs to train a product recommendation model. The training data consists of user interaction logs, which are streamed in real-time and stored in BigQuery. The data scientists require fresh data for daily model retraining to capture trending products. To ensure the training data quality, they need to validate the schema and check for missing values before it's used by the training job. What is the most efficient way to perform these data quality checks within a streaming data pipeline that feeds BigQuery?Operationalizing machine learning models
  7. 157.An online retail company needs to train a machine learning model on a large dataset of customer reviews and product information. The dataset is frequently updated, and the training process is computationally intensive, requiring GPUs. They want a managed service that can automate the entire training workflow, including data preprocessing, model training, and hyperparameter tuning, without managing underlying infrastructure. Which Vertex AI feature is suitable for this scenario?Operationalizing machine learning models
  8. 158.A data science team is building a pipeline to train a new time-series forecasting model daily. The pipeline involves data extraction from Cloud Storage, preprocessing with Dataflow, model training using custom Python code on Vertex AI Training, and model deployment to Vertex AI Endpoints. They need a robust way to orchestrate these steps, handle dependencies, and manage failures. Which Google Cloud service is most suitable for orchestrating this end-to-end ML pipeline?Operationalizing machine learning models
  9. 159.A financial institution uses a machine learning model to detect anomalies in transaction data. The model is deployed on Vertex AI Endpoints. Due to new regulatory requirements, the institution must be able to explain the predictions made by the model for any given transaction. Which Vertex AI feature should be integrated to meet this requirement?Operationalizing machine learning models
  10. 160.A logistics company uses a machine learning model to optimize delivery routes. The model's predictions directly influence fuel consumption and delivery times. They need to understand which input features (e.g., traffic conditions, package weight, delivery time window) are most influential in the model's route recommendations, especially when a route seems suboptimal or generates unexpected results. This understanding is crucial for debugging, building trust, and potentially improving the model. Which Google Cloud service can help them gain this insight?Operationalizing machine learning models
  11. 161.A healthcare provider is developing a machine learning model to assist in disease diagnosis. The model processes sensitive patient data and must adhere to strict data privacy regulations (e.g., HIPAA). During the model training phase, the data engineering team needs to ensure that the training data is anonymized and that no personally identifiable information (PII) is exposed. Which Google Cloud service or technique should be prioritized to de-identify the sensitive data before training?Operationalizing machine learning models
  12. 162.A global e-commerce company is developing a machine learning model to detect fraudulent transactions. They need to ensure that the entire ML lifecycle—from data preparation and model training to deployment and monitoring—is automated, repeatable, and scalable across different regions. The solution must support various ML frameworks and allow for custom code execution within the pipeline steps. Which Google Cloud service is most suitable for building this automated MLOps pipeline?Operationalizing machine learning models
  13. 163.A data engineering team is building an automated ML pipeline for a credit risk assessment model. The pipeline needs to ensure that the training data used for each retraining job is consistent and traceable. Specifically, they need to know exactly which version of the raw data and which preprocessing steps were applied to generate the final training dataset. This is crucial for debugging, auditing, and reproducing model results. Which concept is being addressed here?Operationalizing machine learning models
  14. 164.A data engineering team is building a real-time anomaly detection system. The ML model requires low-latency feature serving for online predictions. Specifically, it needs to access precomputed features (e.g., average transaction amount in the last 5 minutes, number of failed logins in the last hour) that are constantly updated. Which Google Cloud service is best suited for managing and serving these time-sensitive features efficiently?Operationalizing machine learning models
  15. 165.A data science team has developed a new machine learning model to predict customer churn. The model is currently deployed to a Vertex AI Endpoint for online predictions. To ensure the model continues to perform optimally over time, the team needs to regularly retrain it with fresh data and deploy the updated version without causing downtime for the prediction service. Which Vertex AI feature should they leverage to achieve this seamless update process?Operationalizing machine learning models
  16. 166.A telecommunications company is developing an ML model to predict network congestion. The model is trained on streaming network telemetry data. To ensure efficient resource utilization and rapid iteration, they need to manage and version multiple trained models, allowing data scientists to easily discover, share, and deploy specific model versions to different environments (e.g., staging, production). Which Vertex AI service is designed to be the central repository for these trained models?Operationalizing machine learning models
  17. 167.A retail company has deployed a personalized marketing ML model. After a few months, the model's recommendations start to become less effective, leading to a decrease in conversion rates. Upon investigation, they discover that customer preferences and purchasing behaviors have significantly changed since the model was last trained, rendering the original learned patterns less relevant. What phenomenon is their model experiencing?Operationalizing machine learning models
  18. 168.A retail company wants to develop a recommendation engine using machine learning. The data scientists have built a TensorFlow model that needs to be trained on a very large dataset (petabytes) stored in BigQuery. The training process is computationally intensive and requires distributed processing. After training, the model will be deployed for online predictions. Which Google Cloud service combination should the data engineering team use to efficiently train and deploy this model?Operationalizing machine learning models
  19. 169.A global financial institution is developing a credit risk assessment model. The model's predictions are highly sensitive, and any errors can have significant financial and regulatory consequences. They need to ensure that the entire ML development and deployment process, from data ingestion to model serving, is fully auditable and can trace the origin and transformations of every data point and model version. What capability is essential to implement for this requirement?Operationalizing machine learning models
  20. 170.A large-scale logistics company uses an ML model to optimize delivery routes. The model was trained on historical traffic patterns and road conditions. However, new urban development projects and changes in traffic regulations are frequently introduced. The company needs to automatically detect when these changes cause the model's predictions to become less accurate compared to the actual delivery times, even if the input features themselves haven't significantly drifted. What type of monitoring is most crucial to implement?Operationalizing machine learning models
  21. 171.A data engineering team is developing a new machine learning model for fraud detection. The model will be deployed to predict fraudulent transactions in real-time. Due to the critical nature of fraud detection, the team needs to ensure that the model's predictions are highly reliable and that any degradation in performance is detected and addressed immediately. Which Google Cloud service is best suited for monitoring the model's performance in production?Operationalizing machine learning models
  22. 172.A manufacturing company uses an ML model to predict equipment failures. The model was trained on sensor data collected from their machinery. After deploying the model, they observe that while the model performs well on the metrics it was optimized for (e.g., F1-score), the overall business impact (e.g., reduction in unplanned downtime) is not as significant as expected. They suspect that the business objective is not fully aligned with the model's technical optimization metric. What type of monitoring should they implement to address this discrepancy?Operationalizing machine learning models
  23. 173.A research institution needs to process large scientific datasets, often terabytes in size, containing complex, unstructured and semi-structured data (e.g., sensor readings, genomic sequences, research papers). The processing involves custom algorithms implemented in Python, and jobs can run for several hours. They need a cost-effective solution for batch processing that allows them to scale compute resources on demand without managing servers. Which Google Cloud service is most suitable?Designing data processing systems
  24. 174.A data engineering team is designing a new data pipeline for a multinational corporation. The pipeline will process sensitive customer data from various regions, and the data must remain within its geographical region of origin (data residency) to comply with local regulations. Additionally, all data at rest and in transit must be encrypted, and access to the data must be strictly controlled and auditable. Which combination of Google Cloud features and services best addresses these security, privacy, and compliance requirements?Designing data processing systems
  25. 175.A healthcare provider is building a new data lake on Google Cloud to store patient health records (PHR), medical images, and research data. The data lake must enforce strict data governance policies, including data quality checks, metadata management, and lineage tracking, to comply with HIPAA and other healthcare regulations. The provider needs a unified way to discover, manage, and govern diverse data assets across different Google Cloud services. Which Google Cloud service is best suited to meet these data governance requirements?Designing data processing systems
  26. 176.A media company streams live events globally. They need to analyze viewer engagement, concurrent users, and stream quality in real-time to adjust content delivery and troubleshoot issues immediately. The system must handle sudden spikes in viewership (millions of concurrent users) and provide dashboards with metrics updated every few seconds. Which architectural pattern should they implement?Designing data processing systems
  27. 177.A media streaming company is building a new recommendation engine that suggests content to users based on their real-time viewing habits. The engine requires extremely low-latency access to user profiles and viewing history, with consistent read and write performance even under high concurrency. The data model is relatively simple, consisting of key-value pairs and wide-column structures. Which Google Cloud database service should they choose?Designing data processing systems
  28. 178.A global IoT company collects sensor data from millions of devices worldwide. This data is critical for operational monitoring and predictive maintenance. The data arrives as high-volume, low-latency streams and needs to be stored in a time-series optimized database that can handle petabytes of data with very fast writes and efficient queries across time ranges. The solution must also support high availability and replication across multiple regions to ensure business continuity. Which Google Cloud service is best suited for storing this IoT sensor data?Designing data processing systems
  29. 179.A data analytics company has developed a proprietary machine learning model for fraud detection. The model needs to be updated daily with new training data, which amounts to several terabytes. The training process is computationally intensive and takes about 6 hours. The company wants to run this training job efficiently, ensuring that resources are provisioned only when needed and scaled appropriately for the workload, minimizing idle costs. Which Google Cloud service is most suitable for orchestrating and executing this daily batch ML training job?Designing data processing systems
  30. 180.A logistics company needs to process real-time updates from thousands of delivery vehicles. Each vehicle sends location, speed, and status updates every few seconds. This data needs to be ingested, transformed, and then stored for both real-time operational dashboards and historical analysis. The company requires a fully managed, scalable solution that minimizes operational overhead and supports flexible schema evolution. Which set of Google Cloud services would be most appropriate for this data pipeline?Designing data processing systems
  31. 181.A financial institution processes millions of transactions daily. Due to regulatory requirements, all transaction data must be retained for 7 years for auditing purposes, but only the most recent 90 days are frequently accessed for operational reporting. The remaining historical data is rarely accessed, perhaps once or twice a year for compliance audits. The institution wants to minimize storage costs while ensuring data availability and integrity over the entire retention period. Which BigQuery feature should they leverage for cost optimization?Designing data processing systems
  32. 182.A data engineering team is responsible for maintaining several critical data pipelines that feed a central data warehouse. They need a solution to ensure data quality and track data lineage across various transformations. Specifically, they want to automatically discover data assets, understand their relationships, and identify potential data quality issues before data reaches the production data warehouse. Which Google Cloud service is designed to address these data governance and data quality requirements?Designing data processing systems
  33. 183.A large e-commerce company needs to store customer order history, which can grow to petabytes of data over time. This data is primarily used for complex analytical queries, such as identifying purchasing trends, customer segmentation, and quarterly sales reporting. The system must be highly scalable, performant for analytical workloads, and cost-effective for storing massive datasets. Which Google Cloud service is the most appropriate for this data warehousing requirement?Designing data processing systems
  34. 184.A startup is building a new mobile application that generates user activity logs. They expect rapid growth, potentially leading to millions of events per second. The application needs a highly available and durable messaging service to decouple event producers from consumers, allowing for flexible scaling of downstream analytics and machine learning systems. The service must also handle message retention for at least 7 days for reprocessing in case of consumer failures. Which Google Cloud service is the most appropriate for this messaging requirement?Designing data processing systems
  35. 185.A retail company uses BigQuery for its data warehouse and needs to optimize costs. They have identified several tables containing historical data that is accessed infrequently (less than once a month) but must be retained for compliance reasons. These tables do not require the same query performance as frequently accessed data. Which BigQuery feature should be utilized to store this infrequently accessed data more cost-effectively?Designing data processing systems
  36. 186.A financial institution needs to design a data processing system on Google Cloud to handle sensitive customer transaction data. The system must ensure that data at rest is encrypted, and access is strictly controlled based on the principle of least privilege. Data must be available for real-time fraud detection and also for daily batch reporting. The security team mandates that all encryption keys for sensitive data must be managed externally by a dedicated Hardware Security Module (HSM) and never directly exposed to Google Cloud services. Which Google Cloud service combination should be used to meet these requirements, specifically for key management and data storage?Designing data processing systems
  37. 187.A marketing analytics team needs to analyze user clickstream data from their website. The data arrives continuously and needs to be processed to identify user sessions, enrich with CRM data, and then loaded into a data warehouse for daily reporting. The team requires a solution that minimizes operational overhead and ensures that data is processed correctly, even if individual processing steps fail. Which Google Cloud service should be used for the stream processing and transformation part of this pipeline?Designing data processing systems
  38. 188.A data analytics team needs to perform complex, ad-hoc queries on petabytes of historical sales data. The queries often involve large joins and aggregations across multiple tables. The solution must be fully managed, highly scalable, and cost-effective, with pricing based on actual query usage rather than provisioned compute resources. Which Google Cloud data warehousing service is most appropriate for this scenario?Designing data processing systems
  39. 189.A media company collects vast amounts of video metadata, user engagement logs, and content consumption patterns. They need to build a data pipeline to process this data for personalized content recommendations, audience segmentation, and content optimization. The data arrives continuously and requires transformations and aggregations before being stored in a data warehouse. The solution must be cost-effective, serverless, and provide a unified programming model for both batch and streaming data to simplify development and maintenance. Which Google Cloud service best fits the data processing requirements for transformations and aggregations?Designing data processing systems
  40. 190.A large enterprise has a critical batch processing job that runs nightly, transforming and loading data from various sources into their data warehouse. This job is highly sensitive to failures, and any interruption can lead to significant business impact. The design requires the data processing system to automatically recover from node failures, ensure data integrity, and minimize manual intervention. Which design principle is being prioritized, and which Google Cloud feature directly supports this principle for batch jobs?Designing data processing systems
  41. 191.A data engineering team is building a pipeline to process sensitive customer data. They need to ensure that the data is encrypted at rest and in transit, and that access to the data is strictly controlled based on user roles and data classifications. They also need to be able to audit all data access. Which combination of Google Cloud services would best meet these security and compliance requirements?Designing data processing systems
  42. 192.An advertising company needs to analyze user engagement data from their mobile applications. This data arrives in various formats and needs to be standardized, cleaned, and transformed before being loaded into their BigQuery data warehouse for business intelligence and reporting. The data processing pipeline must be scheduled to run daily in a batch manner, be fault-tolerant, and handle schema evolution gracefully without manual intervention. Which Google Cloud service should be used for building this robust batch data pipeline?Designing data processing systems
  43. 193.A global e-commerce company wants to build a new recommendation engine. The engine needs to process user clickstream data in real-time, enrich it with user profile information from a low-latency NoSQL database, and then feed the enriched data to a machine learning model for immediate prediction. The solution must handle millions of events per second and maintain sub-100ms latency for enrichment. Which Google Cloud services should be chosen for the low-latency NoSQL database and the real-time stream processing?Designing data processing systems
  44. 194.A global marketing analytics company needs to process clickstream data from millions of users across various websites and mobile apps. The data is generated continuously, and they need to perform near real-time aggregations (e.g., hourly unique visitors, page views per campaign) to feed into live dashboards. Additionally, this processed data must be stored for long-term historical analysis and machine learning model training. The solution must be cost-effective and highly available. Which approach should be used to build this data pipeline?Designing data processing systems
  45. 195.A data science team is developing a recommendation engine that needs to perform real-time feature lookups for millions of users. The engine requires a database that can store complex, semi-structured user profiles and item metadata, support fast key-value lookups (sub-10ms latency), and scale globally. The data is frequently updated but doesn't require strong transactional consistency across multiple records. Which Google Cloud service is the most suitable for this use case?Designing data processing systems
  46. 196.A startup is building a new mobile application that requires a database to store user preferences and application state. This database needs to offer real-time synchronization across devices, scale effortlessly with user growth, and support flexible, schema-less data structures. The startup aims to minimize operational overhead and prefers a serverless solution. Which Google Cloud database service is the most suitable choice?Designing data processing systems
  47. 197.A global IoT company collects sensor data from millions of devices, generating terabytes of time-series data daily. This data needs to be retained for several years for compliance and long-term analysis, but only the most recent data (last 30 days) requires high-performance, low-latency queries for real-time monitoring dashboards. Older data can be accessed with slightly higher latency but must remain cost-effective. How should this data be stored and queried efficiently on Google Cloud?Designing data processing systems
  48. 198.A retail company processes millions of daily customer transactions. They need to build a data pipeline that can ingest these transactions in real-time, perform immediate validation and enrichment, and then store them in a data warehouse for analytical reporting. The system must be highly scalable to handle peak loads during sales events and guarantee exactly-once processing to prevent data inconsistencies. Which Google Cloud services should the architect recommend for the ingestion and real-time processing components?Designing data processing systems
  49. 199.A financial institution is migrating its on-premises data warehouse to Google Cloud. The existing data warehouse contains petabytes of sensitive customer and transaction data that must comply with strict regulatory requirements for data residency, encryption, and access control. The institution needs to perform complex analytical queries on this data while ensuring cost-effectiveness and scalability. Which Google Cloud service is most appropriate for this data warehousing requirement?Designing data processing systems
  50. 200.A global gaming company collects telemetry data from millions of active users. This data, consisting of events like game progress, purchases, and errors, arrives continuously at extremely high volumes (billions of events per day). The company needs to perform real-time aggregations for leaderboards and fraud detection, and also store the raw data for historical analysis and machine learning model training. The solution must be highly available and cost-effective. Which architecture should the company implement?Designing data processing systems