Professional Data Engineer practice questions
262 free questions with answers and explanations.
- 101.A data engineering team is developing a new real-time analytics pipeline. During testing, they observe that data is occasionally arriving out of order and with duplicate records from upstream sources. This is causing incorrect aggregations in their streaming Dataflow jobs. They need a mechanism to handle these issues reliably before data is written to their analytical database. Which Apache Beam feature, when implemented in their Dataflow pipeline, would directly address these challenges?Building and operationalizing data processing systems
- 102.A data engineering team is building a complex data pipeline that involves ingesting data from on-premises databases, transforming it with custom Python scripts, and loading the results into BigQuery. The pipeline has several interdependent stages, including data validation, enrichment, and aggregation, all scheduled to run daily. They need a managed service that can orchestrate these tasks, handle dependencies, and provide robust scheduling and monitoring capabilities. Which Google Cloud service should they choose?Building and operationalizing data processing systems
- 103.A research institution is building a data lake on Google Cloud using Cloud Storage. They frequently ingest various unstructured and semi-structured datasets from different research partners. To ensure proper data governance and discoverability, they need a centralized metadata catalog that automatically extracts schema and technical metadata, allows for business metadata tagging, and supports data lineage tracking across their datasets. Which Google Cloud service is best suited for this requirement?Managing and securing data
- 104.A data engineering team is designing a new batch processing pipeline on Google Cloud to analyze historical sales data from multiple regional databases. The total data volume is in the hundreds of terabytes, and new data arrives daily, requiring a full refresh of the analytical dataset. The team needs a highly scalable, fully managed service that can perform complex SQL transformations efficiently and integrate well with BigQuery for final analysis. Which Google Cloud service should they choose for the data transformation step?Building and operationalizing data processing systems
- 105.A global ride-sharing company collects real-time GPS coordinates from millions of active vehicles. This data is ingested into Cloud Pub/Sub and then processed by a Dataflow streaming pipeline. The Dataflow job aggregates vehicle locations every 30 seconds to update a real-time map. During peak hours, the number of active vehicles can surge dramatically, causing the Dataflow job to fall behind, leading to increased processing latency and stale map updates. The data engineering team needs to ensure that the Dataflow job can dynamically scale to handle these unpredictable traffic spikes. Which Dataflow feature should they enable and configure?Building and operationalizing data processing systems
- 106.A global media streaming platform uses Pub/Sub for ingesting real-time user activity data. Due to potential network disruptions or subscriber processing delays, they need to ensure that messages are not lost and can be reprocessed if an issue occurs. Specifically, messages should be retained for up to 7 days, and any messages that fail to be acknowledged after multiple delivery attempts should be automatically moved to a separate dead-letter topic for manual inspection and reprocessing. Which Pub/Sub features should they configure?Managing and securing data
- 107.A healthcare provider is storing patient medical records in Cloud Storage. Due to regulatory compliance, these records must be encrypted at rest and accessible only by a specific service account that manages data ingestion. Furthermore, the encryption keys must be rotated annually. Which encryption and access control strategy should be used?Managing and securing data
- 108.A financial services company is migrating its on-premises transactional database to Google Cloud. The database contains highly sensitive customer financial data that requires stringent access controls and encryption at rest. The company needs to ensure that only authorized applications and users can access specific tables and columns, and that all data is encrypted with customer-managed encryption keys (CMEK). Which Google Cloud storage service and access control mechanism should you recommend to meet these requirements?Managing and securing data
- 109.A healthcare provider stores sensitive patient medical images in Cloud Storage. Due to strict HIPAA regulations, these images must be encrypted using keys that are maintained physically isolated in a FIPS 140-2 Level 3 certified hardware security module (HSM). The provider wants to use Google Cloud services where possible but needs to integrate their existing on-premises HSMs for key management. Which encryption approach should they use?Managing and securing data
- 110.A data analytics team is analyzing a large dataset in BigQuery. They have a table with over a billion rows of customer transactions, including `transaction_date` (DATE), `customer_id` (STRING), `product_category` (STRING), and `amount` (NUMERIC). Analysts frequently query the data to filter by `transaction_date` for specific date ranges and also often filter and group by `product_category`. They want to optimize query performance and reduce query costs. How should the table be designed?Building and operationalizing data processing systems
- 111.A data engineering team is building a data pipeline that ingests sensitive customer data from various sources into BigQuery. They need to ensure that the encryption keys used for this data are highly secure and meet stringent compliance requirements, including FIPS 140-2 Level 3. Which type of Cloud Key Management Service (KMS) key should they use?Managing and securing data
- 112.A data engineering team is troubleshooting a Dataflow job that processes real-time sensor data. They observe that some incoming data records are processed with significant delays, leading to outdated results in their analytics dashboard. The Dataflow monitoring interface shows a consistently increasing 'System Lag' metric, while 'Data freshness' is also high. The pipeline uses fixed windows and processes data based on event time. What is the most likely cause of this issue?Building and operationalizing data processing systems
- 113.A gaming company is using Cloud Spanner for its global leaderboard, which stores player IDs and scores. Due to privacy regulations, player IDs must be pseudonymized to prevent direct identification while still allowing for internal analytics that require joining with other datasets. The pseudonymization must be reversible for specific authorized processes (e.g., customer support). How should you implement this privacy requirement?Managing and securing data
- 114.A financial institution needs to store immutable audit logs for regulatory compliance. These logs are written once and rarely accessed after initial ingestion, but must be retained for 7 years. Cost optimization for storage is a critical factor, as the volume is expected to grow into petabytes. Which Cloud Storage class is most appropriate for this use case?Building and operationalizing data processing systems
- 115.A retail company collects customer clickstream data from its website and mobile applications. This data needs to be transformed to enrich user sessions with demographic information from a CRM system, filter out bot traffic, and aggregate events for product recommendation engines. The company requires a serverless, horizontally scalable solution that can handle varying data volumes efficiently. Which Google Cloud service should be used for this transformation?Building and operationalizing data processing systems
- 116.A large enterprise uses BigQuery as its primary data warehouse. They have a dataset containing highly sensitive financial transaction data that needs to be retained for 7 years due to regulatory compliance. However, after 2 years, the data is rarely accessed, and after 5 years, it's almost never accessed. The data engineering team needs a cost-effective solution to manage the lifecycle of this data, minimizing storage costs while still meeting the 7-year retention requirement. Which BigQuery feature combination should they implement?Managing and securing data
- 117.A media company stores large volumes of video archives in Cloud Storage. These archives are frequently accessed for the first 30 days after upload for editing, then rarely accessed for the next 90 days, and finally, after 120 days, they are almost never accessed but must be retained for legal compliance for 5 years. They need to optimize storage costs by automatically transitioning these objects between storage classes. Which Cloud Storage feature should they configure?Managing and securing data
- 118.A media company is building a data pipeline to process user engagement metrics from their streaming platform. They need to aggregate data into 1-hour windows. Due to the distributed nature of their platform, some events might arrive slightly out of order, and occasionally, a few events might be delayed by up to 5 minutes. The team wants to ensure that these slightly delayed events are included in the correct 1-hour window without significantly impacting the overall processing latency for the majority of on-time events. How should they configure their Dataflow windowing and watermark settings?Building and operationalizing data processing systems
- 119.A media company stores large volumes of video archives in Cloud Storage. These archives are rarely accessed after the first 30 days but must be retained for 10 years due to regulatory requirements. The company wants to optimize storage costs while ensuring data availability for the full retention period. What Cloud Storage class and lifecycle management policy should be implemented?Managing and securing data
- 120.A data engineering team is building a new real-time fraud detection system. The system needs to ingest millions of events per second from various sources, process them with minimal latency, and then store the results for immediate lookup. Data loss is unacceptable, and the system must scale dynamically to handle peak loads. Which combination of Google Cloud services would provide the most efficient and reliable solution for ingesting and processing these events?Building and operationalizing data processing systems
- 121.A data analytics team is developing a real-time anomaly detection system using Dataflow. The system processes a stream of events and needs to ensure that each event is processed exactly once, even in the case of worker failures or restarts, to avoid false positives or negatives in anomaly detection. Which Dataflow feature guarantees this critical processing characteristic?Building and operationalizing data processing systems
- 122.A data analytics team is migrating an on-premises Hadoop cluster to Google Cloud. They need to store petabytes of historical log data and frequently accessed reports, ensuring high durability and availability. The data will be accessed by various analytics tools and machine learning workloads. Which storage service should they choose for this data lake?Building and operationalizing data processing systems
- 123.A data engineering team is developing a new real-time fraud detection system. The system ingests a high-volume stream of financial transactions and needs to perform complex, stateful aggregations (e.g., calculating the sum of transactions for a user within the last 5 minutes) before sending alerts. The team requires a managed service that can handle these aggregations with low latency and high reliability, scaling automatically with transaction volume. Which Google Cloud service is the most appropriate for processing this data?Building and operationalizing data processing systems
- 124.A large e-commerce company uses BigQuery as its primary data warehouse. They have a daily batch job that loads millions of new customer orders into a `raw_orders` table. Analysts frequently query this table, filtering by `order_date` and `customer_id`. The table currently has over a billion rows and queries are becoming slow and expensive. The team needs to optimize query performance and reduce costs for these common analytical queries. Which BigQuery feature should be implemented to address this issue?Building and operationalizing data processing systems
- 125.A retail company uses BigQuery for analyzing customer purchasing behavior. To comply with GDPR, they must ensure that customer data is deleted after 7 years, but aggregated, anonymized purchasing trends need to be retained indefinitely. How should you manage the data lifecycle in BigQuery to meet these requirements?Managing and securing data
- 126.A global manufacturing company uses BigQuery for its operational analytics. They have several datasets containing sensitive production data, including intellectual property details. The company needs to implement a solution that allows different teams to access specific subsets of columns within these tables based on their roles, without creating multiple copies of the data. For example, the R&D team needs to see all columns, while the sales team only needs product ID and quantity. Which BigQuery feature should be used to achieve this column-level access control efficiently?Managing and securing data
- 127.A research institution is building a data lake on Google Cloud using Cloud Storage. They frequently receive new datasets from various partners, which are stored in different buckets. To ensure data governance, they need to apply consistent metadata (e.g., data source, department, sensitivity level) to all incoming data files for easier discovery and regulatory compliance. How should the institution implement this metadata management?Managing and securing data
- 128.A data engineering team is building a new batch processing pipeline using Dataflow to analyze daily sales data. The pipeline reads data from Cloud Storage, performs complex transformations, and writes the results to BigQuery. The current Dataflow job is consistently taking longer than expected, impacting the downstream reporting deadlines. The team has observed that the job spends a significant amount of time in the 'Shuffle' phase, indicating high data movement between workers. What is the most effective strategy to optimize the performance of this Dataflow job?Building and operationalizing data processing systems
- 129.A gaming company is experiencing intermittent spikes in latency and error rates in their game analytics pipeline. This pipeline uses Cloud Pub/Sub for ingestion, Dataflow for processing, and BigQuery for storage. They need to quickly identify the root cause of these performance issues. Which Google Cloud service combination provides the most effective solution for monitoring and troubleshooting this multi-service pipeline?Building and operationalizing data processing systems
- 130.A global IoT company collects sensor data from millions of devices worldwide. This data is ingested into Cloud Pub/Sub and then processed by a Dataflow pipeline. The company needs to optimize costs by ensuring that the Dataflow pipeline only uses the necessary compute resources at any given time, dynamically adjusting to varying data ingestion rates throughout the day. Which Dataflow feature should be configured to achieve this cost optimization?Building and operationalizing data processing systems
- 131.A pharmaceutical company stores sensitive clinical trial data in BigQuery. The data must be accessible to researchers for analysis, but direct access to the raw tables is prohibited to prevent accidental exposure of PII. Researchers should only be able to query a subset of the data (e.g., de-identified patient outcomes) that is pre-filtered and aggregated. All access to this derived data must also be auditable. Which BigQuery feature, combined with appropriate logging, should be used to achieve this?Managing and securing data
- 132.A global retail company uses BigQuery for its analytics platform. They have a dataset containing customer personally identifiable information (PII) that must be protected according to GDPR. Data analysts need to query this data for aggregate reports, but must never see the raw PII. The security team wants a solution that allows different levels of access to specific columns based on user roles, without creating multiple copies of the data. Which BigQuery feature should the data engineering team implement?Managing and securing data
- 133.A media company stores large volumes of video assets in Cloud Storage. These assets are frequently accessed for the first 30 days after upload, then accessed infrequently for the next 90 days, and rarely accessed thereafter, but must be retained for 5 years for archival purposes. They want to minimize storage costs while maintaining accessibility. Which Cloud Storage lifecycle management policy would be most cost-effective?Managing and securing data
- 134.A global e-commerce company uses BigQuery for analyzing customer purchasing behavior. They have a dataset containing sensitive customer information, including personally identifiable information (PII) like email addresses and phone numbers. To comply with privacy regulations, they need to ensure that this PII is not directly exposed to analysts but can be used for aggregated reporting. The solution must allow for different levels of obfuscation based on the analyst's role (e.g., some analysts might see only the domain of an email, others only a masked version). Which BigQuery feature should they implement?Managing and securing data
- 135.A data engineering team is building a real-time recommendation engine. They need to store user interaction data (clicks, views, purchases) that is constantly being updated and accessed with low latency. The data model is flexible and may evolve over time. The system needs to support high read and write throughput for millions of users globally, and strong consistency is not always required for every read. Which Google Cloud database service is most suitable?Building and operationalizing data processing systems
- 136.An e-commerce company is building a real-time product recommendation system. The system needs to serve fresh, personalized recommendations with very low latency (under 50ms). The ML model relies on frequently updated user interaction data (e.g., recent clicks, purchases) and product catalog information. Which Google Cloud service is best suited for managing and serving these features to the online prediction model?Operationalizing machine learning models
- 137.A data science team has developed a new image classification model. They want to deploy this model to serve online predictions with low latency. The model is packaged as a custom container image. They anticipate highly variable traffic patterns, with potential spikes, and need the deployment to scale automatically while minimizing costs during idle periods. Which Google Cloud deployment option is best suited for these requirements?Operationalizing machine learning models
- 138.A data science team is developing a critical ML model for fraud detection. They need to ensure that the model development and deployment process is reproducible, auditable, and allows for easy rollback to previous versions if issues arise. Multiple data scientists will be collaborating on the model, experimenting with different algorithms and hyperparameters. Which Google Cloud service should they use to manage their model artifacts and metadata throughout their lifecycle?Operationalizing machine learning models
- 139.A financial institution is building a credit scoring model. They need to ensure that the data used for training the model is consistent, accurate, complete, and timely. Inconsistent data could lead to biased predictions, and missing data could reduce model performance significantly. They want to implement checks throughout their data pipeline, from ingestion to feature engineering, to maintain high data quality for ML. What is a key practice they should adopt to ensure data quality for their ML model?Operationalizing machine learning models
- 140.A healthcare provider is developing a machine learning model to assist in disease diagnosis. Due to strict regulatory compliance (HIPAA, GDPR), they must ensure that no personally identifiable information (PII) or protected health information (PHI) is inadvertently used during model training or exposed during inference. They need a mechanism to identify and redact sensitive data within their datasets before it reaches the ML pipeline. Which Google Cloud service should they integrate into their data preparation workflow?Operationalizing machine learning models
- 141.A pharmaceutical company is training a deep learning model for drug discovery. The training process involves iterating through hundreds of different model architectures, hyperparameter configurations, and datasets. They need a systematic way to track each experiment's inputs (code, configurations, data versions), outputs (metrics, models), and ensure reproducibility. Which Vertex AI service is best suited for managing this iterative and experimental workflow?Operationalizing machine learning models
- 142.A media company uses an ML model to generate personalized content recommendations. The model's performance is critical for user engagement. They notice a sudden drop in recommendation quality, but conventional data drift metrics on input features haven't flagged any significant issues. Upon investigation, they find that while the input data distribution (e.g., user demographics, content categories) hasn't changed, the relationship between user preferences and content features has shifted over time. What type of model degradation is most likely occurring?Operationalizing machine learning models
- 143.A financial institution is developing a machine learning model to detect fraudulent transactions. Due to strict regulatory compliance requirements, all data used for training and inference must be stored within a specific geographical region and adhere to data residency policies. Which Google Cloud service is most appropriate to ensure that the datasets for this ML model are consistently stored and processed within the required region?Operationalizing machine learning models
- 144.A global e-commerce company is building an automated ML pipeline to generate personalized product recommendations. The pipeline involves data ingestion, preprocessing (feature engineering), model training, evaluation, and deployment. Each stage needs to be orchestrated, allowing for retry logic, conditional execution based on previous stage outcomes, and easy visualization of the pipeline's progress. Which Google Cloud service is specifically designed for building and managing such complex, multi-step ML workflows?Operationalizing machine learning models
- 145.A retail company has deployed a personalized marketing ML model. After a few months, the marketing team reports that the model's recommendations are becoming less relevant, and customer engagement is decreasing. Upon investigation, the data engineering team finds that the distribution of customer demographics (e.g., age groups, purchase history patterns) has subtly shifted over time, impacting the model's performance. This phenomenon is known as concept drift. What is the most effective strategy for managing this issue?Operationalizing machine learning models
- 146.A large e-commerce company is building an automated ML pipeline to generate personalized product recommendations. The pipeline involves data ingestion, feature engineering, model training, evaluation, and deployment. They need a robust orchestration service that can manage complex dependencies, handle failures gracefully, and allow for scheduling and monitoring of the entire workflow. Which Google Cloud service is the most suitable for orchestrating this ML pipeline?Operationalizing machine learning models
- 147.A financial services company is developing a machine learning model to predict loan default risk. The model will be integrated into their existing loan application system, which requires real-time predictions with low latency. The company needs to ensure that the model is highly available and can scale automatically to handle fluctuating request volumes, especially during peak application periods. Which Google Cloud service is the most appropriate for deploying this model to meet these requirements?Operationalizing machine learning models
- 148.A media company is building a recommendation system for video content. They need to serve personalized recommendations to users in real-time based on their viewing history and preferences. The features for the recommendation model, such as 'user_watch_count_genre_X' or 'item_popularity_last_24h', need to be consistently computed and served with very low latency for both training and online inference. Which Google Cloud service is best suited to manage and serve these features?Operationalizing machine learning models
- 149.A data engineering team is responsible for managing several machine learning models in production. They need a centralized repository for storing model artifacts, versioning models, and tracking metadata such as training parameters, evaluation metrics, and the datasets used for training. This repository should facilitate collaboration among data scientists and ensure reproducibility. Which Google Cloud service is designed for this purpose?Operationalizing machine learning models
- 150.A retail company uses a machine learning model to optimize inventory levels. The model was trained on historical sales data. Over time, new product lines are introduced, and consumer purchasing habits shift due to market trends. The current model's performance has started to degrade, leading to suboptimal inventory recommendations. What is the most likely cause of this degradation?Operationalizing machine learning models