Step2Study
IT & Technology100% Free

Professional Data Engineer

Practice bank
262 Qs
Real exam
50 Qs
Time limit
120 min
Passing
Candidates need to achieve a minimum score to pass the exam. The passing score is not publicly disclosed.

Exam blueprint

Designing data processing systems
30%
Building and operationalizing data processing systems
25%
Operationalizing machine learning models
15%
Ensuring solution quality
15%
Managing and securing data
15%

Practice

Untimed · instant feedback · 4 practice tests of 90 questions

Questions per test

Custom practice

Flashcard on every question Mental map when you miss

Exam simulation

4 timed tests · 90 questions each · 216 min · pass 70% · 262 questions in the bank

+50 XP per test · +100 XP for a pass

Random simulation (weighted by domain)

Everything is open to everyone. Create a free account to save scores, XP, badges and get progress emails.

Free study resources

All resources →

Study with friends

Challenge a friend to beat your score.

Professional Data Engineer practice test questions

Sample questions from the 262-question bank, with answers and explanations.

All questions
  1. 1. A healthcare provider is building a new system to store anonymized patient records. They need a highly available, globally consistent, transactional database that can scale horizontally to handle millions of records and thousands of concurrent read/write operations. Data integrity and strong consistency are paramount. Which Google Cloud database service should they choose?

    Building and operationalizing data processing systems

    • A. Cloud SQL for PostgreSQL
    • B. Firestore
    • C. BigQuery
    • D. Cloud Spanner
    Show answer

    D. Cloud Spanner

    Cloud Spanner is a globally distributed, strongly consistent, and highly available relational database service. It offers transactional consistency and horizontal scalability, making it ideal for the described requirements of sensitive patient records.

  2. 2. A financial institution needs to process daily batch files containing millions of customer transactions. These files arrive at a specific time each day and must be processed within a strict 2-hour window. The processing involves complex transformations and aggregations, which can be computationally intensive. The company wants to optimize the Dataflow job for both performance (completing within 2 hours) and cost, ensuring that resources are perfectly matched to the workload without over-provisioning. Which Dataflow Prime feature is specifically designed to address this challenge?

    Building and operationalizing data processing systems

    • A. Streaming Engine
    • B. Vertical Autoscaling
    • C. Right-Fitting
    • D. Dataflow Shuffle
    Show answer

    C. Right-Fitting

    Dataflow Prime's Right-Fitting feature intelligently analyzes the pipeline's workload and automatically selects the optimal machine types and resource configurations (CPU, memory, disk) for each stage of the pipeline, ensuring efficient processing within performance targets while minimizing costs, especially for batch jobs with strict deadlines.

  3. 3. A global ride-sharing company collects real-time GPS coordinates from millions of active vehicles. This data needs to be ingested into Google Cloud for immediate processing by a Dataflow streaming pipeline. The ingestion system must handle extreme, unpredictable spikes in message volume, scale automatically without manual intervention, and guarantee high durability to prevent data loss. Which Google Cloud service is best suited to ingest this data reliably and scalably?

    Building and operationalizing data processing systems

    • A. Cloud Storage
    • B. Cloud Spanner
    • C. Cloud SQL
    • D. Cloud Pub/Sub
    Show answer

    D. Cloud Pub/Sub

    Cloud Pub/Sub is a fully managed, globally distributed messaging service designed for real-time data ingestion. It automatically scales to handle millions of messages per second, provides high durability, and decouples publishers from subscribers, making it ideal for this scenario.

  4. 4. A global online gaming company needs to store petabytes of user gameplay data, including session logs, in-game events, and player statistics. This data is characterized by extremely high write throughput (millions of writes per second), low-latency reads for real-time leaderboards and player profiles, and requires horizontal scalability to handle unpredictable traffic spikes. Data consistency is eventually consistent for most use cases, but strong consistency is preferred where possible. Which Google Cloud database service is best suited for this workload?

    Building and operationalizing data processing systems

    • A. Cloud Bigtable
    • B. Cloud SQL
    • C. BigQuery
    • D. Cloud Spanner
    Show answer

    A. Cloud Bigtable

    Cloud Bigtable is a wide-column NoSQL database designed for very large analytical and operational workloads, offering extremely high write and read throughput at low latency, horizontal scalability, and is ideal for time-series data like gameplay logs.

  5. 5. A financial institution is implementing a new data warehousing solution on Google Cloud. They store highly sensitive customer financial data in BigQuery. Due to regulatory requirements, they must ensure that data at rest is encrypted with customer-managed encryption keys (CMEK), and these keys must be automatically rotated annually. Which Google Cloud service should they use to manage and automatically rotate these encryption keys for BigQuery?

    Managing and securing data

    • A. Cloud Storage
    • B. Cloud Key Management Service (KMS)
    • C. Secret Manager
    • D. Identity and Access Management (IAM)
    Show answer

    B. Cloud Key Management Service (KMS)

    Cloud Key Management Service (KMS) is specifically designed for managing cryptographic keys, including customer-managed encryption keys (CMEK), and supports automatic key rotation, which aligns with the regulatory requirement for annual rotation.

  6. 6. A gaming company is experiencing intermittent spikes in latency and error rates in their game analytics pipeline, which uses Cloud Pub/Sub for ingestion and Dataflow for real-time processing. They need a comprehensive solution to quickly detect, diagnose, and alert on these issues across their entire data processing stack. Which combination of Google Cloud services would provide the best solution?

    Building and operationalizing data processing systems

    • A. Cloud Debugger and Cloud Trace
    • B. Cloud Monitoring and Cloud Logging
    • C. Cloud Audit Logs and Cloud Security Command Center
    • D. Cloud Trace and Cloud Logging
    Show answer

    B. Cloud Monitoring and Cloud Logging

    Cloud Monitoring provides metrics and alerting for performance issues like latency and error rates, while Cloud Logging collects and stores detailed logs for diagnosis, making them a powerful combination for operational visibility.

  7. 7. A global manufacturing company uses BigQuery for its operational analytics. They have several datasets containing sensitive production data, including intellectual property (IP) related to product designs. They need to ensure that only specific engineering teams can view the 'design_specifications' column in the `production_designs` table, while other teams can still access other columns in the same table for general reporting. How should they implement this granular access control?

    Managing and securing data

    • A. Implement BigQuery column-level security using policy tags on the 'design_specifications' column.
    • B. Apply BigQuery data masking to the 'design_specifications' column to hide its content from unauthorized users.
    • C. Create an authorized view that includes all columns except 'design_specifications' for general users.
    • D. Store 'design_specifications' in a separate table and restrict access to that table using Cloud IAM.
    Show answer

    A. Implement BigQuery column-level security using policy tags on the 'design_specifications' column.

    BigQuery column-level security, implemented with policy tags, is designed precisely for this scenario. It allows you to restrict access to specific columns within a table based on user roles, ensuring that only authorized users (e.g., specific engineering teams) can view the sensitive column while others can still access the rest of the table.

  8. 8. A financial services company is building a data lake on Google Cloud. They need to store raw, semi-structured transaction data (e.g., JSON, CSV) from various sources, including on-premises databases and third-party APIs. The data volume is expected to grow to petabytes, and it will be accessed by different analytical tools (BigQuery, Spark) for ad-hoc queries and machine learning models. The solution must be highly durable, globally accessible, and cost-effective for long-term storage, with flexible schema support. Which Google Cloud storage service is most appropriate for their data lake's raw data storage layer?

    Building and operationalizing data processing systems

    • A. Bigtable
    • B. Cloud Spanner
    • C. Cloud SQL
    • D. Cloud Storage
    Show answer

    D. Cloud Storage

    Cloud Storage is the ideal choice for a data lake's raw data storage layer. It offers petabyte-scale capacity, high durability, global accessibility, and cost-effectiveness. Its object-based nature makes it schema-agnostic, supporting various file formats (JSON, CSV) and allowing seamless integration with analytical tools like BigQuery (via external tables) and Spark (via Cloud Dataproc).

  9. 9. A data engineering team is developing a complex batch data pipeline that involves data ingestion from various sources (Cloud Storage, external APIs), transformation using custom Python scripts, and loading into BigQuery. The pipeline has multiple dependencies between stages, conditional logic based on data availability, and needs to be scheduled daily. The team wants a fully managed service that allows them to define, schedule, and monitor these workflows using a familiar open-source tool. Which Google Cloud service meets these requirements?

    Building and operationalizing data processing systems

    • A. Cloud Functions
    • B. Cloud Composer
    • C. Cloud Dataflow
    • D. Cloud Run
    Show answer

    B. Cloud Composer

    Cloud Composer is a fully managed workflow orchestration service built on Apache Airflow. It allows data engineers to define complex data pipelines using Python (DAGs), manage dependencies, implement conditional logic, schedule runs, and monitor the entire workflow, perfectly matching the requirements for a managed, open-source orchestration tool.

  10. 10. A data team is developing a new streaming pipeline to process real-time sensor data from IoT devices. The pipeline uses Dataflow to perform aggregations and transformations before storing the results in BigQuery. They notice that during periods of high data volume, the Dataflow job's CPU utilization is consistently high, and the processing latency increases significantly. They have already increased the number of workers, but the problem persists. Upon inspecting the Dataflow UI, they see that a specific `DoFn` responsible for a complex, CPU-intensive calculation is the bottleneck. What is the most effective approach to optimize this bottleneck in the Dataflow pipeline?

    Building and operationalizing data processing systems

    • A. Switch the Dataflow job from streaming mode to batch mode.
    • B. Increase the memory allocated to each Dataflow worker.
    • C. Implement a custom windowing strategy to reduce the frequency of aggregations.
    • D. Refactor the CPU-intensive `DoFn` to be more efficient, potentially offloading parts to a specialized service.
    Show answer

    D. Refactor the CPU-intensive `DoFn` to be more efficient, potentially offloading parts to a specialized service.

    If a specific `DoFn` is consistently causing high CPU utilization and latency even after scaling workers, it indicates an inefficiency in the code itself. Refactoring the `DoFn` to be more computationally efficient or offloading parts of the complex calculation to a specialized service (e.g., a custom API on Cloud Run for ML inference) is the most effective way to address this code-level bottleneck.

  11. 11. A data engineering team is building a new real-time analytics pipeline. During testing, they observe that some data events arrive out of order or with significant delays, leading to inaccurate windowed aggregations. They need to ensure that aggregations correctly account for all data within a specific time window, regardless of when events actually arrive. Which Apache Beam concept is crucial for handling this problem in Dataflow?

    Building and operationalizing data processing systems

    • A. Global Windows
    • B. Side Inputs
    • C. Session Windows
    • D. Watermarks
    Show answer

    D. Watermarks

    Watermarks in Apache Beam (and Dataflow) are used to track event time progress and determine when a window can be considered complete, allowing for the accurate handling of late-arriving data.

  12. 12. A data analytics team is migrating an on-premises data warehouse to Google Cloud. They have petabytes of historical data stored in various formats (CSV, JSON, Parquet) that need to be loaded into BigQuery for analysis. The team wants to perform initial data cleaning and transformation before loading, but they also need to support incremental daily loads. They prefer a serverless approach that minimizes operational overhead. Which Google Cloud service combination would be most suitable for this migration and ongoing data loading?

    Building and operationalizing data processing systems

    • A. Cloud SQL and Cloud Functions
    • B. Cloud Storage (as a data lake) and Dataflow (batch and streaming)
    • C. Cloud Spanner and Data Catalog
    • D. Cloud Pub/Sub and Bigtable
    Show answer

    B. Cloud Storage (as a data lake) and Dataflow (batch and streaming)

    Cloud Storage serves as a cost-effective data lake for storing raw historical data in various formats. Dataflow, in both batch mode for initial historical loads and streaming mode for incremental daily loads, provides serverless, scalable data cleaning and transformation before loading into BigQuery.

  13. 13. A data team is building a complex data pipeline that involves ingesting data from various sources (databases, APIs, files), performing transformations, and loading into BigQuery. The pipeline has dependencies between different stages, requires scheduling, error handling, and robust monitoring. They need to orchestrate this multi-step workflow in a managed, scalable, and fault-tolerant manner. Which Google Cloud service is the most appropriate for orchestrating this pipeline?

    Building and operationalizing data processing systems

    • A. Cloud Scheduler
    • B. Cloud Functions
    • C. Cloud Composer
    • D. Dataflow
    Show answer

    C. Cloud Composer

    Cloud Composer, a managed Apache Airflow service, is designed for orchestrating complex, multi-step data pipelines with dependencies, scheduling, and robust monitoring capabilities, making it ideal for the described scenario.

  14. 14. A data scientist needs to train a machine learning model using a large dataset stored in BigQuery. The training process requires performing complex statistical aggregations and feature engineering steps that are best expressed using SQL, but the dataset is too large to fit into memory on a single machine. The data scientist wants to leverage the scalability of BigQuery for these data preparation steps before exporting the final features for model training in a separate environment. Which BigQuery feature is most suitable for this scenario?

    Building and operationalizing data processing systems

    • A. BigQuery BI Engine
    • B. BigQuery Scripting
    • C. BigQuery ML
    • D. BigQuery Data Transfer Service
    Show answer

    B. BigQuery Scripting

    BigQuery Scripting allows you to write complex multi-statement SQL queries, including control flow statements and variable declarations, directly within BigQuery. This enables sophisticated data preparation, feature engineering, and iterative aggregations on very large datasets, leveraging BigQuery's scalable query engine without needing to move data out or use external processing frameworks.

  15. 15. A data engineering team is building a real-time recommendation engine. They need to store user interaction data (e.g., clicks, views, purchases) and serve personalized recommendations with extremely low latency (sub-10ms) to millions of concurrent users. The data model is relatively simple, consisting of key-value pairs and wide-column structures. High write throughput and read throughput are critical. Which Google Cloud service is best suited for this operational database requirement?

    Building and operationalizing data processing systems

    • A. Cloud Bigtable
    • B. BigQuery
    • C. Cloud SQL
    • D. Cloud Spanner
    Show answer

    A. Cloud Bigtable

    Cloud Bigtable is a fully managed NoSQL wide-column database designed for high throughput and low-latency access to large datasets, making it ideal for real-time recommendation engines and other operational analytics workloads with millions of concurrent users.

  16. 16. A financial institution is building a data warehouse in BigQuery. They need to ensure that data consumed by downstream analytical applications is trustworthy and that its origin and transformations can be tracked. Specifically, they need to visualize the path of data from its ingestion sources through various BigQuery transformations to its final destination tables. Which Google Cloud service provides this capability?

    Managing and securing data

    • A. Cloud Data Catalog Data Lineage
    • B. Cloud Audit Logs
    • C. BigQuery Information Schema
    • D. Cloud Trace
    Show answer

    A. Cloud Data Catalog Data Lineage

    Cloud Data Catalog Data Lineage provides automated tracking and visualization of data movement and transformations across Google Cloud, allowing users to understand the origin and evolution of their data assets.

  17. 17. A data engineering team is developing a new streaming pipeline to process real-time sensor data. The pipeline runs on Dataflow, and the processed data is eventually stored in BigQuery. The team needs to ensure that the Dataflow workers can securely communicate with BigQuery and other Google Cloud services without exposing credentials directly in the code or to the internet. They also want to restrict access based on the principle of least privilege. Which authentication and networking approach should they implement?

    Building and operationalizing data processing systems

    • A. Configure VPC Service Controls around the Dataflow job and BigQuery.
    • B. Assign a custom service account to the Dataflow job with specific IAM roles and enable Private Google Access.
    • C. Embed service account keys directly in the Dataflow job code.
    • D. Use SSH tunneling from Dataflow workers to access BigQuery.
    Show answer

    B. Assign a custom service account to the Dataflow job with specific IAM roles and enable Private Google Access.

    Assigning a custom service account with least privilege IAM roles to the Dataflow job ensures secure, managed authentication. Enabling Private Google Access (or using Private Service Connect) allows workers on a private network to reach Google Cloud services like BigQuery without traversing the public internet, enhancing security.

  18. 18. A global e-commerce company needs to store petabytes of historical transaction data for auditing and compliance purposes. This data is accessed very infrequently, perhaps once or twice a year, but must be retained for at least 10 years. The company prioritizes minimizing storage costs while ensuring data durability and availability when needed. Which Google Cloud Storage class should be used for this data?

    Building and operationalizing data processing systems

    • A. Coldline storage
    • B. Nearline storage
    • C. Archive storage
    • D. Standard storage
    Show answer

    C. Archive storage

    Archive storage is designed for long-term data archival, offering the lowest storage costs and highest durability, ideal for data accessed very infrequently with long retention periods.

  19. 19. A data engineer is implementing a custom data processing job on Compute Engine. This job needs to read data from a Cloud Storage bucket, process it, and then write the results to a BigQuery table. The engineer wants to grant the Compute Engine instance the necessary permissions securely and follow the principle of least privilege. Which Google Cloud identity should be associated with the Compute Engine instance to achieve this?

    Building and operationalizing data processing systems

    • A. A user-managed service account with Storage Object Admin and BigQuery Data Editor roles.
    • B. A user-managed service account with Storage Object Viewer and BigQuery Data Editor roles.
    • C. A Google-managed service account with Project Editor role.
    • D. The default Compute Engine service account with Storage Object Admin and BigQuery Admin roles.
    Show answer

    B. A user-managed service account with Storage Object Viewer and BigQuery Data Editor roles.

    A user-managed service account allows for fine-grained control over permissions. The instance needs `Storage Object Viewer` to read from Cloud Storage and `BigQuery Data Editor` to write to BigQuery, adhering to the principle of least privilege.

  20. 20. A data engineering team is building a complex data pipeline that involves ingesting data from various sources, performing multiple transformation steps (filtering, aggregation, joining), and loading the results into BigQuery. The pipeline has dependencies between stages, conditional execution logic, and requires robust error handling and retry mechanisms. The team needs a fully managed service to define, schedule, and monitor these workflows, leveraging Python for custom logic. Which Google Cloud service is best suited for orchestrating this data pipeline?

    Building and operationalizing data processing systems

    • A. Cloud Run
    • B. Cloud Scheduler
    • C. Cloud Functions
    • D. Cloud Composer
    Show answer

    D. Cloud Composer

    Cloud Composer, based on Apache Airflow, is a fully managed workflow orchestration service. It is specifically designed to programmatically author, schedule, and monitor complex data pipelines with dependencies, conditional logic, and robust error handling, using Python.

  21. 21. A media streaming service is migrating its on-premises user activity database to Google Cloud. The database contains millions of user profiles and preferences, requiring low-latency reads and writes for real-time personalization features. The data schema is highly flexible and subject to frequent changes as new personalization features are introduced. Which Google Cloud database service is the most appropriate choice for this migration?

    Building and operationalizing data processing systems

    • A. Cloud Spanner
    • B. Firestore
    • C. Cloud SQL
    • D. BigQuery
    Show answer

    B. Firestore

    Firestore is a NoSQL document database designed for flexible, scalable, and low-latency access to data, particularly well-suited for mobile, web, and IoT applications requiring real-time synchronization and flexible schemas. Its document model accommodates schema changes easily, and its real-time capabilities are ideal for personalization features.

  22. 22. A large media company manages petabytes of video assets that are accessed frequently for the first 30 days after upload, then rarely (once a quarter or less) for the next 5 years, and finally almost never (once a year or less) for long-term archival. They need a cost-effective storage solution on Google Cloud that automatically transitions data between storage classes based on its age and access patterns, minimizing storage costs while ensuring availability when needed. Which Cloud Storage feature should they implement?

    Building and operationalizing data processing systems

    • A. Object Lifecycle Management
    • B. Object Versioning
    • C. Signed URLs
    • D. Customer-Supplied Encryption Keys
    Show answer

    A. Object Lifecycle Management

    Cloud Storage Object Lifecycle Management allows you to define rules to automatically transition objects between storage classes (e.g., from Standard to Nearline to Coldline to Archive) or delete them based on conditions like age, versioning, or creation date. This is ideal for optimizing storage costs based on access patterns.

  23. 23. A logistics company uses a BigQuery data warehouse for analyzing shipment data. They notice that queries filtering on `delivery_date` and `warehouse_id` columns are consistently slow and expensive, even though these columns are frequently used together in WHERE clauses. The table contains billions of rows and is partitioned by `shipment_date`. Which BigQuery feature should they implement to improve query performance and reduce costs for these specific queries?

    Building and operationalizing data processing systems

    • A. Export the filtered data to Cloud Storage and query from there.
    • B. Implement clustering on `delivery_date` and `warehouse_id`.
    • C. Create a materialized view on the `delivery_date` and `warehouse_id` columns.
    • D. Change the table partitioning to `delivery_date`.
    Show answer

    B. Implement clustering on `delivery_date` and `warehouse_id`.

    Clustering in BigQuery organizes data within each partition based on the specified columns, significantly improving performance and reducing costs for queries that filter or aggregate on those clustered columns.

  24. 24. A gaming company uses Cloud Spanner for its global leaderboard, which stores player IDs and high scores. To comply with privacy regulations, they need to ensure that player IDs are pseudonymized when used for analytics, but can be reversed to their original form for customer support purposes. The pseudonymization process must be consistent across all systems and securely managed. Which Google Cloud service should be used to manage this reversible pseudonymization?

    Managing and securing data

    • A. Cloud Data Loss Prevention (DLP) API
    • B. Secret Manager
    • C. Cloud Key Management Service (KMS)
    • D. Cloud Identity and Access Management (IAM)
    Show answer

    C. Cloud Key Management Service (KMS)

    Cloud Key Management Service (KMS) can be used to manage symmetric encryption keys. These keys can then be used by applications to encrypt (pseudonymize) and decrypt (reverse pseudonymize) sensitive data like player IDs, ensuring consistency and secure key management.

  25. 25. A global ride-sharing company collects real-time GPS coordinates from millions of active vehicles. This data needs to be ingested immediately, with very low latency, and then processed by a stream analytics engine to detect anomalies and optimize routes. The data volume can fluctuate significantly, peaking during rush hours. The company requires a messaging service that can handle high throughput, ensure message durability, and scale automatically without manual intervention. Which Google Cloud service is most suitable for ingesting this data?

    Building and operationalizing data processing systems

    • A. Cloud Storage
    • B. Cloud SQL
    • C. Cloud Pub/Sub
    • D. BigQuery
    Show answer

    C. Cloud Pub/Sub

    Cloud Pub/Sub is a fully managed, real-time messaging service designed for high-throughput, low-latency data ingestion and delivery. It automatically scales to handle fluctuating data volumes and ensures message durability, making it ideal for streaming data from millions of devices.

Professional Data Engineer flashcards

Tap a card to flip it. 170 flashcards in the full deck.

  • Google Cloud Spanner

    Flip card

    A globally distributed, horizontally scalable, and strongly consistent relational database service, designed for mission-critical applications requiring high availability and transactional consistency.

    • Global strong consistency (ACID transactions)
    • Horizontal scalability across regions and continents
    • High availability with automatic failover
    Study this card →
  • Dataflow Prime Right-Fitting

    Flip card

    Right-Fitting is a Dataflow Prime feature that automatically and intelligently selects the optimal virtual machine shapes and resource configurations for each stage of a Dataflow pipeline, balancing performance and cost.

    • Part of Dataflow Prime, an advanced Dataflow offering.
    • Analyzes pipeline workload characteristics.
    • Dynamically allocates appropriate CPU, memory, and disk.
    Study this card →
  • Cloud Pub/Sub

    Flip card

    Cloud Pub/Sub is a fully managed, real-time messaging service that enables asynchronous communication between applications, designed for highly scalable and durable event ingestion and delivery.

    • Globally distributed and highly available.
    • Automatically scales to handle millions of messages/sec.
    • Guarantees at-least-once message delivery.
    Study this card →
  • Cloud Bigtable

    Flip card

    Cloud Bigtable is a fully managed, scalable NoSQL wide-column database service for large analytical and operational workloads, offering high throughput and low latency.

    • Ideal for time-series, marketing, financial, and IoT data.
    • Supports millions of reads/writes per second.
    • Sub-10ms latency for typical operations.
    Study this card →
  • Cloud KMS for CMEK

    Flip card

    Cloud KMS is a Google Cloud service for managing cryptographic keys, including customer-managed encryption keys (CMEK), which can be used to encrypt data in other Google Cloud services.

    • Supports symmetric and asymmetric keys.
    • Integrates with many Google Cloud services for CMEK.
    • Provides automatic key rotation policies.
    Study this card →
  • Cloud Monitoring & Logging

    Flip card

    A suite of Google Cloud services that provides comprehensive observability for applications and infrastructure, including metrics, logs, and alerts.

    • Cloud Monitoring for metrics, dashboards, and alerts.
    • Cloud Logging for centralized log collection and analysis.
    • Essential for detecting and diagnosing operational issues.
    Study this card →
  • BigQuery Column-level Security

    Flip card

    BigQuery column-level security allows you to define fine-grained access control on specific columns within a table. By assigning policy tags to columns and granting users access to these tags, you can restrict who can view or query the data in those sensitive columns.

    • Restricts access to entire columns, not just their masked values.
    • Uses policy tags to categorize and control access to sensitive columns.
    • Integrated with Cloud IAM for granting access to policy tags.
    Study this card →
  • Cloud Storage for Data Lakes

    Flip card

    Cloud Storage serves as a foundational component for data lakes on Google Cloud, providing scalable, durable, and cost-effective object storage for raw and semi-structured data.

    • Schema-agnostic, supports all file types
    • Integrates with BigQuery, Dataflow, Dataproc
    • Offers different storage classes for cost optimization
    Study this card →
  • Cloud Composer (Apache Airflow)

    Flip card

    A fully managed workflow orchestration service on Google Cloud, based on Apache Airflow, enabling users to programmatically author, schedule, and monitor complex data pipelines as Directed Acyclic Graphs (DAGs).

    • Uses Python for defining workflows (DAGs).
    • Handles scheduling, dependencies, and conditional logic.
    • Provides a web UI for monitoring and management.
    Study this card →
  • Dataflow CPU Bottleneck Optimization

    Flip card

    Strategies to resolve performance bottlenecks in Dataflow pipelines caused by CPU-intensive operations, primarily focusing on optimizing the inefficient code within `DoFn`s or offloading specialized computations.

    • High CPU utilization points to inefficient code or complex calculations.
    • Scaling workers helps with parallelism, but not intrinsic code efficiency.
    • Refactoring `DoFn`s or offloading work are key solutions.
    Study this card →
  • Apache Beam Watermarks

    Flip card

    A system-generated timestamp that tracks the progress of event time in a streaming pipeline, indicating when all data up to a certain point is expected to have arrived.

    • Crucial for correct windowed aggregations in streaming.
    • Helps determine when a window can be closed.
    • Manages out-of-order and late-arriving data.
    Study this card →
  • Data Lake with Cloud Storage & Dataflow

    Flip card

    A common pattern for building a scalable data lake on Google Cloud using Cloud Storage for raw data storage and Dataflow for flexible, serverless ETL/ELT processing.

    • Cloud Storage provides cost-effective, scalable object storage for raw data.
    • Dataflow supports both batch and streaming transformations.
    • Serverless operations minimize overhead.
    Study this card →
  • BigQuery Scripting

    Flip card

    A BigQuery feature that allows writing complex, multi-statement SQL queries with control flow (e.g., `DECLARE`, `SET`, `IF`, `LOOP`) directly within the BigQuery engine.

    • Enables advanced data preparation and feature engineering
    • Leverages BigQuery's distributed execution
    • Keeps data processing within the data warehouse
    Study this card →
  • Cloud Bigtable for Real-time Analytics

    Flip card

    Google Cloud's fully managed NoSQL wide-column database, optimized for large analytical and operational workloads requiring high throughput and low-latency access.

    • High read/write throughput (millions of ops/sec)
    • Sub-10ms latency for point reads
    • Ideal for time-series, IoT, and operational analytics
    Study this card →
  • Cloud Data Catalog Data Lineage

    Flip card

    Cloud Data Catalog Data Lineage provides automated tracking and visualization of data movement and transformations across various Google Cloud services, helping users understand data origin and trust.

    • Automates lineage collection.
    • Visualizes data flow and transformations.
    • Supports BigQuery, Pub/Sub, Dataflow, and more.
    Study this card →
  • Service Account-based Auth with Private IP (Dataflow)

    Flip card

    Using a Google Cloud service account with specific IAM roles for Dataflow worker authentication, combined with Private Google Access or Private Service Connect for secure, private network communication to GCP services.

    • Managed identity for GCP services
    • Principle of least privilege (IAM roles)
    • Private network access to GCP APIs (no public internet)
    Study this card →
  • Cloud Storage Archive Class

    Flip card

    Google Cloud Storage Archive class is a highly durable, low-cost storage class for data accessed less than once a year, ideal for long-term archives and disaster recovery.

    • Lowest storage cost among GCS classes.
    • High durability (99.999999999% annual durability).
    • Higher retrieval costs and latency compared to other classes.
    Study this card →
  • Google Cloud Service Account

    Flip card

    Special Google accounts used by applications or Compute Engine instances to make authorized API calls.

    • Acts as an identity for non-human components.
    • Permissions are granted via IAM roles.
    • User-managed service accounts allow for custom, fine-grained permissions.
    Study this card →
  • Cloud Composer

    Flip card

    A fully managed workflow orchestration service built on Apache Airflow, allowing you to programmatically author, schedule, and monitor complex data pipelines.

    • Uses Python for defining workflows (DAGs)
    • Manages dependencies between tasks
    • Provides a rich UI for monitoring and management
    Study this card →
  • Firestore

    Flip card

    A flexible, scalable NoSQL document database for mobile, web, and server development, offering real-time data synchronization.

    • NoSQL document-oriented
    • Real-time data synchronization
    • Flexible schema and highly scalable
    Study this card →
  • Cloud Storage Object Lifecycle Management

    Flip card

    A Cloud Storage feature that allows you to define rules to automatically perform actions on objects, such as changing their storage class or deleting them, based on conditions.

    • Automates cost optimization by moving data to cheaper classes
    • Rules can be based on age, creation date, number of versions, etc.
    • Applies at the bucket level
    Study this card →
  • BigQuery Clustering

    Flip card

    A BigQuery feature that organizes data within each partition based on the values of one or more specified columns, improving query performance and reducing cost for filtering and aggregation.

    • Data is physically co-located within partitions.
    • Optimizes queries with `WHERE` clauses and aggregations on clustered columns.
    • Can be combined with partitioning for multi-level optimization.
    Study this card →
  • Reversible Pseudonymization with KMS

    Flip card

    Reversible pseudonymization using Cloud KMS involves encrypting sensitive identifiers with a symmetric encryption key managed by KMS. This allows the data to be de-identified for analytics while retaining the ability to decrypt it back to its original form for specific use cases.

    • Uses symmetric encryption keys from KMS.
    • Applications perform encryption/decryption using the key.
    • Ensures consistent pseudonymization across systems.
    Study this card →
  • Cloud Spanner with Cloud KMS External Key Manager (EKM)

    Flip card

    Cloud Spanner can integrate with Cloud KMS to use Customer-Managed Encryption Keys (CMEK). When combined with Cloud KMS External Key Manager (EKM), this allows Cloud Spanner to encrypt data at rest using keys that are physically managed in the customer's own external key management systems, such as on-premises FIPS 140-2 Level 3 certified HSMs.

    • Enables Cloud Spanner to use customer-managed keys from external HSMs.
    • Keys are stored and controlled outside Google Cloud infrastructure.
    • Meets strict compliance for key sovereignty and physical isolation.
    Study this card →

Questions are original practice items written to match the published exam objectives. Step2Study is not affiliated with or endorsed by any certification body.