Professional Data Engineer flashcards
170 free flashcards. Tap a card to flip it.
Cloud Storage Archive Class
Flip cardThe lowest-cost Cloud Storage class designed for long-term data archiving with access frequencies of less than once per year.
- Lowest storage cost
- Highest retrieval cost and access latency
- Ideal for regulatory compliance, disaster recovery, and deep archives
Memory trick: Standard is quick, Nearline is less, Coldline's cool, Archive's best for rest.
Reversible Pseudonymization with Cloud KMS (Deterministic Encryption)
Flip cardReversible pseudonymization using Cloud KMS involves encrypting sensitive identifiers (like PII) with a deterministic encryption algorithm. This means the same input always produces the same encrypted output, allowing for consistent lookups or joins on the pseudonymized data while enabling authorized decryption back to the original identifier when needed.
- Uses deterministic encryption for consistent pseudonymized values.
- Enables joins and analytics on pseudonymized data.
- Encryption keys are securely managed by Cloud KMS.
- Allows authorized reversal to the original data for specific purposes (e.g., support).
Memory trick: KMS gives a consistent mask, when you need to unmask, it's there.
Cloud Spanner with Cloud KMS External Key Manager (EKM)
Flip cardCloud Spanner can integrate with Cloud KMS to use Customer-Managed Encryption Keys (CMEK). When combined with Cloud KMS External Key Manager (EKM), this allows Cloud Spanner to encrypt data at rest using keys that are physically managed in the customer's own external key management systems, such as on-premises FIPS 140-2 Level 3 certified HSMs.
- Enables Cloud Spanner to use customer-managed keys from external HSMs.
- Keys are stored and controlled outside Google Cloud infrastructure.
- Meets strict compliance for key sovereignty and physical isolation.
- Cloud Spanner data is encrypted/decrypted using these external keys.
Memory trick: Spanner stretches globally, but the key stays home.
Cloud Pub/Sub
Flip cardA global, fully managed, real-time messaging service that allows you to send and receive messages between independent applications.
- Asynchronous messaging
- Decouples senders and receivers
- Scales automatically to handle high throughput
Memory trick: Pub/Sub is the express lane for your streaming data.
Reversible Pseudonymization with KMS
Flip cardReversible pseudonymization using Cloud KMS involves encrypting sensitive identifiers with a symmetric encryption key managed by KMS. This allows the data to be de-identified for analytics while retaining the ability to decrypt it back to its original form for specific use cases.
- Uses symmetric encryption keys from KMS.
- Applications perform encryption/decryption using the key.
- Ensures consistent pseudonymization across systems.
Memory trick: Keys Can Securely Convert, Keep Reversible.
BigQuery Clustering
Flip cardA BigQuery feature that organizes data within each partition based on the values of one or more specified columns, improving query performance and reducing cost for filtering and aggregation.
- Data is physically co-located within partitions.
- Optimizes queries with `WHERE` clauses and aggregations on clustered columns.
- Can be combined with partitioning for multi-level optimization.
Memory trick: Cluster your data for faster access.
Cloud Storage Object Lifecycle Management
Flip cardA Cloud Storage feature that allows you to define rules to automatically perform actions on objects, such as changing their storage class or deleting them, based on conditions.
- Automates cost optimization by moving data to cheaper classes
- Rules can be based on age, creation date, number of versions, etc.
- Applies at the bucket level
Memory trick: Lifecycle management helps your objects gracefully age and save money.
Firestore
Flip cardA flexible, scalable NoSQL document database for mobile, web, and server development, offering real-time data synchronization.
- NoSQL document-oriented
- Real-time data synchronization
- Flexible schema and highly scalable
Memory trick: Database choice: Schema flexibility and access speed guide the way.
Dataflow Exactly-Once
Flip cardA processing guarantee in Dataflow (and Beam) that ensures each data element is processed and affects the output exactly one time, even with failures or retries.
- Crucial for correctness in systems like fraud detection
- Achieved through internal mechanisms like unique IDs and state management
- Prevents duplicate outputs and side effects
Memory trick: Exactly once means no duplicates, ever.
Cloud Composer
Flip cardA fully managed workflow orchestration service built on Apache Airflow, allowing you to programmatically author, schedule, and monitor complex data pipelines.
- Uses Python for defining workflows (DAGs)
- Manages dependencies between tasks
- Provides a rich UI for monitoring and management
Memory trick: Composer conducts your data symphony.
Google Cloud Service Account
Flip cardSpecial Google accounts used by applications or Compute Engine instances to make authorized API calls.
- Acts as an identity for non-human components.
- Permissions are granted via IAM roles.
- User-managed service accounts allow for custom, fine-grained permissions.
Memory trick: Service accounts are the robots with specific job badges.
Service Account-based Auth with Private IP (Dataflow)
Flip cardUsing a Google Cloud service account with specific IAM roles for Dataflow worker authentication, combined with Private Google Access or Private Service Connect for secure, private network communication to GCP services.
- Managed identity for GCP services
- Principle of least privilege (IAM roles)
- Private network access to GCP APIs (no public internet)
Memory trick: Service accounts on private roads, keep data safe.
Cloud Data Catalog Data Lineage
Flip cardCloud Data Catalog Data Lineage provides automated tracking and visualization of data movement and transformations across various Google Cloud services, helping users understand data origin and trust.
- Automates lineage collection.
- Visualizes data flow and transformations.
- Supports BigQuery, Pub/Sub, Dataflow, and more.
Memory trick: Lineage Links Latest Location.
Cloud Bigtable for Real-time Analytics
Flip cardGoogle Cloud's fully managed NoSQL wide-column database, optimized for large analytical and operational workloads requiring high throughput and low-latency access.
- High read/write throughput (millions of ops/sec)
- Sub-10ms latency for point reads
- Ideal for time-series, IoT, and operational analytics
Memory trick: Bigtable handles big data fast, with wide columns for real-time blast.
BigQuery Scripting
Flip cardA BigQuery feature that allows writing complex, multi-statement SQL queries with control flow (e.g., `DECLARE`, `SET`, `IF`, `LOOP`) directly within the BigQuery engine.
- Enables advanced data preparation and feature engineering
- Leverages BigQuery's distributed execution
- Keeps data processing within the data warehouse
Memory trick: Scripting orchestrates your BigQuery data dance.
Cloud Composer (Apache Airflow)
Flip cardA fully managed workflow orchestration service built on Apache Airflow, enabling programmatic authoring, scheduling, and monitoring of complex data pipelines.
- DAG (Directed Acyclic Graph) for workflow definition
- Managed Airflow environment
- Supports various operators for different services
Memory trick: Composer conducts the orchestra of data tasks.
Dataflow Worker Sizing and Autoscaling
Flip cardDataflow automatically scales the number of worker VMs based on workload. `maxNumWorkers` sets the upper limit for this scaling, influencing the pipeline's ability to handle peak loads and maintain performance.
- Autoscaling adjusts worker count dynamically
- `maxNumWorkers` defines the upper limit for scaling
- Proper sizing is crucial for performance and cost
Memory trick: Dataflow: Scale up for speed.
Data Lake with Cloud Storage & Dataflow
Flip cardA common pattern for building a scalable data lake on Google Cloud using Cloud Storage for raw data storage and Dataflow for flexible, serverless ETL/ELT processing.
- Cloud Storage provides cost-effective, scalable object storage for raw data.
- Dataflow supports both batch and streaming transformations.
- Serverless operations minimize overhead.
Memory trick: Storage for the lake, Dataflow for the transformation stream.
Apache Beam Watermarks
Flip cardA system-generated timestamp that tracks the progress of event time in a streaming pipeline, indicating when all data up to a certain point is expected to have arrived.
- Crucial for correct windowed aggregations in streaming.
- Helps determine when a window can be closed.
- Manages out-of-order and late-arriving data.
Memory trick: The watermark marks the flow of time.
Dataflow CPU Bottleneck Optimization
Flip cardStrategies to resolve performance bottlenecks in Dataflow pipelines caused by CPU-intensive operations, primarily focusing on optimizing the inefficient code within `DoFn`s or offloading specialized computations.
- High CPU utilization points to inefficient code or complex calculations.
- Scaling workers helps with parallelism, but not intrinsic code efficiency.
- Refactoring `DoFn`s or offloading work are key solutions.
Memory trick: If the DoFn is slow, refactor and offload for a better flow.
Cloud Storage for Data Lakes
Flip cardUtilizing Google Cloud Storage as the foundational layer for a data lake due to its scalability, durability, cost-effectiveness, and ability to store any data format.
- Object storage, schema-agnostic
- Unlimited scalability
- Integration with various GCP analytics services (Dataflow, BigQuery, Dataproc)
Memory trick: Cloud Storage holds all the lake's treasures.
Stream Processing with Autoscaling
Flip cardA data processing pattern that continuously processes data as it arrives, leveraging autoscaling to dynamically adjust resources to handle fluctuating data volumes and maintain low latency.
- Processes data in real-time or near real-time
- Automatically scales resources based on workload
- Critical for applications requiring immediate insights
Memory trick: Stream autoscaling: A flexible river for real-time flow.
Dataflow Worker Sizing
Flip cardDataflow workers are Compute Engine VMs that execute pipeline steps. Their `Machine Type` defines CPU and memory resources, which are critical for performance and preventing resource exhaustion.
- Machine Type determines CPU and RAM per worker.
- Insufficient RAM leads to 'Out of memory' errors.
- Choosing the right machine type is crucial for stability and cost.
Memory trick: Dataflow's power: Match the machine's mind to the data's might.
Service Account Auth with Private Google Access
Flip cardThis setup enables secure and private communication between GCP resources (like Dataflow workers in a private VPC) and Google APIs (like BigQuery), without exposing traffic to the public internet.
- Service accounts provide identity for applications to authenticate to GCP services.
- Private Google Access allows instances in a private subnet to reach Google APIs.
- Ensures data privacy and security.
- Eliminates the need for public IPs or VPNs for GCP-to-GCP communication.
Memory trick: Service accounts unlock the private access path.
Cloud Spanner
Flip cardA globally distributed, strongly consistent, relational database service designed for mission-critical applications requiring high availability, horizontal scalability, and ACID transactions.
- Unique combination of relational structure and global scalability.
- Offers strong consistency across regions.
- Ideal for financial services, gaming, and healthcare applications.
Memory trick: Spanner spans the globe with strong ACID relations.
Cloud Pub/Sub Automatic Scaling for Ingestion
Flip cardCloud Pub/Sub's ability to automatically and dynamically adjust its capacity to handle varying message volumes, ensuring consistent performance during peak ingestion loads without manual intervention.
- Handles millions of messages/second
- Low latency message delivery
- No manual provisioning or scaling required
Memory trick: Pub/Sub's flow catches all the real-time buzz.
Dataflow Stateful Processing
Flip cardA Dataflow (Apache Beam) capability that allows PTransforms to maintain and update per-key state across elements and over time, enabling complex streaming aggregations and pattern detection.
- Essential for continuous aggregations (e.g., running counts, averages).
- State is associated with individual keys.
- Supports timers for time-based state management.
Memory trick: Stateful processing remembers the past of each data item.
Dataflow Prime Right-Fitting
Flip cardA Dataflow Prime feature that intelligently analyzes pipeline characteristics and data patterns to automatically select the optimal worker type and resource configuration (CPU, memory, storage) for each stage of the pipeline, optimizing for both performance and cost.
- Automated resource optimization
- Reduces manual tuning efforts
- Aims for optimal cost-performance balance
Memory trick: Prime right-fits the puzzle for perfect costs.
Google Cloud Composer
Flip cardGoogle Cloud Composer is a fully managed workflow orchestration service built on Apache Airflow, enabling users to author, schedule, and monitor complex data pipelines as Directed Acyclic Graphs (DAGs).
- Managed Apache Airflow service
- Orchestrates complex multi-step workflows
- Supports Python for defining DAGs
- Integrated with other Google Cloud services
Memory trick: Composer: Conducts the data symphony.
Dataflow Late Data Handling
Flip cardMechanisms in Dataflow (Apache Beam) to manage data elements that arrive after the system's watermark has passed, ensuring they are still processed correctly within their intended time windows.
- Allowed lateness specifies how long to wait for late data.
- Triggers define when to emit results, including updated results for late data.
- Crucial for accurate event-time aggregations in real-world streaming scenarios.
Memory trick: Don't be late to the party, but if you are, we'll still let you in for a bit.
BigQuery Partitioning and Clustering
Flip cardBigQuery partitioning segments tables by a date, timestamp, or integer column, reducing data scanned for filtered queries. Clustering further sorts data within partitions by specified columns, improving performance for filters and aggregations.
- Partitioning reduces data scanned by segmenting data.
- Clustering sorts data within partitions, optimizing scan for filters/aggregations.
- Combining both offers significant performance and cost benefits for large tables.
Memory trick: Partition by time, cluster by common queries.
Pub/Sub Ingestion Monitoring
Flip cardMonitoring key Pub/Sub subscription metrics to detect if messages are being consumed effectively by downstream services like Dataflow, preventing data loss.
- `oldest_unacked_message_age` indicates message backlog
- `num_unacked_messages` shows pending messages
- High values suggest consumer lag or issues
Memory trick: Follow the data flow; check ingestion first for missing pieces.
Dataflow Exactly-Once Processing
Flip cardA Dataflow streaming pipeline guarantee that ensures each unique input record affects the output exactly one time, even in the presence of system failures or retries, typically achieved using unique event IDs and stateful processing.
- Critical for financial transactions, IoT, and accurate aggregations.
- Requires unique event identifiers.
- Leverages Dataflow's fault tolerance and state management.
Memory trick: Unique IDs and Dataflow's state ensure it's done exactly once.
Cloud Monitoring & Logging
Flip cardA suite of Google Cloud services that provide comprehensive observability for applications and infrastructure, including metrics, logs, and alerts.
- Cloud Monitoring for metrics, dashboards, and alerts
- Cloud Logging for centralized log collection and analysis
- Essential for troubleshooting and performance optimization
Memory trick: Monitor and Log to keep your Dataflow flowing.
Dataflow Autoscaling
Flip cardA Dataflow feature that automatically adjusts the number of worker instances in a pipeline based on the current workload, optimizing for throughput, latency, and cost.
- Dynamically adds/removes workers
- Handles fluctuating data volumes
- Optimizes resource utilization and cost
Memory trick: Autoscaling expands and contracts like a data accordion.
Cloud Audit Logs for Data Governance
Flip cardCloud Audit Logs record administrative activities and data access events across Google Cloud services. They are crucial for security, auditing, and compliance, enabling real-time monitoring and alerting on sensitive data access.
- Captures 'Admin Activity', 'Data Access', and 'System Event' logs.
- Data Access logs are disabled by default for BigQuery and need explicit enabling.
- Logs can be exported to Cloud Storage, BigQuery, or Pub/Sub via sinks.
- Essential for demonstrating compliance and detecting security incidents.
Memory trick: Audit Logs flow to Pub/Sub, then a Function calls the SOC club!
Cloud Storage Customer-Supplied Encryption Keys (CSEK)
Flip cardA Google Cloud Storage encryption option where the customer generates and manages their own encryption keys and provides them to Google Cloud Storage for data encryption and decryption.
- Customer retains full control over key management.
- Keys are provided with each request to Cloud Storage.
- Offers the highest level of control over encryption keys.
Memory trick: Google keys easy, KMS keys managed, but Supplied keys are fully customer-handed.
Google Cloud Spanner
Flip cardA globally distributed, horizontally scalable, and strongly consistent relational database service, designed for mission-critical applications requiring high availability and transactional consistency.
- Global strong consistency (ACID transactions)
- Horizontal scalability across regions and continents
- High availability with automatic failover
Memory trick: Spanner: Spans the globe with strong transactions.
Cloud Bigtable
Flip cardCloud Bigtable is a fully managed, scalable NoSQL wide-column database service for large analytical and operational workloads, offering high throughput and low latency.
- Ideal for time-series, marketing, financial, and IoT data.
- Supports millions of reads/writes per second.
- Sub-10ms latency for typical operations.
- Horizontally scalable, integrates with Hadoop, Dataflow, etc.
Memory trick: Big gaming data, Bigtable's big speed.
Cloud KMS for CMEK
Flip cardCloud KMS is a Google Cloud service for managing cryptographic keys, including customer-managed encryption keys (CMEK), which can be used to encrypt data in other Google Cloud services.
- Supports symmetric and asymmetric keys.
- Integrates with many Google Cloud services for CMEK.
- Provides automatic key rotation policies.
Memory trick: Keys Keep Many Secrets Securely.
BigQuery Column-level Security
Flip cardBigQuery column-level security allows you to define fine-grained access control on specific columns within a table. By assigning policy tags to columns and granting users access to these tags, you can restrict who can view or query the data in those sensitive columns.
- Restricts access to entire columns, not just their masked values.
- Uses policy tags to categorize and control access to sensitive columns.
- Integrated with Cloud IAM for granting access to policy tags.
- Allows different teams to access different subsets of columns in the same table.
Memory trick: Column Security: Tag the column, lock the view.
Organization Policy Service (Resource Location)
Flip cardA Google Cloud service that allows administrators to enforce programmatic restrictions on how resources are configured and deployed across an organization, including restricting where resources can be created (resource location restriction).
- Enforces policies at the organization, folder, or project level.
- Uses constraints to define allowed or disallowed behaviors.
- The 'Resource location restriction' constraint specifically controls data residency.
- Provides centralized, automated governance.
Memory trick: Organization Policy sets the borders, Resource Location constraint draws the line, so data stays where it's meant to shine.
BigQuery Data Masking
Flip cardBigQuery data masking allows you to obscure sensitive data in a column, making it unreadable to unauthorized users while still allowing authorized users to see the original data. This is applied at query time based on policies.
- Part of BigQuery column-level security.
- Applies dynamic masking based on roles/permissions.
- No data duplication required.
- Supports various masking rules (e.g., default, email, number).
Memory trick: Masked data, quick and neat, BigQuery column-level can't be beat!
Cloud IAM and Cloud Audit Logs for Data Governance
Flip cardCloud IAM defines who has access to which Google Cloud resources and what actions they can perform. Cloud Audit Logs, particularly Data Access logs, record data access attempts, providing a comprehensive audit trail of user activities on data, which together form a robust solution for data governance, security, and compliance.
- IAM specifies permissions (e.g., BigQuery Data Viewer, BigQuery Data Editor).
- Cloud Audit Logs record administrative, system, and data access events.
- Data Access logs capture detailed user activity on data, including queries.
- Essential for meeting compliance standards like GDPR and HIPAA.
Memory trick: IAM sets the rules, Audit Logs watch the game.
Cloud KMS HSM Keys
Flip cardCloud Key Management Service (KMS) provides various key types, including Hardware Security Module (HSM) keys, which are backed by FIPS 140-2 Level 3 validated HSMs for enhanced security and compliance.
- Manages cryptographic keys for various Google Cloud services.
- HSM keys offer strong, hardware-backed protection for encryption keys.
- Meets strict regulatory and compliance requirements.
- Integrates with other services like Cloud Storage for CMEK.
Memory trick: For keys so secure, KMS with HSM is the cure!
Cloud Audit Logs for BigQuery Data Access
Flip cardCloud Audit Logs record administrative activities, system events, and data access on Google Cloud. For BigQuery, enabling Data Access logs captures detailed information about queries, including who ran them, the query text, and the resources (tables/columns) accessed, crucial for auditing and compliance.
- Records API calls that read or modify data.
- Captures user identity, operation, and affected resources.
- Essential for security, auditing, and compliance.
- Data Access logs are usually disabled by default for BigQuery and need explicit activation.
Memory trick: To know *who* touched *what* data, *audit* the *access*.
Organization Policy Constraints (Resource Location)
Flip cardOrganization Policy Constraints are rules that apply across a Google Cloud organization, folders, or projects to enforce specific behaviors, such as restricting where resources can be created (Resource Location Restriction).
- Enforced at the organization, folder, or project level.
- Prevents creation of resources in non-compliant locations.
- Crucial for data residency and compliance (e.g., GDPR, HIPAA).
- Overrides individual user permissions for resource creation.
Memory trick: Organization Policy's Constraint, keeps your data's location without taint!
Dataflow Watermark Lag
Flip cardA state in Dataflow streaming pipelines where the system's understanding of event time (watermark) falls significantly behind the actual current event time, leading to increased 'System Latency' and 'Data freshness' issues.
- Indicates a backlog of unprocessed event-time data
- Causes stale results in dashboards
- Often due to bottlenecks in processing, not necessarily resource exhaustion
Memory trick: Slow watermark means data's stuck in time's bottleneck.
Cloud Logging and Cloud Monitoring
Flip cardCore Google Cloud services for observing and managing operations, providing centralized log collection and analysis (Logging) and comprehensive metrics, dashboards, and alerting (Monitoring).
- Cloud Logging aggregates logs from all GCP services.
- Cloud Monitoring collects metrics, provides dashboards, and configures alerts.
- Essential for troubleshooting, performance analysis, and operational health.
Memory trick: Log everything, then Monitor the metrics.
Watermarks and Allowed Lateness
Flip cardIn Apache Beam, watermarks are a measure of processing progress based on event time, and allowed lateness defines a grace period after a window's watermark passes during which late data can still be processed for that window.
- Watermarks indicate event-time completeness
- Allowed lateness handles data arriving after its window's end
- Prevents indefinite delays while ensuring data accuracy
Memory trick: Watermarks and Lateness: The clock and the grace period.
Cloud Interconnect (Dedicated)
Flip cardCloud Interconnect (Dedicated) provides a direct physical connection between an on-premises data center and Google's network, bypassing the public internet.
- Private, dedicated connection.
- High bandwidth and low latency.
- Enhanced security and compliance.
- Requires physical presence at a Google colocation facility.
Memory trick: Private paths for precious packets, bypassing the public's prying eyes.
BigQuery Data Lifecycle Management
Flip cardThe process of managing BigQuery data from creation to archival or deletion, often using features like table expiration and data retention policies to optimize storage costs and meet compliance.
- Table expiration automatically deletes tables or partitions.
- Data retention policies define minimum data age before deletion or transition.
- Can be set at dataset or table level.
Memory trick: Expire the temporary, retain the old, manage with policies, and save on the gold.
Pub/Sub Message Retention & Dead-Letter Topics
Flip cardPub/Sub subscriptions allow configuring message retention duration to automatically purge messages after a specified time. Dead-letter topics provide a mechanism to catch messages that cannot be successfully processed, preventing loss and enabling debugging or reprocessing.
- Message retention is configured at the subscription level (up to 7 days, or 31 days with topic retention).
- Dead-letter topics are configured on subscriptions.
- Messages sent to a dead-letter topic are not lost.
- Useful for compliance and robust message processing.
Memory trick: Retain the messages, but don't let them die, use a dead-letter, for failures high!
Cloud Spanner with CMEK
Flip cardCloud Spanner is a globally distributed, strongly consistent, and highly available relational database service that supports customer-managed encryption keys (CMEK) for data at rest.
- Enterprise-grade, highly scalable relational database.
- Strong global consistency.
- Supports CMEK for enhanced security and compliance.
Memory trick: Spanner Secures Critical Customer Keys Globally.
BigQuery Partitioning for Cost
Flip cardUsing BigQuery table partitioning, especially by date, to leverage automatic cost reductions for older, less-accessed data that qualifies for long-term storage pricing.
- Divides table into smaller segments
- Data in partitions older than 90 days gets cheaper pricing
- Improves query performance by pruning partitions
Memory trick: Partition your BigQuery tables to pay less for old logs.
Cloud IAM & Cloud Audit Logs for Data Governance
Flip cardCloud IAM provides granular access control to Google Cloud resources, while Cloud Audit Logs record administrative activities and data access events, forming a crucial foundation for data governance and compliance.
- Cloud IAM defines 'who' can do 'what' on 'which' resources.
- Cloud Audit Logs capture Admin Activity, Data Access, and System Event logs.
- Data Access logs record read/write operations on user-provided data.
Memory trick: IAM grants the keys, Audit Logs record every turn, ensuring governance and trust are earned.
Dataflow Allowed Lateness
Flip cardA Dataflow windowing option that specifies how long the system should wait for late-arriving data after a window's watermark has passed its end, allowing late events to be processed.
- Configured per window.
- Ensures critical late data is not discarded.
- Can lead to multiple emissions for a window (early, on-time, late).
Memory trick: Don't be late for the allowed lateness party!
Dataflow Fixed Windows with Allowed Lateness
Flip cardA Dataflow strategy where data is grouped into non-overlapping, fixed-duration windows, and a specified 'allowed lateness' period is added to the window's end to process late-arriving elements.
- Fixed windows group data into consistent time intervals.
- Allowed lateness defines a grace period after the watermark passes the window end.
- Events arriving within the allowed lateness are processed as 'late' elements.
- Events arriving after the allowed lateness are typically discarded.
Memory trick: Fixed window, plus a little grace, late data finds its place.
Cloud SQL with Customer-Managed Encryption Keys (CMEK)
Flip cardCloud SQL instances can be configured to use Customer-Managed Encryption Keys (CMEK) from Cloud Key Management Service (KMS). This allows customers to control the encryption keys used for their data at rest in Cloud SQL, providing enhanced security and meeting specific compliance requirements.
- Encrypts data at rest in Cloud SQL using keys from Cloud KMS.
- Customer controls key lifecycle, permissions, and rotation.
- Integrated directly with Cloud SQL for seamless encryption/decryption.
- Essential for compliance with regulations requiring customer key management.
Memory trick: SQL's data is safe with KMS's custom key.
Cloud Data Fusion
Flip cardA fully managed, cloud-native data integration service built on open-source CDAP, enabling ETL/ELT pipelines with a visual interface.
- Visual, code-free data pipeline creation
- Supports various data sources and sinks
- Ideal for batch ETL/ELT and data quality tasks
Memory trick: Visual flows make data fuse seamlessly.