Professional Data Engineer flashcards
170 free flashcards. Tap a card to flip it.
Cloud Bigtable for Time-Series Data
Flip cardCloud Bigtable is a petabyte-scale, low-latency, wide-column NoSQL database optimized for large analytical and operational workloads, especially well-suited for time-series data, IoT, and high-throughput applications.
- Handles millions of writes/second.
- Consistent low-latency performance.
- Ideal for time-series, IoT, and operational analytics.
- Scales to petabytes of data.
Memory trick: Bigtable: The 'Big Clock' for your time-series data.
Dataflow for Large-scale Batch Processing
Flip cardA fully managed Google Cloud service for executing Apache Beam pipelines, enabling highly scalable, fault-tolerant, and cost-effective batch data processing for petabyte-scale datasets.
- Serverless and auto-scaling.
- Supports complex transformations (Apache Beam).
- Built-in fault tolerance and exactly-once processing (for streaming).
- Ideal for large-scale ETL, ELT, and data-intensive computations.
Memory trick: For a massive data factory, you need a 'Dataflow' river that can carry huge loads, process them, and handle any bumps along the way.
Secure Data Lake on Cloud Storage
Flip cardBuilding a compliant data lake on Google Cloud Storage involves using its object storage capabilities combined with robust security features like CMEK for encryption, VPC Service Controls for perimeter security, and Cloud Audit Logs for transparency and accountability.
- Cloud Storage is scalable for diverse data lake data.
- CMEK provides customer control over encryption keys.
- VPC Service Controls protect against data exfiltration.
- Cloud Audit Logs record administrative activity and data access.
Memory trick: Secure Lake, Controlled Access, Audited Trail.
Cloud Pub/Sub for Real-time Ingestion
Flip cardCloud Pub/Sub is a fully managed, scalable, and asynchronous messaging service that decouples senders and receivers, ensuring reliable message delivery for real-time data streams.
- At-least-once delivery guarantee.
- Scales automatically to handle high message volumes.
- Messages are persisted until acknowledged by subscribers.
Memory trick: Pub/Sub is the 'Post Office' for streaming data, always delivering.
Real-time Fraud Detection Architecture
Flip cardA system designed to identify and flag fraudulent activities as they occur, typically involving high-throughput ingestion, low-latency stream processing, and robust data storage for analysis.
- Requires services capable of handling millions of events/second.
- Processing must occur with sub-second latency.
- Data needs to be durable and available for historical analysis.
- Often combines a message queue, stream processor, and data warehouse.
Memory trick: Fraudulent transactions are like a fast-moving river; you need a strong boat to catch them, a net to filter, and a big lake to store evidence.
Cloud Bigtable for Operational Time-Series
Flip cardA highly scalable, low-latency NoSQL wide-column database on Google Cloud, optimized for large operational and analytical workloads, especially suitable for time-series data, IoT, and financial data.
- Petabyte-scale, fully managed NoSQL database.
- Ultra-low latency reads and writes (milliseconds).
- Ideal for high-throughput time-series data (IoT, sensor, GPS).
- Supports simple key-value lookups and wide-column structures.
Memory trick: For a 'Big Table' of fast-moving car data, you need a database that can write and read at lightning speed.
Cloud Pub/Sub for Scalable Real-time Ingestion
Flip cardCloud Pub/Sub is a fully managed, global messaging service that enables asynchronous communication between applications, designed for highly scalable and reliable real-time data ingestion.
- Automatic scaling for unpredictable traffic.
- Decouples senders and receivers.
- At-least-once delivery guarantee.
- Low operational overhead.
Memory trick: Pub/Sub is the 'Post Office' for your growing log messages.
Hybrid Real-time & Batch Data Pipeline
Flip cardA robust data pipeline combining streaming ingestion and processing for real-time insights with batch processing for historical analysis and machine learning, leveraging services like Pub/Sub, Dataflow, Bigtable, and BigQuery.
- Pub/Sub for high-volume, real-time ingestion.
- Dataflow for flexible, scalable streaming and batch processing.
- Cloud Bigtable for low-latency operational data (aggregations).
- BigQuery for petabyte-scale historical data and ML training.
Memory trick: Pub/Sub streams, Dataflow transforms, Bigtable shows NOW, BigQuery learns LATER.
BigQuery for Data Warehousing
Flip cardA serverless, highly scalable, and cost-effective enterprise data warehouse designed for petabyte-scale analytics, offering robust security and compliance features.
- Fully managed and serverless.
- Scales automatically to petabytes.
- Optimized for analytical queries (OLAP).
- Strong security features (encryption, IAM, data residency controls).
Memory trick: For a huge data vault, you need a Big Query to find treasures, with strong locks and guards.
Pub/Sub and Dataflow for Streaming
Flip cardPub/Sub provides a global, scalable, and durable message queue for ingesting real-time event streams, while Dataflow offers a serverless platform for executing stream processing pipelines with advanced features like exactly-once processing and auto-scaling.
- Pub/Sub handles high-throughput message ingestion.
- Dataflow supports both batch and streaming processing.
- Dataflow streaming pipelines offer exactly-once processing guarantees.
Memory trick: Publish and Flow for Fast Data.
Hybrid Time-Series Storage
Flip cardA data architecture that combines different storage solutions to optimize for both low-latency access to recent data and cost-effective, long-term storage of historical data, often involving a fast operational database and an analytical data warehouse.
- Balances performance (low-latency) with cost-effectiveness.
- Typically uses a NoSQL database for recent, 'hot' data.
- Leverages a data warehouse or object storage for 'cold' historical data.
- Common pattern for IoT, monitoring, and financial time-series data.
Memory trick: Time-series data is like a river: the fresh water needs quick access, but the old water can sit in a big, cheap reservoir.
Firestore for Real-time Feature Serving
Flip cardA flexible, scalable NoSQL document database on Google Cloud, providing low-latency access to semi-structured data, suitable for real-time application features like recommendation engines and user profiles.
- NoSQL document model, flexible schema.
- Scales globally and is fully managed.
- Sub-10ms latency for point reads/writes.
- Supports real-time synchronization and offline capabilities.
Memory trick: To feed a 'Fire' fast recommendation engine with user 'Stories', you need a database that's quick, flexible, and always available.
Streaming Analytics Pipeline
Flip cardAn architecture designed to ingest, process, and analyze continuous streams of data in near real-time, often combining a message queue, a stream processing engine, and an analytical data store.
- Handles high-volume, continuous data streams.
- Provides low-latency processing and aggregations.
- Supports both real-time dashboards and historical analysis.
- Commonly used for clickstream, IoT, and financial data.
Memory trick: To analyze every click in real-time and keep it forever, you need a data river that flows fast into a huge, smart lake.
Bigtable for Low-Latency NoSQL
Flip cardCloud Bigtable is a fully managed, scalable NoSQL wide-column database service designed for large analytical and operational workloads, offering very low latency for high-throughput reads and writes, making it ideal for real-time applications.
- High throughput and low latency for petabyte-scale data.
- Ideal for time-series data, marketing data, financial data, and IoT data.
- Supports open-source HBase API.
- Fully managed with automatic scaling and replication.
Memory trick: BigTable + Dataflow = Fast Recommendations.
Dataflow for Batch Pipelines
Flip cardGoogle Cloud Dataflow is a fully managed, serverless service that executes Apache Beam pipelines, providing a powerful and flexible platform for building robust, fault-tolerant batch data pipelines with advanced features like schema evolution handling and auto-scaling.
- Unified programming model (Apache Beam) for batch and streaming.
- Serverless and auto-scaling.
- Fault-tolerant with exactly-once processing for certain operations.
- Excellent for complex ETL/ELT transformations and aggregations.
Memory trick: Data Flow, Batch by Batch, Clean and Go.
Secure Data Handling with CMEK, IAM, and Audit Logs
Flip cardThis combination provides robust security for sensitive data on Google Cloud: encryption at rest with customer key control (CMEK), granular access management (IAM), and comprehensive logging of data access and administrative actions (Cloud Audit Logs).
- CMEK provides customer control over encryption keys for data at rest.
- IAM enables fine-grained role-based access control.
- Cloud Audit Logs record all access and admin activities for compliance and forensics.
Memory trick: Keys, Roles, and Logs: The KRL of data security.
Dataflow (Apache Beam)
Flip cardGoogle Cloud Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, enabling unified programming for both batch and streaming data processing with auto-scaling and high performance.
- Unified programming model for batch and streaming.
- Serverless and fully managed service.
- Auto-scaling for dynamic workloads.
- Supports complex data transformations and aggregations.
Memory trick: Data Flow, One Code, Both Streams.
Serverless Batch Processing
Flip cardExecuting large-scale batch data transformations and aggregations using a fully managed, auto-scaling, and pay-per-use serverless service.
- Minimizes operational overhead.
- Scales automatically based on workload.
- Cost-effective with pay-as-you-go pricing.
Memory trick: Dataflow: Data flows, serverless, saving cash.
Serverless Container Workloads
Flip cardRunning containerized applications in a fully managed, automatically scaling, and pay-per-use serverless environment.
- Ideal for spiky workloads and microservices.
- Scales to zero, meaning no cost when idle.
- Supports any language or library that can be containerized.
Memory trick: Cloud Run: Containers that Scale and Save Cash.
Time-Series Hot/Cold Path
Flip cardAn architecture pattern for time-series data where recent, frequently accessed data resides in a low-latency store (hot path) and older, less frequently accessed data is moved to a cost-effective archival store (cold path).
- Hot path for immediate operational queries.
- Cold path for long-term historical analysis.
- Optimizes for both performance and cost.
Memory trick: Pub/Sub for Push, Bigtable for Hot, BigQuery for Cold Storage.
Real-time Streaming Pipeline
Flip cardA data pipeline designed to ingest, process, and analyze data in real-time, often for immediate insights or operational decisions.
- Uses Pub/Sub for scalable ingestion.
- Leverages Dataflow for stream processing.
- Stores processed data in BigQuery for analytics.
Memory trick: Pub/Sub Pushes, Dataflow Digests, BigQuery's Best for Data Storage.
Cloud Storage Data Lake
Flip cardUsing Cloud Storage as the primary, cost-effective, and flexible storage layer for raw, unstructured, and semi-structured data in a data lake architecture.
- Stores objects of any size and type.
- Offers high durability, availability, and scalability.
- Provides various storage classes for cost optimization.
Memory trick: Cloud Storage: The Lake's Foundation, holding all the Data's nation.
Dataflow for Batch Processing
Flip cardA fully managed service for executing Apache Beam pipelines on Google Cloud, supporting large-scale batch data transformations with automatic scaling and fault tolerance.
- Serverless and auto-scaling.
- Supports Apache Beam SDK.
- Fault-tolerant execution.
Memory trick: Dataflow's Beam of light transforms batches with serverless might.
Global Relational Database
Flip cardA globally distributed, strongly consistent, and highly available relational database with unlimited scale, suitable for critical operational workloads.
- Offers strong consistency across regions.
- Provides high availability and unlimited scalability.
- Supports SQL queries and relational schemas.
Memory trick: Spanner: Spanning the globe with Strong, Scalable, Speedy data.
Secure & Compliant Data System
Flip cardDesigning data systems on Google Cloud that meet strict regulatory requirements for data residency, encryption, access control, auditability, and cost-effectiveness at scale.
- Uses BigQuery for petabyte-scale analytics.
- Leverages KMS for encryption key management.
- Implements IAM for granular access control.
- Relies on Cloud Audit Logs for comprehensive auditability.
Memory trick: Securely Analyze, Encrypt, Control, and Log all Financial Data for Compliance.
Bigtable for Operational Analytics
Flip cardA fully managed, petabyte-scale NoSQL wide-column database optimized for high-throughput, low-latency operational analytics, suited for time-series and complex aggregations.
- Handles millions of ops/sec with low latency.
- Ideal for time-series and operational data.
- Scales to petabytes of data.
Memory trick: Bigtable: Big throughput, Big data, Big analytics.
Hybrid Real-time & Batch Pipeline
Flip cardA data pipeline combining real-time stream processing for immediate insights with batch processing for comprehensive historical analysis, designed for high availability and scalability.
- Uses Pub/Sub for real-time ingestion.
- Leverages Dataflow for unified stream and batch processing.
- Stores data in BigQuery for analytics and Cloud Storage for raw/long-term.
Memory trick: Pub/Sub, Dataflow, BigQuery, Storage: The Hybrid's Core for Gaming Data.
BigQuery Materialized Views
Flip cardPre-computed views in BigQuery that store the results of a query, accelerating subsequent queries that use the same underlying logic.
- Accelerates complex queries (joins, aggregations).
- Automatically maintained by BigQuery.
- Reduces query costs and latency.
Memory trick: Materialized views are like having the answer key for frequently asked questions, so you don't have to re-solve them every time.
Dataflow Exactly-Once Semantics
Flip cardA guarantee in Dataflow streaming pipelines that each data element is processed and committed exactly one time, even when failures occur.
- Prevents data duplication and data loss.
- Achieved through checkpointing and persistent state.
- Essential for accurate real-time analytics and financial transactions.
Memory trick: Dataflow Delivers Diligently.
BigQuery Time Travel
Flip cardBigQuery Time Travel allows you to access data from any point within the past 7 days, enabling point-in-time recovery and historical analysis without explicit backups.
- Default retention period of 7 days (can be configured for less).
- Recovers from accidental deletions or table updates.
- Does not require explicit backups or manual configuration.
- Accessible via SQL queries using FOR SYSTEM_TIME AS OF.
Memory trick: Time Travel: Go back in time, fix the data crime.
Dataproc Preemptible Workers
Flip cardLower-cost Dataproc workers that can be shut down (preempted) by Google Cloud if resources are needed elsewhere, suitable for fault-tolerant batch workloads.
- Offer significant cost savings (typically 60-91% off).
- Ideal for batch jobs that can tolerate interruptions.
- Dataproc automatically reschedules tasks from preempted workers.
Memory trick: Preemptively Pick Pennies.
BigQuery Right to Be Forgotten & Row-Level Security
Flip cardThe ability to delete specific customer data upon request using DML statements, combined with granular access control to individual data rows via BigQuery row-level security.
- DML DELETE statements permanently remove data.
- Row-level security policies filter rows visible to users.
- Essential for GDPR, CCPA, and other data privacy regulations.
Memory trick: Rows Remove Rights, Roles Restrict.
Dataproc for ML Experimentation
Flip cardA fully managed service on Google Cloud for running Apache Spark, Hadoop, and other open-source data tools, ideal for flexible and scalable data processing in ML workflows.
- Managed Spark, Hadoop, Presto clusters.
- Rapid cluster provisioning/de-provisioning.
- Scalable and cost-effective for experimentation.
- Leverages familiar open-source tools.
Memory trick: Dataproc is like a flexible workbench for data scientists: you can quickly grab any open-source tool (Spark, Hadoop) and spin up a cluster to test your ideas.
Google Cloud Dataplex
Flip cardAn intelligent data fabric that unifies distributed data, automates data quality and governance, and enables data discovery across an organization's data landscape.
- Provides a single pane of glass for managing data lakes, data warehouses, and data marts.
- Automates data quality, metadata management, and governance.
- Facilitates data discovery and self-service analytics.
Memory trick: Dataplex Delivers Data's Domain.
GCP Monitoring & Alerting for Streaming Data
Flip cardUsing Google Cloud services to collect logs, extract metrics, and trigger alerts based on real-time streaming data for anomaly detection.
- Cloud Logging for log collection.
- Log-based metrics for data extraction.
- Cloud Monitoring for alerts and dashboards.
Memory trick: For real-time anomaly detection, think of logs as raw ingredients, log-based metrics as your recipe, and Cloud Monitoring as the chef's alarm when something's burning.
Cloud Key Management Service (KMS)
Flip cardA cloud-hosted key management service that lets you manage cryptographic keys for your cloud services in the same way you manage keys on-premises.
- Supports symmetric and asymmetric encryption.
- Integrates with many Google Cloud services for CMEK.
- Provides auditing and access control for keys.
Memory trick: Keys Keep Cloud Secrets Safe.
BigQuery Flex Slots
Flip cardAn hourly commitment option for BigQuery Flat-Rate pricing that allows users to purchase dedicated compute capacity (slots) on an hourly basis.
- Offers more flexibility than annual/monthly commitments.
- Provides cost predictability and dedicated performance for peak loads.
- Ideal for fluctuating workloads or testing Flat-Rate pricing.
Memory trick: Flex For Fluctuating Futures.
BigQuery Autoscaling Slots
Flip cardA feature of BigQuery's flat-rate pricing that automatically adjusts the number of slots (compute capacity) allocated to a project or organization based on real-time workload demand.
- Dynamically scales compute capacity up or down.
- Aims to maintain consistent query performance.
- Available with BigQuery flat-rate pricing.
Memory trick: Scale Smoothly, Save Smartly.
BigQuery Flat-Rate Pricing (Slots)
Flip cardA BigQuery pricing model where users commit to a fixed amount of dedicated compute capacity (slots) per second, providing predictable performance and costs.
- Guarantees a consistent level of query performance.
- Ideal for large, enterprise-level workloads with high concurrency.
- Cost is predictable, regardless of data scanned (after commitment).
Memory trick: Slots Scale Swiftly.
Data Transformation Traceability
Flip cardThe ability to track and audit all changes to data transformation logic and schema, ensuring immutability and compliance.
- Version control for pipeline code.
- CI/CD for controlled deployments.
- Auditable history of logic changes.
Memory trick: To audit your data transformations, treat your code like a legal document: version it, sign it (CI/CD), and keep every draft.
Google Cloud Data Catalog
Flip cardA fully managed, scalable metadata management service that helps organizations discover, understand, and manage all their data assets.
- Centralized metadata repository.
- Enables data discovery and search.
- Supports policy tagging and lineage.
Memory trick: Data Catalog is like the library's card catalog for your cloud data: it helps you find, understand, and manage everything.
Apache Avro
Flip cardA row-oriented remote procedure call and data serialization framework developed within the Apache Hadoop project.
- Includes schema in the data file, enabling schema evolution.
- Supports rich data structures and efficient binary serialization.
- Widely used in big data ecosystems for data storage and interchange.
Memory trick: Formats Fuel Future Files.
GCP Monitoring for Streaming Data Quality
Flip cardLeveraging Cloud Monitoring and Cloud Logging to collect metrics and logs from streaming pipelines to detect and alert on data quality issues in real-time.
- Custom metrics track key performance indicators (KPIs) like event counts, latencies, error rates.
- Log-based metrics and alerts can detect specific data patterns or errors in logs.
- Integrates with Pub/Sub, Dataflow, and other streaming services.
Memory trick: Monitor Metrics, Log Lapses.
BigQuery Pricing Models
Flip cardBigQuery offers different pricing models to accommodate various workloads and cost predictability needs.
- On-demand: Pay per query (data scanned).
- Flat-rate: Pay for dedicated processing capacity (slots).
- Storage: Pay for data stored.
Memory trick: BigQuery pricing is like choosing between a taxi (on-demand) or renting a car for a month (flat-rate).
BigQuery On-demand with Autoscaling Slots
Flip cardCombining BigQuery on-demand pricing with autoscaling slots provides a cost-effective and highly responsive solution for highly variable workloads by paying for query data processed and dynamically adjusting processing capacity.
- On-demand: Pay for data processed by queries (first 1 TB free/month).
- Autoscaling Slots: Automatically adjust processing capacity (slots) based on workload demand.
- Ideal for unpredictable, bursty workloads.
- Optimizes costs by only paying for what's used, scaling down to zero during idle times.
Memory trick: Auto-scale the demand, pay as you land.
Cloud DLP for Sensitive Data Discovery
Flip cardCloud Data Loss Prevention (DLP) is a fully managed service designed to discover, classify, and protect sensitive data at scale.
- Identifies over 150 types of sensitive data (e.g., PII, financial, credentials).
- Works across various data sources (Cloud Storage, BigQuery, Datastores, images, text).
- Offers de-identification techniques like tokenization, masking, and redaction.
Memory trick: Discover, Classify, Protect: DLP is the data's best detective.
Cloud DLP De-identification
Flip cardA set of techniques within Cloud Data Loss Prevention (DLP) to transform sensitive data into a less sensitive format while preserving utility, with options for re-identification.
- Includes tokenization, pseudonymization, format-preserving encryption.
- Used to comply with data privacy regulations like HIPAA, GDPR.
- Can be configured to allow controlled re-identification of data.
Memory trick: DLP Detects, De-identifies, Deciphers.
BigQuery Multi-Region Datasets
Flip cardBigQuery datasets configured to store data across multiple geographical regions, providing high availability and disaster recovery.
- Data automatically replicated by BigQuery.
- Protects against regional outages.
- Ensures data durability and availability.
Memory trick: BigQuery's multi-region is like having data mirrored in two different continents, so if one falls, the other instantly takes over.
BigQuery Dataset Location
Flip cardA configuration setting for a BigQuery dataset that determines the geographic location where the dataset's data is stored and processed.
- Can be set to a specific region (e.g., `us-east1`) or a multi-region (e.g., `US`, `EU`).
- Enforces data residency requirements.
- Cannot be changed after the dataset is created.
Memory trick: Location Limits Logic.
Secure & Compliant Data Warehouse
Flip cardDesigning a data warehouse system on Google Cloud that adheres to strict regulatory requirements for encryption, access control, data residency, and auditability.
- BigQuery offers CMEK, IAM, data location, and audit logs.
- CMEK ensures encryption key control by the customer.
- IAM provides granular access management.
- Data location settings enforce residency requirements.
Memory trick: Encrypt, Control, Locate, Audit - The four pillars of compliance.