AWS Certified Data Engineer – Associate flashcards
118 free flashcards. Tap a card to flip it.
Apache Parquet
Flip cardApache Parquet is a free and open-source columnar storage format designed for efficient data storage and retrieval, especially for big data processing frameworks.
- Columnar storage format.
- Optimized for analytical queries.
- Supports complex nested data structures.
- Offers efficient compression and encoding schemes.
Memory trick: Columnar's the key, for speed you see, Parquet sets data free.
Amazon S3 for Data Lakes
Flip cardAmazon S3 is a highly scalable, durable, and cost-effective object storage service widely used as the foundation for data lakes on AWS.
- Stores unstructured data (images, videos, logs, etc.)
- Offers various storage classes for cost optimization
- Provides high durability and availability
- Integrates with other AWS analytics services
Memory trick: Scalable Storage Solutions are Key for Data Lakes.
AWS Glue Data Catalog
Flip cardThe AWS Glue Data Catalog is a persistent metadata store for all your data assets, regardless of where they are located. It allows you to store, annotate, and share metadata, enabling various AWS services to discover and query your data.
- Centralized metadata repository
- Integrates with S3, RDS, Redshift, etc.
- Supports schema discovery (Crawlers)
- Used by Athena, Redshift Spectrum, EMR, Glue ETL
Memory trick: Glue Guides Global Data.
Snappy Compression
Flip cardSnappy is a fast compression/decompression library developed by Google, designed for high-speed data processing rather than maximum compression ratio.
- Optimized for speed (fast compression and decompression).
- Lower CPU overhead during processing.
- Good compression ratio, though not the highest.
- Widely used in big data ecosystems (e.g., Hadoop, Spark, Parquet).
Memory trick: Snappy's swift, for queries to lift, GZIP's tight, but slows the light.
Amazon Kinesis Data Streams
Flip cardAmazon Kinesis Data Streams is a massively scalable, durable, real-time data streaming service that continuously captures gigabytes of data per second from hundreds of thousands of sources.
- Real-time data ingestion and processing.
- Scalable to handle high-throughput data streams.
- Data retention up to 365 days.
- Supports multiple consumers reading from the same stream.
Memory trick: Kinesis streams, for real-time dreams, no data's lost, it always gleams.
AWS Lake Formation
Flip cardA service that helps you build, secure, and manage data lakes by simplifying security, governance, and auditing, offering fine-grained access control.
- Centralized data lake security and governance
- Fine-grained access control (table, column, row level)
- Integrates with AWS Glue Data Catalog
- Simplifies permissions for analytics services
Memory trick: Lake Formation is the 'Lake Guard' for your data lake's security and rules.
EFS for AWS Lambda
Flip cardAmazon EFS (Elastic File System) can be mounted to AWS Lambda functions, providing persistent, scalable, and shared file storage for code, libraries, and temporary state that exceeds Lambda's ephemeral storage limits.
- Provides up to 10 GB of ephemeral storage for /tmp, EFS extends this.
- Persistent and shared file system across Lambda invocations and functions.
- Low-latency file access.
- Scales automatically to petabytes.
Memory trick: EFS to Lambda, like a shared drive, keeps your functions alive.
S3 Glacier Deep Archive
Flip cardAmazon S3 Glacier Deep Archive is Amazon S3's lowest-cost storage class, designed for long-term archival of data that is accessed rarely (e.g., once or twice a year) and can tolerate retrieval times of several hours.
- Lowest cost S3 storage class
- Designed for long-term archival (7-10+ years)
- Retrieval times in hours
- High durability (11 nines)
Memory trick: Deep Archive: deep savings for deep sleep data.
S3 Encryption with KMS (SSE-KMS)
Flip cardServer-Side Encryption for Amazon S3 objects using keys managed by AWS Key Management Service (KMS), providing customer control and auditing of encryption keys.
- S3 encrypts objects server-side
- Uses AWS KMS for key management
- Customer controls KMS key policies and rotation
- Key usage is auditable via AWS CloudTrail
Memory trick: KMS gives you the 'Key Master' control for your S3 encryption.
Amazon Timestream
Flip cardA fast, scalable, and serverless time-series database service that makes it easy to store and analyze petabytes of time-series data.
- Purpose-built for time-series data (IoT, operational apps)
- High throughput ingestion and efficient storage
- Serverless and fully managed
- SQL-like query language with time-series functions
Memory trick: Timestream is the 'Time' machine for your 'Stream' of IoT data.
Amazon S3 Standard
Flip cardAmazon Simple Storage Service (S3) Standard is an object storage service offering high durability, availability, and performance for frequently accessed data.
- Object storage for unstructured data
- Virtually unlimited scalability
- Designed for 99.999999999% (11 nines) durability
- Low-latency access
Memory trick: S3 stores everything, simply and securely.
Redshift DISTSTYLE KEY
Flip cardAmazon Redshift's DISTSTYLE KEY distributes rows of a table across compute nodes based on the hash value of a chosen column. This is crucial for optimizing join performance in data warehouses.
- Co-locates data for efficient joins.
- Choose a column frequently used in join predicates (e.g., foreign key).
- Minimizes data movement (shuffling) during query execution.
- Ideal for large fact tables in star schemas.
Memory trick: Key for joins, all for small, even for none, auto for all.
S3 Lifecycle Policies
Flip cardS3 Lifecycle policies define rules to automatically transition objects to different S3 storage classes or expire them after a specified time, optimizing storage costs based on access patterns.
- Automate storage class transitions
- Optimizes storage costs
- Based on object age or access patterns
- Can also be used to expire objects
Memory trick: Standard to IA, then to Glacier for cheap, deep archival.
AWS Glue ETL Jobs
Flip cardAWS Glue ETL jobs are serverless Apache Spark-based jobs that enable data engineers to extract, transform, and load data at scale, automatically handling provisioning, setup, and scaling of compute resources.
- Serverless Spark environment
- Automates ETL processes
- Integrates with Glue Data Catalog
- Ideal for data transformation, compaction, format conversion
Memory trick: Glue ETL jobs transform data, serverless and grand.
S3 Glacier Instant Retrieval
Flip cardS3 Glacier Instant Retrieval is an archive storage class that delivers the lowest cost storage for data that is rarely accessed, but requires immediate retrieval (in milliseconds).
- Lowest cost storage for infrequently accessed data.
- Retrieval in milliseconds.
- Designed for archive data that needs immediate access.
- Offers 11 nines of durability.
Memory trick: Standard's fast, IA's a blend, Glacier's slow, but low cost to the end.
Amazon S3 for Unstructured Data
Flip cardAmazon S3 (Simple Storage Service) is an object storage service offering industry-leading scalability, data availability, security, and performance. It is ideal for storing unstructured data like images, videos, and backups.
- Object storage for any data type
- High durability and availability
- Scalable to petabytes/exabytes
- Pay-as-you-go pricing
Memory trick: S3 Stores Simple Stuff Superbly.
Amazon S3 Versioning
Flip cardAmazon S3 Versioning allows you to keep multiple versions of an object in the same bucket, providing protection against accidental overwrites, deletions, and enabling easy recovery of previous object states.
- Retains multiple object versions
- Protects against accidental overwrites/deletions
- Can recover previous object states
- Uses version IDs to distinguish objects
Memory trick: Versioning keeps every version, never losing a thing.
Redshift KEY Distribution
Flip cardAmazon Redshift KEY distribution ensures that rows with the same value in the specified distribution column are stored on the same compute node. This is crucial for optimizing join performance between large fact tables and dimension tables.
- Distributes data based on a column's value
- Optimizes join performance by co-locating data
- Used for large fact tables joined with dimension tables
- Choose a column with high cardinality for even distribution
Memory trick: Keys Keep Joins Quick.
Amazon Aurora
Flip cardAmazon Aurora is a MySQL and PostgreSQL-compatible relational database built for the cloud, combining the performance and availability of traditional enterprise databases with the simplicity and cost-effectiveness of open-source databases.
- Relational database service
- ACID compliant
- High performance and scalability
- Compatible with MySQL and PostgreSQL
Memory trick: ACID needs Aurora, Analytics needs Redshift, NoSQL needs DynamoDB.
S3 Data Partitioning
Flip cardOrganizing data in Amazon S3 by creating a hierarchical directory structure based on column values, enabling query engines to filter data more efficiently.
- Improves query performance in services like Athena
- Reduces the amount of data scanned, lowering costs
- Commonly partitions by time (year, month, day) or other categorical data
- Data is stored in folders like `s3://bucket/table/year=YYYY/month=MM/day=DD/`
Memory trick: Divide and conquer your data with partitions to make queries faster and cheaper.
MSCK REPAIR TABLE
Flip cardThe `MSCK REPAIR TABLE` command (available in Athena/Presto/Hive) scans the underlying data location (e.g., S3) for a partitioned table and updates the AWS Glue Data Catalog with any new partitions it discovers.
- Used to add new partitions to the Glue Data Catalog.
- Scans the S3 path of a table based on its partition scheme.
- Essential for improving query performance by making all data discoverable.
- More efficient for partition updates than running a full Glue Crawler.
Memory trick: MSCK repair, makes partitions fair, for queries to declare.
Data Lake Querying with Athena
Flip cardAmazon Athena is an interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It's serverless, so there's no infrastructure to manage, and you pay only for the queries you run.
- Serverless interactive query service
- Queries data directly in S3
- Uses standard SQL
- Pay-per-query pricing model
Memory trick: Athena Answers Any Analytics on S3.
S3 Lifecycle Policies for Cost Optimization
Flip cardS3 Lifecycle policies define rules to automatically transition objects to different S3 storage classes or expire them, based on age or other criteria. This helps optimize storage costs by moving data to more cost-effective classes as access patterns change.
- Automates object transitions and expirations
- Reduces storage costs
- Based on object age or access patterns
- Applies to specific prefixes or entire buckets
Memory trick: Standard, IA, Glacier: Scale, Access, Archive.
Redshift Encryption (KMS & SSL)
Flip cardAmazon Redshift can encrypt data at rest using AWS KMS for key management and enforce SSL/TLS for data in transit, ensuring comprehensive protection for sensitive data.
- Data at rest encrypted by default
- Use KMS for customer-managed keys, rotation, and auditing
- SSL/TLS for data in transit encryption
- Crucial for regulatory compliance with sensitive data
Memory trick: KMS for rest, SSL for transit, keep data best.
Redshift DISTSTYLE ALL
Flip cardA Redshift table distribution style that copies the entire table to every compute node, typically used for small dimension tables to optimize join performance.
- Copies full table to all compute nodes
- Eliminates data movement for joins with large fact tables
- Best for small dimension tables (few hundred MBs/thousands of rows)
- Increases storage footprint, so not suitable for large tables
Memory trick: For small tables, 'ALL' the nodes need 'ALL' the data for super-fast joins.
Amazon EFS for AWS Lambda
Flip cardAmazon EFS provides scalable, shared, and persistent file storage that can be mounted by AWS Lambda functions, enabling them to process large files and share data across invocations.
- Persistent storage for Lambda functions
- Shared access across multiple Lambda functions
- Scalable and elastic file system
- Supports file system semantics (NFS)
Memory trick: EFS is the 'Elastic File Share' for your Lambda data needs.
AWS Glue ETL
Flip cardA serverless data integration service that makes it easy to discover, prepare, move, and combine data for analytics, machine learning, and application development.
- Fully managed, serverless, and scales dynamically.
- Uses Apache Spark for powerful data transformations.
- Includes Glue Data Catalog for metadata management.
Memory trick: Glue: The serverless sticky solution for your data transformations.
AWS Budgets for Cost Control
Flip cardAn AWS service that allows users to set custom budgets to track their costs and usage, and to receive alerts when actual or forecasted amounts exceed the budgeted thresholds.
- Tracks costs, usage, or reservation utilization/coverage.
- Can be scoped to specific services, linked accounts, or cost allocation tags.
- Sends alerts via SNS or email when thresholds are breached.
Memory trick: Budget with Tags: Track your pipeline's cash, not just its lags.
AWS DataSync for Large-Scale Migration & Sync
Flip cardA service that automates and accelerates the online transfer of large amounts of data between on-premises storage and AWS storage services, supporting initial migration and continuous synchronization.
- High-performance, online data transfer
- Supports file shares (NFS, SMB) and object storage
- Automates, accelerates, and synchronizes data for ongoing changes
Memory trick: DataSync carries big data loads, on-prem to cloud, secure roads.
Amazon EMR with Apache Spark
Flip cardA managed cluster platform that simplifies running big data frameworks, such as Apache Spark, Hadoop, Presto, and Hive, on AWS to process and analyze vast amounts of data.
- Provides a managed Apache Spark (or other framework) cluster.
- Offers full control over the cluster and framework configuration.
- Scalable resources, can be provisioned and de-provisioned as needed.
- Ideal for complex, custom, and batch-oriented big data processing.
Memory trick: EMR is your 'Personal Lab' for Big Data experiments.
Step Functions Monitoring/Alerting
Flip cardMonitor AWS Step Functions pipelines using CloudWatch Alarms on execution metrics (e.g., failures, duration) and integrate with Amazon SNS for centralized, proactive alerting.
- CloudWatch is the primary monitoring service.
- Alarms trigger on metric thresholds.
- SNS provides flexible notification channels.
Memory trick: CloudWatch Alarms for Step Functions, SNS for the news.
Glue Job Troubleshooting
Flip cardThe primary step for troubleshooting AWS Glue job failures is to examine CloudWatch Logs for detailed error messages, stack traces, and resource utilization metrics to identify the root cause.
- CloudWatch Logs are central for Glue job diagnostics.
- 'Non-zero exit code' is a generic error needing log analysis.
- Resource constraints (memory, disk) are common causes.
Memory trick: When Glue jobs crash, the logs hold the truth.
AWS Glue Streaming ETL for IoT
Flip cardA serverless, Apache Spark-based service tailored for continuous, real-time transformation, enrichment, and schema-aware processing of high-volume streaming data from sources like IoT devices.
- Serverless and scalable for real-time streams
- Automatic schema inference and evolution handling
- Supports complex transformations and data enrichment
Memory trick: Glue streams IoT data, transforming schema-aware.
AWS DataSync
Flip cardAn online data transfer service that simplifies, automates, and accelerates moving data between on-premises storage and Amazon S3, Amazon EFS, or Amazon FSx for Windows File Server.
- Accelerates data transfers over WANs.
- Provides automatic retry, encryption, and integrity verification.
- Managed service, eliminating custom scripting and infrastructure.
- Ideal for large files and datasets, ensuring reliable delivery.
Memory trick: DataSync is the 'Bridge' for on-prem to S3 data.
AWS Storage Gateway (File Gateway)
Flip cardA hybrid cloud storage service that provides on-premises applications with file-based access to cloud storage (S3) using standard file protocols like NFS and SMB.
- Acts as a cache for frequently accessed data on-premises.
- Translates standard file protocols (SMB/NFS) to S3 object storage.
- Enables easy integration of on-premises applications with S3.
Memory trick: File Gateway bridges on-prem shares to S3, making it feel local.
Kinesis Data Streams + Glue Streaming ETL
Flip cardA powerful combination for real-time data ingestion and serverless, continuous transformation of streaming data, enabling delivery to various AWS destinations.
- Kinesis Data Streams for real-time ingestion
- Glue Streaming ETL for serverless, continuous transformations
- Supports multiple destinations like S3 and Redshift
Memory trick: Kinesis streams flow, Glue transforms, data glows.
Glue Job Bookmarks
Flip cardAWS Glue Job Bookmarks enable incremental processing by tracking previously processed data, ensuring that subsequent job runs only process new or changed data from the source, reducing processing time and cost.
- Tracks processed data for incremental runs.
- Reduces data scanned and processed.
- Lowers Glue ETL costs and execution time.
Memory trick: Bookmarks make Glue smart, only processing the new parts.
Amazon Kinesis Data Streams for Event Ingestion
Flip cardA real-time data streaming service capable of ingesting high-volume, low-latency event data from numerous sources, providing ordered and durable storage for subsequent processing.
- Real-time, high-volume ingestion
- Low latency and durable storage
- Ordered records within a shard
Memory trick: Kinesis streams game events in order.
AWS Glue ETL Batch Processing
Flip cardA serverless service that allows for scheduled, cost-effective batch transformations of data from various sources (like S3) into optimized formats (like Parquet) for analytical workloads.
- Serverless and pay-per-use
- Handles large batch datasets
- Supports diverse data formats and complex transformations
Memory trick: Glue processes nightly batches, making data fit for matches.
Redshift Managed Scaling
Flip cardAmazon Redshift Managed Scaling automatically adjusts the number of nodes in a Redshift cluster up or down based on workload demand, optimizing performance and cost by right-sizing the cluster capacity.
- Automates cluster resizing (add/remove nodes).
- Optimizes for both peak and off-peak workloads.
- Reduces costs by preventing over-provisioning.
Memory trick: Redshift Managed Scaling is the 'Smart Size' for your wallet and speed.
AWS Glue Security Configurations
Flip cardAWS Glue Security Configurations allow you to specify encryption settings for data at rest (S3, CloudWatch logs, Job bookmarks) and data in transit within the Glue environment, using AWS KMS keys.
- Centralized encryption management for Glue.
- Encrypts S3 data at rest, CloudWatch logs, and Job bookmarks.
- Encrypts data in transit within Glue processing.
- Uses AWS KMS for key management.
Memory trick: Glue's Secure Config: Key to Protecting Data's Journey.
Kinesis Data Analytics for Apache Flink
Flip cardA fully managed service for processing and analyzing streaming data in real time using Apache Flink, enabling complex, stateful transformations and custom code execution.
- Fully managed Apache Flink
- Real-time, stateful stream processing
- Supports custom Java/Scala/Python code
Memory trick: Flink flows with smart real-time insights.
Step Functions Long-Term Auditing
Flip cardTo retain AWS Step Functions execution history beyond 90 days for auditing, stream CloudWatch Logs (which capture execution details) to Amazon S3 and manage retention using S3 Lifecycle policies.
- Default Step Functions retention is 90 days.
- CloudWatch Logs capture detailed execution events.
- S3 Lifecycle policies manage long-term data retention cost-effectively.
Memory trick: CloudWatch logs everything, S3 stores it forever, cheap.
Step Functions Retry/Catch
Flip cardBuilt-in error handling mechanisms in AWS Step Functions that allow individual states to automatically re-attempt execution (Retry) or transition to a different state upon encountering specific errors (Catch).
- Retry policies include max attempts, interval, and backoff rate.
- Catchers can define specific error codes to handle.
- Prevents entire workflow failures due to transient issues.
Memory trick: When a Step fails, Catch it and Retry, or choose a new path to try.
AWS Glue Job Bookmarks
Flip cardA feature in AWS Glue that helps process incremental data by tracking the state of previously processed data, allowing ETL jobs to process only new or changed data on subsequent runs.
- Saves processing state information in a persistent store (DynamoDB).
- Reduces processing time and DPU costs.
- Supports various data sources (S3, JDBC).
- Must be enabled and configured in the Glue job properties.
Memory trick: Bookmarks mark the spot, so old data gets forgot. New data only, fast the job will trot.
AWS IoT Core with IoT Rules Engine
Flip cardA fully managed service that enables connected devices to easily and securely interact with cloud applications and other devices, with a Rules Engine to filter, transform, and route messages.
- Connects millions of IoT devices securely.
- IoT Rules Engine allows filtering, transforming, and routing messages.
- Supports multiple destinations (S3, Kinesis, Lambda, DynamoDB, etc.).
- Fully managed and highly scalable for high-velocity IoT data.
Memory trick: IoT Core is the 'Traffic Cop' for device data.
MWAA Centralized Alerting
Flip cardLeveraging AWS CloudWatch Logs, Alarms, EventBridge, and Lambda to create a scalable and centralized mechanism for monitoring and alerting on MWAA DAG failures with detailed contextual information.
- MWAA logs are automatically sent to CloudWatch Logs.
- CloudWatch Alarms can detect specific error patterns in logs.
- EventBridge routes events from Alarms to target services.
- Lambda functions can parse event data and send rich notifications.
Memory trick: MWAA logs to CloudWatch, alarms to EventBridge, Lambda delivers Slack's rich message.
AWS Transfer Family + AWS Glue ETL
Flip cardA common pattern for ingesting files from external partners via SFTP/FTPS/FTP directly to S3, followed by serverless batch ETL using AWS Glue to transform and standardize the data.
- AWS Transfer Family for managed file transfer protocols (SFTP, FTPS, FTP)
- AWS Glue ETL for serverless batch data transformation
- Directly integrates with Amazon S3 for source and target
Memory trick: Transfer Family gets files, Glue transforms them to golden piles.
Amazon Athena
Flip cardAn interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It is serverless and pays per query.
- Serverless SQL query service
- Queries data directly in S3
- Pay-per-query pricing model
Memory trick: Athena queries S3 data like a wise owl.
Amazon Kinesis Data Analytics for Apache Flink
Flip cardA fully managed service that allows you to process and analyze streaming data in real time using Apache Flink.
- Supports complex stream processing, aggregations, and windowing.
- Automatically scales to handle varying data throughput.
- Integrates with Kinesis Data Streams and Firehose for input/output.
Memory trick: Flink analyzes streams fast, finding patterns and anomalies.
Glue I/O Optimization
Flip cardOptimizing I/O-bound AWS Glue ETL jobs, especially when data is in S3, involves reducing the amount of data read by partitioning and leveraging predicate pushdown.
- High I/O Wait and low CPU indicate I/O bottleneck.
- Partitioning S3 data reduces data scanned.
- Predicate pushdown filters data at the source.
- Improves job performance and reduces costs.
Memory trick: Slow Glue Job? Check your S3 partitions, then push down your predicates to speed up the data flow.
AWS Glue Data Quality (Deequ)
Flip cardA feature within AWS Glue that allows users to define and run data quality rules directly within their Glue ETL jobs to validate data at various stages of the pipeline.
- Open-source Deequ library integrated into Glue.
- Supports various data quality checks (e.g., completeness, uniqueness, consistency, validity).
- Can be configured to fail jobs or trigger alerts based on rule violations.
Memory trick: Deequ Detects Data Discrepancies, Delivering Dependable Data.
AWS Glue Streaming ETL
Flip cardA serverless data integration service that allows you to continuously process streaming data for ETL (Extract, Transform, Load) operations.
- Processes data from Kinesis and Kafka.
- Supports complex transformations, joins, and filtering.
- Automatically scales and manages infrastructure.
Memory trick: Glue streams transform, enrich, and filter data on the fly.
EventBridge Dead-Letter Queues (DLQ)
Flip cardA feature in Amazon EventBridge that allows undelivered events (after all retry attempts) to be sent to an SQS queue or SNS topic for later inspection, analysis, and reprocessing, ensuring no data loss.
- Configured at the target level for an EventBridge rule.
- Captures events that fail delivery after retries.
- Typically uses Amazon SQS as the underlying queue.
- Essential for building robust, fault-tolerant event-driven architectures.
Memory trick: EventBridge DLQ: Dead-Letter Queue keeps failed events, not lost due.
AWS Pipeline Monitoring Stack
Flip cardA combination of AWS services used to observe, track, and alert on the health and performance of data pipelines, covering execution, resource utilization, and data flow.
- CloudWatch for metrics, logs, and alarms.
- SNS for sending notifications.
- X-Ray for distributed tracing across services.
- EventBridge for event-driven automation and notifications.
Memory trick: CloudWatch's Metrics and Alarms are the Eyes, SNS the Voice, X-Ray the X-ray Vision.
AWS Snowball Edge
Flip cardA data migration and edge computing device that securely transfers large amounts of data to and from AWS.
- Available in various capacities (e.g., Storage Optimized, Compute Optimized).
- Used for offline data transfer when network transfer is impractical.
- Provides tamper-resistant enclosures and end-to-end encryption.
Memory trick: Snowball Edge: Your physical truck for massive data moves.
Amazon Kinesis Data Firehose
Flip cardA fully managed service for delivering real-time streaming data to destinations such as Amazon S3, Amazon Redshift, Amazon OpenSearch Service, and Splunk.
- Fully managed and scales automatically.
- Supports various destinations (S3, Redshift, etc.).
- Offers built-in data transformation (e.g., format conversion, Lambda for custom logic).
- Batching, compression, and encryption are handled automatically.
Memory trick: Firehose 'sprays' your data to multiple destinations.
Glue Redshift Write Optimization
Flip cardOptimize AWS Glue writes to Amazon Redshift by using 'use_s3_dist_copy' and 'num_partitions' to stage data in S3 and leverage Redshift's parallel COPY command, reducing cluster load.
- S3 staging + Redshift COPY is highly efficient.
- Reduces direct Redshift cluster load.
- Parameters: use_s3_dist_copy, num_partitions/num_files.
Memory trick: Glue's best Redshift friend is S3, for a fast 'COPY' party.
AWS Database Migration Service (DMS)
Flip cardA cloud service that makes it easy to migrate relational databases, data warehouses, NoSQL databases, and other types of data stores.
- Supports homogeneous and heterogeneous migrations.
- Enables continuous data replication (Change Data Capture).
- Automates schema and data type conversions.
Memory trick: DMS moves and mirrors databases, even different types.
Glue ETL for Data Lake Refinement
Flip cardUtilizing AWS Glue ETL jobs to transform raw data in a data lake into optimized formats (e.g., Parquet, ORC) and structures (e.g., partitioning) for improved query performance and reduced costs.
- Serverless Apache Spark environment.
- Integrates with Glue Data Catalog for schema management.
- Supports various data sources and targets.
- Ideal for batch processing and large-scale transformations.
Memory trick: Glue's ETL is the Gold Standard for transforming your data lake's CSVs to Parquet, saving bucks and boosting speed.