Free knowledge base

AWS Certified Data Engineer – Associate — key terms, tricks & tips

Everything from the course in one searchable place: 224 entries. Use it to review before a practice test or look up a word you forgot.

224 results

Key term

Exam Guide

Official AWS document detailing exam scope and objectives.

Getting Started: Exam Essentials

Key term

Domain

A major knowledge area covered by the exam.

Getting Started: Exam Essentials

Key term

Weighting

The percentage of questions from a specific domain.

Getting Started: Exam Essentials

Key term

Multiple-choice

Question with one correct answer out of several options.

Getting Started: Exam Essentials

Key term

Multiple-response

Question with two or more correct answers to select.

Getting Started: Exam Essentials

Key term

Scaled score

Raw score converted to a standard scale (100-1000).

Getting Started: Exam Essentials

Key term

Unscored questions

Experimental questions not affecting your final score.

Getting Started: Exam Essentials

Memory trick

Understanding the DEA-C01 Exam Structure

Remember 'DATA' for the domains: Data Ingestion (34%), Data Storage (26%), Data Orchestration (25%), Data Security (15%).

Getting Started: Exam Essentials

Exam tip

Understanding the DEA-C01 Exam Structure

The DEA-C01 exam is 130 minutes long and contains 65 questions. The passing score is 700 out of 1000. Memorize these exact numbers.

Getting Started: Exam Essentials

Common mistake

Understanding the DEA-C01 Exam Structure

Not reviewing the official Exam Guide, which is the definitive source for exam content.

Getting Started: Exam Essentials

Common mistake

Understanding the DEA-C01 Exam Structure

Ignoring the domain weightings, leading to over-studying less important topics.

Getting Started: Exam Essentials

Common mistake

Understanding the DEA-C01 Exam Structure

Misinterpreting multiple-response questions and selecting too few or too many answers.

Getting Started: Exam Essentials

Key term

Learning Style

Preferred method for absorbing and processing information.

Getting Started: Exam Essentials

Key term

AWS Skill Builder

Official AWS platform for digital training and labs.

Getting Started: Exam Essentials

Key term

Spaced Repetition

Reviewing information at increasing intervals for retention.

Getting Started: Exam Essentials

Key term

AWS Free Tier

Free usage limits for many AWS services to gain experience.

Getting Started: Exam Essentials

Key term

Whitepaper

In-depth technical document on AWS services or architecture.

Getting Started: Exam Essentials

Memory trick

Effective Study Strategies and Resources

PLAN: P-Practice, L-Labs, A-AWS Resources, N-Notes.

Getting Started: Exam Essentials

Exam tip

Effective Study Strategies and Resources

The DEA-C01 exam heavily emphasizes practical application. Memorizing facts is not enough; you must understand how services interact and when to use them. Look for scenario-based questions.

Getting Started: Exam Essentials

Common mistake

Effective Study Strategies and Resources

Relying solely on third-party practice tests without understanding the underlying concepts from official AWS documentation.

Getting Started: Exam Essentials

Common mistake

Effective Study Strategies and Resources

Neglecting hands-on labs and practical experience, which is crucial for scenario-based exam questions.

Getting Started: Exam Essentials

Common mistake

Effective Study Strategies and Resources

Cramming all material at the last minute instead of using spaced repetition and regular review.

Getting Started: Exam Essentials

Key term

Batch Processing

Processing data in large, discrete chunks at scheduled intervals.

Data Ingestion and Transformation Fundamentals

Key term

Streaming Processing

Continuously processing data records as they are generated in real-time.

Data Ingestion and Transformation Fundamentals

Key term

Latency

The delay between data generation and its availability for processing or analysis.

Data Ingestion and Transformation Fundamentals

Key term

Throughput

The rate at which data can be processed or transferred over a period.

Data Ingestion and Transformation Fundamentals

Key term

Amazon Kinesis

A family of services for collecting, processing, and analyzing real-time streaming data.

Data Ingestion and Transformation Fundamentals

Key term

AWS DMS

AWS Database Migration Service, for migrating databases to AWS quickly and securely.

Data Ingestion and Transformation Fundamentals

Key term

AWS Glue

A serverless data integration service for ETL (Extract, Transform, Load) jobs.

Data Ingestion and Transformation Fundamentals

Key term

AWS Transfer Family

Securely transfer files over SFTP, FTPS, and FTP directly into and out of S3.

Data Ingestion and Transformation Fundamentals

Memory trick

Choosing Data Ingestion Services (Batch/Streaming)

BATS (Batch) are for Big, Archived, Timed, Scheduled data. STREAMS are for Swift, Timely, Real-time, Event-driven, Always-on data.

Data Ingestion and Transformation Fundamentals

Exam tip

Choosing Data Ingestion Services (Batch/Streaming)

The exam often presents scenarios focusing on latency requirements. If the scenario mentions 'real-time,' 'immediate insights,' or 'seconds/milliseconds,' think streaming services like Kinesis. If it mentions 'daily reports,' 'nightly jobs,' or 'periodic analysis,' think batch services like Glue or DMS.

Data Ingestion and Transformation Fundamentals

Common mistake

Choosing Data Ingestion Services (Batch/Streaming)

Using streaming services for non-real-time batch data, leading to unnecessary cost and complexity.

Data Ingestion and Transformation Fundamentals

Common mistake

Choosing Data Ingestion Services (Batch/Streaming)

Attempting to use batch services for real-time analytics, resulting in unacceptable data latency.

Data Ingestion and Transformation Fundamentals

Common mistake

Choosing Data Ingestion Services (Batch/Streaming)

Not considering data volume and velocity when choosing a service, leading to scalability issues or over-provisioning.

Data Ingestion and Transformation Fundamentals

Key term

Kinesis Data Streams

Service for capturing, storing, and processing real-time data streams.

Data Ingestion and Transformation Fundamentals

Key term

Kinesis Firehose

Fully managed service for delivering streaming data to destinations.

Data Ingestion and Transformation Fundamentals

Key term

Shard

A base throughput unit in Kinesis Data Streams.

Data Ingestion and Transformation Fundamentals

Key term

Change Data Capture (CDC)

Process of identifying and capturing changes in a database.

Data Ingestion and Transformation Fundamentals

Key term

Replication Instance

EC2 instance that performs data migration tasks in DMS.

Data Ingestion and Transformation Fundamentals

Key term

Amazon S3

Object storage service, often used as a data lake landing zone.

Data Ingestion and Transformation Fundamentals

Memory trick

Implementing Data Ingestion Solutions

Kinesis is for Kicking off streams. DMS is for Database Migrations. S3 is for Storing everything.

Data Ingestion and Transformation Fundamentals

Exam tip

Implementing Data Ingestion Solutions

The exam often tests your ability to distinguish between Kinesis Data Streams (for custom applications and multiple consumers) and Kinesis Firehose (for direct delivery to specific destinations with less setup). Pay attention to 'real-time processing' vs. 'delivery to S3/Redshift'.

Data Ingestion and Transformation Fundamentals

Common mistake

Implementing Data Ingestion Solutions

Confusing Kinesis Data Streams with Kinesis Firehose: KDS is for custom, real-time processing by applications, while Firehose is for managed delivery to specific destinations.

Data Ingestion and Transformation Fundamentals

Common mistake

Implementing Data Ingestion Solutions

Underestimating S3's role in ingestion: While not a streaming service, S3 is a primary target for many ingestion services and a critical staging area.

Data Ingestion and Transformation Fundamentals

Common mistake

Implementing Data Ingestion Solutions

Using DMS for general file transfer: DMS is specifically for database migration and replication, not for moving arbitrary files.

Data Ingestion and Transformation Fundamentals

Key term

Amazon EMR

Managed cluster platform for big data processing with open-source frameworks.

Data Ingestion and Transformation Fundamentals

Key term

AWS Lambda

Serverless compute service for event-driven, short-lived code execution.

Data Ingestion and Transformation Fundamentals

Key term

Data Catalog

Metadata repository for data assets, often used with AWS Glue.

Data Ingestion and Transformation Fundamentals

Key term

ETL

Extract, Transform, Load: a data integration process.

Data Ingestion and Transformation Fundamentals

Key term

Serverless

Cloud execution model where the provider manages servers.

Data Ingestion and Transformation Fundamentals

Key term

Spark

Open-source distributed processing system for big data.

Data Ingestion and Transformation Fundamentals

Memory trick

Selecting Data Transformation Services

Think of the services as tools in a kitchen: Lambda is a sharp paring knife for quick, precise cuts. Glue is a powerful food processor for large batches. EMR is a full-blown commercial kitchen with specialized equipment for complex, custom recipes.

Data Ingestion and Transformation Fundamentals

Exam tip

Selecting Data Transformation Services

The exam often presents scenarios requiring you to choose the 'most appropriate' or 'most cost-effective' service. Look for keywords like 'batch processing,' 'real-time,' 'petabytes,' 'serverless,' 'open-source frameworks,' and 'event-driven' to guide your selection.

Data Ingestion and Transformation Fundamentals

Common mistake

Selecting Data Transformation Services

Using Lambda for large-scale, long-running ETL jobs, leading to execution timeouts and increased costs.

Data Ingestion and Transformation Fundamentals

Common mistake

Selecting Data Transformation Services

Choosing EMR when a simpler, serverless option like Glue would suffice, resulting in unnecessary operational overhead and cost.

Data Ingestion and Transformation Fundamentals

Common mistake

Selecting Data Transformation Services

Not leveraging the AWS Glue Data Catalog when using Glue, missing out on centralized metadata management.

Data Ingestion and Transformation Fundamentals

Key term

ELT

Extract, Load, Transform; raw data loaded, then transformed in target.

Data Ingestion and Transformation Fundamentals

Key term

Data Cleaning

Process of fixing or removing incorrect, corrupted, or incomplete data.

Data Ingestion and Transformation Fundamentals

Key term

Data Enrichment

Adding value to data by combining it with other relevant datasets.

Data Ingestion and Transformation Fundamentals

Key term

Data Aggregation

Summarizing data, often by grouping and applying functions like sum or average.

Data Ingestion and Transformation Fundamentals

Key term

AWS EMR

Managed Hadoop framework for processing large datasets with Spark, Hive, etc.

Data Ingestion and Transformation Fundamentals

Key term

Data Lake

Centralized repository storing all data, structured and unstructured, at any scale.

Data Ingestion and Transformation Fundamentals

Memory trick

Implementing Data Transformation Solutions (ETL/ELT)

Think of ETL as a chef: ingredients (data) are prepped and cooked (transformed) in the kitchen (staging area) before being served (loaded). ELT is like a buffet: all raw ingredients are put out first (loaded), and then guests (analysts) pick and prepare (transform) what they want at their table.

Data Ingestion and Transformation Fundamentals

Exam tip

Implementing Data Transformation Solutions (ETL/ELT)

The exam often tests your ability to differentiate between ETL and ELT scenarios. Look for keywords like 'pre-processed,' 'on-premises data warehouse,' or 'strict schema' for ETL. For ELT, look for 'raw data lake,' 'schema-on-read,' 'flexible transformations,' or 'cloud-native scalability.'

Data Ingestion and Transformation Fundamentals

Common mistake

Implementing Data Transformation Solutions (ETL/ELT)

Confusing ETL and ELT: Remember the order of 'T' and 'L' is the key difference.

Data Ingestion and Transformation Fundamentals

Common mistake

Implementing Data Transformation Solutions (ETL/ELT)

Assuming one approach is always superior: The best choice depends heavily on the specific use case, data volume, and target system.

Data Ingestion and Transformation Fundamentals

Common mistake

Implementing Data Transformation Solutions (ETL/ELT)

Underestimating the complexity of transformations: Data cleaning and standardization are often the most time-consuming parts of any data pipeline.

Data Ingestion and Transformation Fundamentals

Key term

Schema Evolution

Ability of a data system to adapt to changes in data structure over time.

Data Ingestion and Transformation Fundamentals

Key term

Data Quality Framework

Set of rules and processes to ensure data accuracy, consistency, and completeness.

Data Ingestion and Transformation Fundamentals

Key term

Data Masking

Replacing sensitive data with realistic, fictitious data for security.

Data Ingestion and Transformation Fundamentals

Key term

Tokenization

Replacing sensitive data with a non-sensitive, random equivalent (token).

Data Ingestion and Transformation Fundamentals

Key term

Predicate Pushdown

Optimization where filters are applied early in query processing, reducing data scanned.

Data Ingestion and Transformation Fundamentals

Key term

Columnar Storage

Data storage format where data is stored by columns, optimizing analytical queries.

Data Ingestion and Transformation Fundamentals

Memory trick

Advanced Data Transformation Patterns & Tools

CDC: 'C'atch 'D'ata 'C'hanges. DMS is your 'D'ata 'M'ovement 'S'ervice for this!

Data Ingestion and Transformation Fundamentals

Exam tip

Advanced Data Transformation Patterns & Tools

Memorize the core capabilities of AWS DMS for CDC, including supported sources and targets. Understand how AWS Glue Data Catalog facilitates schema evolution and data quality management. Be prepared for questions on optimizing Glue/EMR jobs using partitioning and file formats.

Data Ingestion and Transformation Fundamentals

Common mistake

Advanced Data Transformation Patterns & Tools

Ignoring schema evolution, leading to broken pipelines when source schemas change.

Data Ingestion and Transformation Fundamentals

Common mistake

Advanced Data Transformation Patterns & Tools

Not implementing data quality checks, resulting in unreliable analytics and insights.

Data Ingestion and Transformation Fundamentals

Common mistake

Advanced Data Transformation Patterns & Tools

Performing full data loads instead of CDC for incremental updates, wasting resources and time.

Data Ingestion and Transformation Fundamentals

Key term

Amazon RDS

Managed relational database service for structured, transactional data.

Data Storage and Management Strategies

Key term

Amazon DynamoDB

Fully managed NoSQL database for high-performance, low-latency apps.

Data Storage and Management Strategies

Key term

Object Storage

Stores data as objects, highly scalable, ideal for unstructured data.

Data Storage and Management Strategies

Key term

Relational Database

Stores structured data in tables with predefined schemas, ACID compliant.

Data Storage and Management Strategies

Key term

NoSQL Database

Non-relational database, flexible schema, high scalability and performance.

Data Storage and Management Strategies

Key term

Durability

The likelihood of data remaining intact and uncorrupted over time.

Data Storage and Management Strategies

Memory trick

Identifying and Choosing Data Storage Services

Remember 'S.R.D.' for Storage, Relational, Dynamo. S3 for Storage of anything; RDS for Relational, structured data; DynamoDB for Dynamic, high-speed NoSQL.

Data Storage and Management Strategies

Exam tip

Identifying and Choosing Data Storage Services

The exam often presents scenarios and asks you to choose the MOST appropriate storage service. Pay close attention to keywords like 'unstructured data,' 'data lake,' 'backups' (S3); 'transactional,' 'ACID,' 'joins,' 'complex queries' (RDS); and 'low latency,' 'high throughput,' 'key-value,' 'serverless' (DynamoDB).

Data Storage and Management Strategies

Common mistake

Identifying and Choosing Data Storage Services

Using RDS for petabytes of unstructured log data, leading to massive costs and performance issues.

Data Storage and Management Strategies

Common mistake

Identifying and Choosing Data Storage Services

Choosing DynamoDB for complex analytical queries requiring multi-table joins, as its query capabilities are limited.

Data Storage and Management Strategies

Common mistake

Identifying and Choosing Data Storage Services

Storing frequently accessed, small transactional data in S3, which is not optimized for low-latency, high-volume reads/writes of small objects.

Data Storage and Management Strategies

Key term

Amazon Redshift

A fully managed, petabyte-scale cloud data warehouse for analytical workloads.

Data Storage and Management Strategies

Key term

Massively Parallel Processing (MPP)

Architecture distributing data and query processing across multiple computing nodes.

Data Storage and Management Strategies

Key term

Amazon OpenSearch Service

Fully managed service for deploying, operating, and scaling OpenSearch clusters.

Data Storage and Management Strategies

Key term

Full-Text Search

A technique for searching a single computer-stored document or a collection in a database.

Data Storage and Management Strategies

Key term

Kibana

An open-source data visualization dashboard for OpenSearch and Elasticsearch.

Data Storage and Management Strategies

Key term

Data Distribution Style

How Redshift distributes data across compute nodes (EVEN, ALL, KEY).

Data Storage and Management Strategies

Memory trick

Implementing Data Warehousing and Search Solutions

Redshift is for 'Reports' (R) and 'Deep' analysis (D). OpenSearch is for 'Operational' (O) insights and 'Quick' searches (Q).

Data Storage and Management Strategies

Exam tip

Implementing Data Warehousing and Search Solutions

The exam often presents scenarios requiring you to choose the most cost-effective and performant AWS service. For Redshift, keywords like 'historical analysis,' 'complex SQL queries,' 'business intelligence,' and 'petabyte-scale data warehousing' are strong indicators. For OpenSearch, look for 'real-time search,' 'log analytics,' 'operational intelligence,' 'full-text search,' and 'unstructured/semi-structured data.'

Data Storage and Management Strategies

Common mistake

Implementing Data Warehousing and Search Solutions

Using Redshift for real-time, low-latency lookups on individual records, which is better suited for services like DynamoDB or RDS.

Data Storage and Management Strategies

Common mistake

Implementing Data Warehousing and Search Solutions

Using OpenSearch for complex, multi-table joins and aggregations typical of traditional data warehousing, where Redshift excels.

Data Storage and Management Strategies

Common mistake

Implementing Data Warehousing and Search Solutions

Underestimating the importance of data distribution and sort keys in Redshift, or shard allocation in OpenSearch, leading to poor query performance.

Data Storage and Management Strategies

Key term

Partitioning

Dividing data into smaller, logical segments for performance.

Data Storage and Management Strategies

Key term

Compression

Reducing data size to save storage and speed up transfer.

Data Storage and Management Strategies

Key term

Columnar Format

Storing data column by column, optimizing analytical queries.

Data Storage and Management Strategies

Key term

Parquet

A popular columnar storage file format for big data.

Data Storage and Management Strategies

Key term

ORC

Optimized Row Columnar, another efficient columnar format.

Data Storage and Management Strategies

Key term

GZIP

A widely used compression algorithm with high compression ratio.

Data Storage and Management Strategies

Key term

Snappy

A fast compression/decompression algorithm, good for analytics.

Data Storage and Management Strategies

Key term

Distribution Key

In Redshift, determines how data is spread across nodes.

Data Storage and Management Strategies

Memory trick

Data Partitioning, Compression, and Optimization

P.C.F. for Performance, Cost, Flexibility! Partition, Compress, and choose the right File format to optimize your data.

Data Storage and Management Strategies

Exam tip

Data Partitioning, Compression, and Optimization

The exam often tests your understanding of how partitioning, compression, and file formats impact query performance and cost. Look for keywords like 'reduce scan,' 'lower costs,' or 'improve analytics speed.' Memorize that Parquet and ORC are columnar, and GZIP/Snappy are common compression types.

Data Storage and Management Strategies

Common mistake

Data Partitioning, Compression, and Optimization

Not partitioning data, leading to full table/object scans and high costs.

Data Storage and Management Strategies

Common mistake

Data Partitioning, Compression, and Optimization

Using row-oriented formats (CSV, JSON) for analytical workloads instead of columnar formats.

Data Storage and Management Strategies

Common mistake

Data Partitioning, Compression, and Optimization

Choosing a compression algorithm that is too slow for the required decompression speed.

Data Storage and Management Strategies

Common mistake

Data Partitioning, Compression, and Optimization

Over-partitioning data, which can lead to a large number of small files and overhead.

Data Storage and Management Strategies

Key term

AWS Glue Catalog

A persistent, serverless metadata store for AWS data assets.

Data Storage and Management Strategies

Key term

Metadata

Data about data, such as schema, data types, and location.

Data Storage and Management Strategies

Key term

Crawler

An AWS Glue component that scans data stores to infer schemas.

Data Storage and Management Strategies

Key term

Database (Glue)

A logical grouping of tables within the AWS Glue Catalog.

Data Storage and Management Strategies

Key term

Table (Glue)

A metadata object in Glue Catalog defining a dataset's schema.

Data Storage and Management Strategies

Key term

Classifier

Helps Glue Crawlers understand specific data formats.

Data Storage and Management Strategies

Memory trick

Data Cataloging & Metadata Management (AWS Glue Catalog)

Imagine a GLUE stick holding together all your data's 'CLUES' (Catalog, Location, Understanding, Everything) in one central place!

Data Storage and Management Strategies

Exam tip

Data Cataloging & Metadata Management (AWS Glue Catalog)

The exam often tests your understanding of Glue Catalog's role as a *central metadata repository* and its *integration points* with services like Athena, Redshift Spectrum, and EMR. Memorize that Glue Catalog stores *table definitions, schema, and physical location*.

Data Storage and Management Strategies

Common mistake

Data Cataloging & Metadata Management (AWS Glue Catalog)

Confusing Glue Catalog with a data storage service; it only stores metadata, not the actual data.

Data Storage and Management Strategies

Common mistake

Data Cataloging & Metadata Management (AWS Glue Catalog)

Underestimating the importance of crawlers for automated schema discovery and keeping the catalog up-to-date.

Data Storage and Management Strategies

Common mistake

Data Cataloging & Metadata Management (AWS Glue Catalog)

Not realizing that Glue Catalog is used by many other AWS analytics services, not just Glue ETL.

Data Storage and Management Strategies

Key term

IAM

AWS Identity and Access Management; controls access to AWS services and resources.

Data Governance and Security Best Practices

Key term

IAM Policy

JSON document defining permissions for AWS resources and actions.

Data Governance and Security Best Practices

Key term

IAM Role

Identity that grants temporary permissions, assumed by users or services.

Data Governance and Security Best Practices

Key term

AWS Lake Formation

Service for building, securing, and managing data lakes with fine-grained access.

Data Governance and Security Best Practices

Key term

Glue Data Catalog

Central metadata repository for data lakes, used by Lake Formation.

Data Governance and Security Best Practices

Key term

Principle of Least Privilege

Security best practice: grant only minimum necessary permissions.

Data Governance and Security Best Practices

Key term

Fine-grained Access Control

Permissions at granular levels like table, column, row, or cell.

Data Governance and Security Best Practices

Memory trick

Implementing Data Access Control and Authentication

IAM is the 'bouncer' for the AWS club (services), Lake Formation is the 'bouncer' for the VIP room (data tables).

Data Governance and Security Best Practices

Exam tip

Implementing Data Access Control and Authentication

The exam often tests the distinction between IAM and Lake Formation. Remember: IAM controls access to AWS services themselves, while Lake Formation controls access to the data within the Glue Data Catalog, often at a more granular level.

Data Governance and Security Best Practices

Common mistake

Implementing Data Access Control and Authentication

Over-privileging IAM users or roles, granting more permissions than necessary.

Data Governance and Security Best Practices

Common mistake

Implementing Data Access Control and Authentication

Confusing IAM policies with Lake Formation permissions; they serve different layers of control.

Data Governance and Security Best Practices

Common mistake

Implementing Data Access Control and Authentication

Not using IAM roles for services, instead embedding credentials or using IAM users directly.

Data Governance and Security Best Practices

Key term

Data Encryption

Transforming data into an unreadable format to protect confidentiality, reversible with a key.

Data Governance and Security Best Practices

Key term

Data at Rest

Data stored persistently in storage devices or databases.

Data Governance and Security Best Practices

Key term

Data in Transit

Data actively moving over a network connection between systems.

Data Governance and Security Best Practices

Key term

AWS KMS

AWS Key Management Service, for creating and managing cryptographic keys.

Data Governance and Security Best Practices

Key term

Server-Side Encryption (SSE)

Encryption performed by the service receiving the data (e.g., S3, RDS).

Data Governance and Security Best Practices

Key term

Client-Side Encryption (CSE)

Encryption performed by the client application before sending data to a service.

Data Governance and Security Best Practices

Key term

TLS (Transport Layer Security)

Cryptographic protocol ensuring secure communication over a computer network.

Data Governance and Security Best Practices

Memory trick

Implementing Data Encryption and Data Masking

KMS: Keep My Secrets Secure. Remember KMS is your central hub for managing encryption keys across AWS.

Data Governance and Security Best Practices

Exam tip

Implementing Data Encryption and Data Masking

The exam frequently tests on the different types of S3 Server-Side Encryption (SSE-S3, SSE-KMS, SSE-C) and when to use each. Memorize their key management differences.

Data Governance and Security Best Practices

Common mistake

Implementing Data Encryption and Data Masking

Forgetting to encrypt data in transit, leaving it vulnerable during network transfers.

Data Governance and Security Best Practices

Common mistake

Implementing Data Encryption and Data Masking

Using masked data in production environments, which defeats the purpose of masking and can lead to data integrity issues.

Data Governance and Security Best Practices

Common mistake

Implementing Data Encryption and Data Masking

Not properly managing KMS key access, potentially allowing unauthorized decryption or preventing authorized access.

Data Governance and Security Best Practices

Key term

AWS CloudTrail

Service that records API calls and events for governance, compliance, and operational auditing.

Data Governance and Security Best Practices

Key term

Amazon CloudWatch

Monitoring and observability service that collects logs, metrics, and events.

Data Governance and Security Best Practices

Key term

Management Events

CloudTrail events that record control plane operations (e.g., creating resources).

Data Governance and Security Best Practices

Key term

Data Events

CloudTrail events that record data plane operations (e.g., S3 object access, DynamoDB item changes).

Data Governance and Security Best Practices

Key term

CloudWatch Logs

Component of CloudWatch for centralizing, monitoring, and storing logs from various sources.

Data Governance and Security Best Practices

Key term

CloudWatch Metrics

Time-ordered set of data points published to CloudWatch for monitoring.

Data Governance and Security Best Practices

Key term

CloudWatch Alarms

Automatically performs actions based on a metric exceeding a threshold.

Data Governance and Security Best Practices

Key term

Log File Integrity

CloudTrail feature that verifies log files have not been tampered with.

Data Governance and Security Best Practices

Memory trick

Implementing Data Auditing and Logging (CloudTrail, CloudWatch)

Imagine a 'Trail' of breadcrumbs (CloudTrail) showing every step taken in your AWS account, and a 'Watch' (CloudWatch) that constantly checks those breadcrumbs for anything unusual, sounding an 'Alarm' if it sees a problem!

Data Governance and Security Best Practices

Exam tip

Implementing Data Auditing and Logging (CloudTrail, CloudWatch)

The exam often tests your understanding of which service logs *API calls* (CloudTrail) versus which service *monitors and centralizes logs* (CloudWatch Logs) and *creates alarms* (CloudWatch Alarms). Remember CloudTrail for 'who did what' and CloudWatch for 'how is it performing and alert me'.

Data Governance and Security Best Practices

Common mistake

Implementing Data Auditing and Logging (CloudTrail, CloudWatch)

Not enabling CloudTrail data events for critical data services like S3 or DynamoDB, leading to blind spots.

Data Governance and Security Best Practices

Common mistake

Implementing Data Auditing and Logging (CloudTrail, CloudWatch)

Failing to secure CloudTrail S3 buckets, making logs vulnerable to tampering or deletion.

Data Governance and Security Best Practices

Common mistake

Implementing Data Auditing and Logging (CloudTrail, CloudWatch)

Ignoring CloudWatch alarms, or not setting up appropriate alarms for critical security events.

Data Governance and Security Best Practices

Common mistake

Implementing Data Auditing and Logging (CloudTrail, CloudWatch)

Not implementing proper log retention policies, leading to either excessive storage costs or non-compliance.

Data Governance and Security Best Practices

Key term

Data Retention Policy

Rules defining how long data must be kept and how it's disposed.

Data Governance and Security Best Practices

Key term

S3 Lifecycle Policy

Automates object transitions between S3 storage classes and object expiration.

Data Governance and Security Best Practices

Key term

S3 Glacier

Low-cost Amazon S3 storage class for archiving data with retrieval times from minutes to hours.

Data Governance and Security Best Practices

Key term

S3 Glacier Deep Archive

Lowest-cost Amazon S3 storage class for long-term archiving with retrieval times within 12 hours.

Data Governance and Security Best Practices

Key term

Data Expiration

The process of permanently deleting data after its retention period ends.

Data Governance and Security Best Practices

Key term

AWS Backup

Centralized service to manage backups across multiple AWS services.

Data Governance and Security Best Practices

Memory trick

Implementing Data Retention and Lifecycle Management

Think of S3 Lifecycle policies as a 'Stairway to Heaven (and then deletion)' for your data: Standard -> IA -> Glacier -> Deep Archive -> Delete. Each step down is cheaper but slower to retrieve.

Data Governance and Security Best Practices

Exam tip

Implementing Data Retention and Lifecycle Management

For S3 Lifecycle policies, remember that the minimum storage duration for S3 Standard-IA and S3 One Zone-IA is 30 days. For S3 Glacier and S3 Glacier Deep Archive, the minimum billable storage duration is 90 days and 180 days, respectively. Deleting objects before these minimums incurs a pro-rated charge.

Data Governance and Security Best Practices

Common mistake

Implementing Data Retention and Lifecycle Management

Forgetting to account for minimum storage durations when setting up S3 Lifecycle policies, leading to unexpected costs.

Data Governance and Security Best Practices

Common mistake

Implementing Data Retention and Lifecycle Management

Not periodically reviewing and updating retention policies as business or regulatory requirements change.

Data Governance and Security Best Practices

Common mistake

Implementing Data Retention and Lifecycle Management

Assuming all data in a data lake has the same retention requirements; different layers (raw, refined) often have different rules.

Data Governance and Security Best Practices

Key term

Orchestration

Coordinating multiple steps in a workflow to achieve a goal.

Data Operations and Monitoring Excellence

Key term

Scheduling

Automating the execution of tasks or workflows at specific times.

Data Operations and Monitoring Excellence

Key term

AWS Step Functions

Serverless workflow service for orchestrating AWS services.

Data Operations and Monitoring Excellence

Key term

State Machine

A visual representation of a workflow in Step Functions.

Data Operations and Monitoring Excellence

Key term

Apache Airflow

Open-source platform to programmatically author, schedule, monitor workflows.

Data Operations and Monitoring Excellence

Key term

MWAA

Amazon Managed Workflows for Apache Airflow; managed Airflow service.

Data Operations and Monitoring Excellence

Key term

DAG (Directed Acyclic Graph)

A collection of tasks with dependencies, defining an Airflow workflow.

Data Operations and Monitoring Excellence

Memory trick

Implementing Data Pipeline Orchestration and Scheduling

STEP up for Serverless, MWAA for Managed Airflow Automation.

Data Operations and Monitoring Excellence

Exam tip

Implementing Data Pipeline Orchestration and Scheduling

For the DEA-C01 exam, distinguish between Step Functions and MWAA. Step Functions is serverless, visual, and event-driven, great for coordinating AWS services. MWAA is a managed Apache Airflow, code-driven (DAGs), best for complex batch scheduling or Airflow migrations.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Pipeline Orchestration and Scheduling

Using Step Functions for extremely complex, long-running batch ETLs that are better suited for MWAA's native Airflow capabilities.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Pipeline Orchestration and Scheduling

Trying to manage an Apache Airflow cluster manually on EC2 when MWAA offers a fully managed solution.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Pipeline Orchestration and Scheduling

Not considering the cost implications: Step Functions charges per state transition, while MWAA charges for environment uptime and resources.

Data Operations and Monitoring Excellence

Key term

Data Quality Dimensions

Characteristics like accuracy, completeness, consistency, timeliness, and validity.

Data Operations and Monitoring Excellence

Key term

AWS Glue Data Quality

Service for defining and evaluating data quality rules on Glue Data Catalog tables.

Data Operations and Monitoring Excellence

Key term

Amazon SNS

A messaging service used for sending notifications to multiple subscribers.

Data Operations and Monitoring Excellence

Key term

Data Lineage

The lifecycle of data, including its origin, transformations, and destinations.

Data Operations and Monitoring Excellence

Key term

Data Profiling

Analyzing data to discover its structure, content, and quality.

Data Operations and Monitoring Excellence

Key term

Data Governance

Overall management of data availability, usability, integrity, and security.

Data Operations and Monitoring Excellence

Memory trick

Implementing Data Quality Monitoring and Alerting

To remember the Data Quality Dimensions: A C C T V. Think 'A C C T V' for 'Accurate, Complete, Consistent, Timely, Valid' data, like a security camera watching over your data.

Data Operations and Monitoring Excellence

Exam tip

Implementing Data Quality Monitoring and Alerting

The exam often asks about specific AWS services for data quality. Memorize that AWS Glue Data Quality is the primary service for defining and evaluating rules directly on your data catalog tables. Look for keywords like 'data quality rules,' 'data catalog integration,' or 'data quality scores.'

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Quality Monitoring and Alerting

Not defining clear data quality metrics and thresholds before implementing monitoring.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Quality Monitoring and Alerting

Failing to integrate data quality checks early in the data pipeline, allowing bad data to propagate.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Quality Monitoring and Alerting

Setting up alerts without a clear plan for who receives them and how issues will be remediated.

Data Operations and Monitoring Excellence

Key term

CloudWatch

AWS service for monitoring resources and applications.

Data Operations and Monitoring Excellence

Key term

CloudTrail

AWS service for logging API calls and events.

Data Operations and Monitoring Excellence

Key term

Dead-Letter Queue (DLQ)

Queue for messages that failed processing.

Data Operations and Monitoring Excellence

Key term

Idempotency

Operation yielding same result if executed multiple times.

Data Operations and Monitoring Excellence

Key term

Observability

Ability to infer internal state from external outputs.

Data Operations and Monitoring Excellence

Key term

Retry Logic

Mechanism to reattempt failed operations.

Data Operations and Monitoring Excellence

Key term

Circuit Breaker

Pattern to prevent cascading failures in distributed systems.

Data Operations and Monitoring Excellence

Key term

SNS

Amazon Simple Notification Service for sending messages.

Data Operations and Monitoring Excellence

Memory trick

Implementing Data Pipeline Troubleshooting and Error Handling

To remember troubleshooting steps: 'M-D-I-D-I-T-R' – Monitor, Detect, Investigate, Diagnose, Implement, Test, Recover. It's like a doctor's visit for your data!

Data Operations and Monitoring Excellence

Exam tip

Implementing Data Pipeline Troubleshooting and Error Handling

The exam often tests your knowledge of specific AWS services for monitoring and logging. Memorize that CloudWatch is for metrics and logs, CloudTrail for API calls and auditing, and S3 access logs for S3 bucket activity. Understand when to use DLQs for error handling.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Pipeline Troubleshooting and Error Handling

Not implementing comprehensive logging at all stages of the pipeline, making root cause analysis difficult.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Pipeline Troubleshooting and Error Handling

Ignoring the importance of dead-letter queues (DLQs) for failed messages, leading to data loss or stalled pipelines.

Data Operations and Monitoring Excellence

Common mistake

Implementing Data Pipeline Troubleshooting and Error Handling

Failing to set up proactive alarms and notifications, resulting in delayed detection of critical issues.

Data Operations and Monitoring Excellence

Key term

S3 Lifecycle Policies

Rules to automate object transitions and expirations in S3.

Data Operations and Monitoring Excellence

Key term

Intelligent-Tiering

S3 storage class that automatically moves data to cost-effective tiers.

Data Operations and Monitoring Excellence

Key term

Spot Instances

Unused EC2 capacity available at a discount, suitable for fault-tolerant workloads.

Data Operations and Monitoring Excellence

Key term

AWS Cost Explorer

A tool to visualize, understand, and manage AWS costs and usage over time.

Data Operations and Monitoring Excellence

Key term

AWS Budgets

Service to set custom budgets and receive alerts when costs exceed thresholds.

Data Operations and Monitoring Excellence

Key term

VPC Endpoints

Private connection to AWS services from within a VPC, reducing data transfer costs.

Data Operations and Monitoring Excellence

Key term

Data Transfer Costs

Charges for moving data between AWS services, regions, or to the internet.

Data Operations and Monitoring Excellence

Memory trick

Cost Optimization for Data Solutions

COST: C - Compute (right-size), O - Object Storage (lifecycle), S - Spot Instances (save money), T - Transfer (minimize egress).

Data Operations and Monitoring Excellence

Exam tip

Cost Optimization for Data Solutions

The exam often tests knowledge of S3 storage classes and their cost implications. Memorize the typical use cases and cost differences between S3 Standard, S3 Standard-IA, S3 One Zone-IA, S3 Glacier, and S3 Glacier Deep Archive. Also, understand when to use lifecycle policies and Intelligent-Tiering.

Data Operations and Monitoring Excellence

Common mistake

Cost Optimization for Data Solutions

Leaving idle resources running (e.g., EMR clusters, RDS instances) when not in use.

Data Operations and Monitoring Excellence

Common mistake

Cost Optimization for Data Solutions

Storing infrequently accessed data in expensive storage classes like S3 Standard.

Data Operations and Monitoring Excellence

Common mistake

Cost Optimization for Data Solutions

Ignoring data transfer costs, especially for cross-region or internet egress.

Data Operations and Monitoring Excellence

Common mistake

Cost Optimization for Data Solutions

Not using tagging, making it difficult to attribute costs to specific teams or projects.

Data Operations and Monitoring Excellence