AWS Certified Data Engineer – Associate practice questions
214 free questions with answers and explanations.
- 201.A global e-commerce company needs to process customer orders in near real-time. Each order record, averaging 1 KB in size, needs to be ingested, transformed to a standardized format, and then loaded into a data warehouse for immediate business intelligence reporting. The company anticipates peak traffic of 5,000 orders per second. Which AWS data ingestion and transformation services should be combined to meet these requirements most efficiently?Data Ingestion and Transformation
- 202.A global IoT company collects sensor data from millions of devices worldwide. The data is high-volume, low-latency, and needs to be ingested, filtered, and routed to various AWS services based on rules (e.g., critical alerts to Amazon SNS, telemetry data to Amazon S3, operational metrics to Amazon Kinesis Data Streams). The solution must be fully managed, highly scalable, and support complex rule-based processing without requiring server management. Which AWS service is the MOST appropriate for this ingestion and initial routing?Data Ingestion and Transformation
- 203.A cybersecurity firm needs to ingest security logs from various on-premises firewalls and intrusion detection systems (IDS). These logs are generated continuously and need to be analyzed in near real-time for threat detection. The firm requires a solution that can reliably collect data from on-premises, encrypt it in transit, and deliver it to an Amazon S3 bucket for long-term storage, and also to an Amazon OpenSearch Service domain for real-time analysis. The solution should be simple to deploy and manage. Which AWS service is the MOST appropriate for ingesting this data?Data Ingestion and Transformation
- 204.A financial institution needs to audit all access to sensitive customer data stored in an Amazon S3 bucket. All read and write operations on this S3 bucket must be captured, and these audit logs need to be delivered to an Amazon Redshift data warehouse for long-term retention and compliance analysis. The solution must be automated, reliable, and ensure that no audit events are missed. Which combination of AWS services should be used to achieve this?Data Ingestion and Transformation
- 205.A retail company needs to analyze customer purchase data to identify trends and personalize recommendations. The data is currently stored in a relational database (Amazon RDS for PostgreSQL) and needs to be extracted, transformed, and loaded into an Amazon Redshift data warehouse on a daily basis. The transformation involves aggregating sales, joining with customer demographics, and cleaning inconsistent product descriptions. The solution must be fully managed, scalable, and support complex SQL-based transformations. Which AWS service is the MOST appropriate for this transformation?Data Ingestion and Transformation
- 206.A financial institution needs to process large volumes of historical market data, approximately 500 TB, currently stored on-premises across various network-attached storage (NAS) devices. The data needs to be moved to Amazon S3 for long-term archival and eventual analysis. The transfer must be highly secure, reliable, and optimized for large datasets over a standard internet connection, minimizing manual effort. Which AWS service is the MOST suitable for this one-time bulk migration?Data Ingestion and Transformation
- 207.A global media company receives large video files (up to 50 GB each) from content creators located worldwide. These files arrive daily and need to be securely ingested into Amazon S3 for archival and subsequent processing by AWS Media Services. The content creators prefer using standard file transfer protocols. The ingestion solution must be highly available and eliminate the need for managing compute instances or complex network configurations. Which combination of AWS services should the company use?Data Ingestion and Transformation
- 208.A global manufacturing company needs to ingest real-time operational metrics from thousands of factory floor sensors. These sensors generate high-volume, low-latency data streams that must be processed for immediate anomaly detection and stored for historical analysis. The solution must be fully managed and scalable to handle millions of data points per second. Which AWS service is the MOST appropriate for ingesting this data?Data Ingestion and Transformation
- 209.A research institution collects high-resolution satellite imagery, with individual image files often exceeding 100 GB. The data is generated in remote locations with intermittent and limited internet connectivity. These files must eventually be stored in Amazon S3 for long-term archival and processing. The solution must ensure data integrity, be physically secure during transit, and minimize reliance on unreliable network connections. Which AWS service is the MOST appropriate for ingesting this data?Data Ingestion and Transformation
- 210.A data engineering team is designing a data lake on AWS. They need to store streaming data from various IoT devices, which arrives at a high velocity and requires immediate processing for real-time analytics. The solution must be highly scalable and durable. Which AWS service is most appropriate for ingesting this data?Data Storage and Management
- 211.A data architect is designing a data lake on Amazon S3 for an e-commerce company. The data includes customer orders, product catalogs, and website clickstreams. To optimize query performance and reduce scanning costs for analytical queries, the architect wants to store the data in a columnar format that supports predicate pushdown and efficient compression. Which file format should be recommended?Data Storage and Management
- 212.A data engineering team manages a large data lake on Amazon S3. They use AWS Glue Data Catalog to store metadata for querying with Amazon Athena. Occasionally, new partitions are added to S3 by an external process, but these new partitions are not immediately discoverable by Athena queries. The team needs a mechanism to automatically update the Glue Data Catalog with these new partitions without manual intervention. Which solution should they implement?Data Storage and Management
- 213.A data engineering team is designing a data lake on AWS. They need to store streaming data from various IoT devices for real-time analytics and subsequent batch processing. The solution must be fully managed, scalable, and capable of delivering data to Amazon S3, Amazon Redshift, and Amazon OpenSearch Service without requiring custom application development. Which AWS service is best suited for this requirement?Data Storage and Management
- 214.A data engineering team is building a serverless data processing pipeline using AWS Lambda. The Lambda functions need to access a shared, persistent file system to store intermediate processing results and configuration files that are too large to fit in Lambda's ephemeral storage. This file system must be accessible concurrently by multiple Lambda invocations and scale automatically. Which AWS storage service should they use?Data Storage and Management