AWS Certified Data Engineer – Associate practice questions

214 free questions with answers and explanations.

Practice test
  1. 51.A data engineering team is responsible for a critical data pipeline that extracts data from an on-premises database, transforms it using AWS Glue, and loads it into an Amazon Redshift data warehouse. The pipeline runs daily. Recently, stakeholders have reported that the reports generated from Redshift are occasionally missing data, but the pipeline status in AWS Glue always shows 'Succeeded'. The team needs to implement a solution to detect these data anomalies proactively and receive notifications when they occur. Which AWS service should the team use to address this requirement efficiently?Data Operations and Monitoring
  2. 52.A retail company collects point-of-sale (POS) data from thousands of stores. Each store sends daily batch files (CSV format) to a central S3 bucket. The data needs to be aggregated by product, anonymized for customer information, and then partitioned by date and store ID into Parquet format for efficient querying in a data lake. The processing should be scheduled nightly and be cost-optimized, paying only for compute capacity used during processing. Which AWS service is most suitable for this batch transformation?Data Ingestion and Transformation
  3. 53.A global e-commerce company wants to centralize customer clickstream data from various regional websites into a single Amazon S3 data lake. The data arrives as high-volume, continuous streams of semi-structured JSON events. Before landing in S3, the company needs to enrich these events by joining them with customer demographic data from a DynamoDB table and filter out sensitive PII. The solution must be serverless, scalable, and minimize operational overhead. Which AWS service combination is best suited for this real-time ingestion and transformation pipeline?Data Ingestion and Transformation
  4. 54.A media company needs to process petabytes of video analytics data stored in Amazon S3. The data is in a variety of complex nested JSON and Parquet formats, and analysts require interactive query capabilities without loading data into a traditional database. The solution must support schema evolution and be cost-effective for ad-hoc queries on large datasets. Which AWS service is most appropriate for transforming and querying this data?Data Ingestion and Transformation
  5. 55.A pharmaceutical company needs to ingest experimental results from laboratory instruments into an Amazon S3 data lake. The data is generated as various file types (CSV, XML, proprietary binary formats) and stored in network shares (SMB) on-premises. The transfer needs to be initiated by lab technicians on demand, requiring a user-friendly interface. Data volumes can vary from small to large files, and secure transfer is paramount. Which AWS service provides a simple, secure, and on-demand way for users to transfer files from on-premises SMB shares to S3?Data Ingestion and Transformation
  6. 56.A global e-commerce company wants to centralize customer clickstream data from various regional websites for real-time analytics and personalized recommendations. The data arrives continuously and needs to be transformed (e.g., anonymize PII, enrich with product catalog data) before being loaded into a data lake in Amazon S3 and a data warehouse in Amazon Redshift. The solution must be fully managed and support both real-time ingestion and transformation. Which AWS service combination is most suitable?Data Ingestion and Transformation
  7. 57.A global logistics company receives shipment tracking updates from thousands of partners in various formats (CSV, JSON, XML) via SFTP. These updates arrive continuously throughout the day, but do not require immediate real-time processing. The company needs to consolidate these files into a centralized Amazon S3 bucket, convert them to Parquet format, and then catalog their schemas in the AWS Glue Data Catalog for subsequent analysis. Which combination of AWS services would be most efficient for this ingestion and initial transformation pipeline?Data Ingestion and Transformation
  8. 58.A healthcare provider receives patient data from various on-premises systems, including electronic health records (EHR) and laboratory information systems (LIS). This data, typically in structured formats like CSV or flat files, needs to be ingested into Amazon S3 daily. The total daily volume is around 5 TB, and the transfer must be reliable, secure, and easily scheduled. The provider wants to minimize operational overhead. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  9. 59.A data analytics company uses AWS Glue to build and manage a data lake on Amazon S3. They have multiple Glue ETL jobs that transform raw data into a refined, queryable format. The company is experiencing high operational costs related to Glue job executions, particularly for jobs that process large datasets daily. The team has already optimized the Spark code within the Glue jobs. They now need to identify and implement a strategy to further reduce Glue ETL costs by minimizing the amount of data processed by each job run, without impacting data freshness or completeness. Which approach should they prioritize?Data Operations and Monitoring
  10. 60.A media company needs to process petabytes of video analytics data stored in Amazon S3. The data consists of billions of small JSON files, each containing metadata and events captured from video streams. Analysts need to run ad-hoc, interactive queries on this data to identify trends and anomalies. The solution must be serverless, cost-effective for intermittent querying, and support standard SQL. Which AWS service should be used?Data Ingestion and Transformation
  11. 61.A financial institution needs to process large volumes of historical transaction data stored in Amazon S3. This data, in Parquet format, needs to be joined with customer master data from a relational database, transformed, and then loaded into a data warehouse for business intelligence reporting. The processing must be serverless, cost-effective, and able to scale dynamically based on data volume. Which AWS service is most suitable for this data transformation task?Data Ingestion and Transformation
  12. 62.A healthcare provider receives patient data from various on-premises systems, including electronic health records (EHR) and medical imaging archives. This data, totaling hundreds of terabytes, needs to be securely and efficiently transferred to an Amazon S3 bucket in a central AWS account. Transfers occur periodically, often overnight, and must resume automatically if interrupted. Which AWS service is the most suitable for this large-scale, periodic data ingestion?Data Ingestion and Transformation
  13. 63.A data engineering team uses AWS Glue ETL jobs to process large datasets daily. Recently, they observed that some Glue jobs are taking significantly longer to complete, occasionally leading to missed SLAs. Upon investigation, they found that the CPU utilization of the Glue job workers is consistently low, while the I/O Wait time is frequently high, indicating that the jobs are bottlenecked by data access. The data is stored in Amazon S3. What is the most effective approach to optimize these Glue jobs by addressing the I/O bottleneck?Data Operations and Monitoring
  14. 64.A global IoT company collects sensor data from millions of devices worldwide. The data is generated continuously, with high velocity and varying schemas depending on the device type. The company needs to transform this raw, semi-structured data into a structured format (e.g., JSON to Parquet) and enrich it with device metadata (e.g., location, ownership) stored in a DynamoDB table. The transformation must occur near real-time, be serverless, and automatically adapt to schema variations without manual intervention. Which AWS service is most appropriate for this real-time, schema-aware transformation?Data Ingestion and Transformation
  15. 65.A startup is building a new mobile application and anticipates millions of concurrent players. The application will generate high-volume, low-latency game event data (e.g., player actions, scores, in-game purchases). This data needs to be captured and made available for real-time analytics and subsequent batch processing. The solution must provide strong ordering guarantees for events within a player session and be highly scalable. Which AWS service is best suited for ingesting this raw event data?Data Ingestion and Transformation
  16. 66.A gaming company collects telemetry data from millions of players, generating petabytes of event data daily. This data is ingested into Amazon Kinesis Data Streams. The company needs to perform real-time aggregations (e.g., calculating average session duration per game, total active players) and detect anomalies in player behavior. The results of these aggregations and anomaly detections should be continuously published to a dashboard for operational monitoring. Which AWS service is best suited for this real-time stream processing and analytics?Data Ingestion and Transformation
  17. 67.A media company has petabytes of historical video content stored in on-premises tape archives. They want to migrate this data to Amazon S3 Glacier Deep Archive for cost-effective long-term storage. Due to the massive volume and the offline nature of the tape archives, a network-based transfer is not feasible. Which AWS service should they use for this one-time, large-scale data ingestion?Data Ingestion and Transformation
  18. 68.A data engineering team operates a critical data pipeline that ingests data into Amazon S3, processes it with AWS Glue, and loads it into Amazon Redshift. The pipeline runs hourly. The team has implemented comprehensive monitoring using AWS CloudWatch for metrics and alarms. To proactively identify and address potential cost overruns, they need to monitor the spending specifically associated with this pipeline and receive alerts if costs exceed predefined thresholds. Which AWS service is best suited for setting up cost-based alerts for this specific pipeline?Data Operations and Monitoring
  19. 69.A data analytics team needs to build a robust and fault-tolerant pipeline to process customer interaction data from various sources (web, mobile, CRM). The data is in various formats (JSON, CSV) and needs to be standardized, de-duplicated, and enriched with geographical information. The transformed data must be stored in a data warehouse (Amazon Redshift) for reporting. The team requires full control over the Spark environment for custom libraries and complex logic, and the ability to scale resources dynamically based on data volume. Which AWS service should they use for the transformation phase?Data Ingestion and Transformation
  20. 70.A marketing agency needs to process customer interaction data from various social media platforms. The data arrives in diverse formats (JSON, XML, CSV) and requires complex transformations, including data parsing, standardization, deduplication, and sentiment analysis using custom Python scripts. The processing needs to be flexible, allowing for custom code execution, and scalable to handle spikes in data volume. The transformed data should be stored in a data lake in Parquet format. Which AWS service provides the most flexibility and control for custom code-driven data transformation at scale?Data Ingestion and Transformation
  21. 71.A data engineering team manages several critical data pipelines using AWS Step Functions to orchestrate various AWS Glue jobs, Lambda functions, and custom containerized tasks. They need a centralized system to monitor the health and performance of these pipelines, detect anomalies, and receive alerts for failures or performance degradation. The solution must provide a single pane of glass for all state machine executions and integrate seamlessly with existing notification mechanisms. Which AWS service combination should they implement?Data Operations and Monitoring
  22. 72.A global logistics company receives shipment tracking updates from thousands of partners in various formats (CSV, JSON, XML) via SFTP. The data needs to be collected, validated, and transformed into a standardized Parquet format for analysis in an S3 data lake. The solution must provide a fully managed SFTP endpoint and integrate seamlessly with serverless ETL services. Which AWS service combination should be used for data ingestion and initial transformation?Data Ingestion and Transformation
  23. 73.A global e-commerce company wants to centralize customer clickstream data from various regional websites. The data arrives as JSON objects, approximately 1KB each, with a peak ingestion rate of 50,000 records per second. The company needs to transform this data by filtering out bot traffic and enriching it with customer demographic information from a DynamoDB table before storing it in an Amazon S3 data lake in Parquet format. The solution must provide near real-time processing with low latency. Which combination of AWS services should be used?Data Ingestion and Transformation
  24. 74.A global IoT company collects sensor data from millions of devices worldwide. The data is small, JSON-formatted, and arrives continuously at a high velocity. The company needs to ingest this data, apply immediate filtering (e.g., remove malformed records), and route valid records to an Amazon S3 data lake for long-term storage and a separate Amazon Kinesis Data Stream for real-time anomaly detection. The solution must be fully managed, scalable, and minimize operational overhead. Which AWS service combination should be used for this ingestion and initial routing?Data Ingestion and Transformation
  25. 75.A large manufacturing company needs to ingest operational sensor data from thousands of industrial machines located on factory floors. The data is generated continuously, with each machine producing small, frequent data points (e.g., temperature, pressure, vibration). The company requires real-time analytics on this streaming data to detect anomalies and predict maintenance needs. The solution must be highly scalable and cost-effective. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  26. 76.A data engineering team needs to build an ETL pipeline for a marketing analytics platform. The pipeline must process large volumes of historical customer interaction data (several terabytes) stored in Amazon S3, apply complex business logic (e.g., deduplication, aggregation, sentiment analysis using external libraries), and then load the refined data into an Amazon Redshift data warehouse. The solution needs to be scalable, cost-effective, and provide a serverless approach for ETL job execution. Which AWS service is most appropriate for the transformation phase?Data Ingestion and Transformation
  27. 77.A data engineering team needs to build an ETL pipeline for a marketing analytics platform. Raw campaign data arrives daily as CSV files in an Amazon S3 bucket. The team needs to clean, normalize, and enrich this data by joining it with customer segmentation data from an Amazon RDS PostgreSQL database. The transformed data should be stored in Parquet format in another S3 bucket, partitioned by date, and made available for querying by Amazon Athena. The solution needs to be fully managed, serverless, and support complex data transformations without managing underlying compute infrastructure. Which AWS service is the most appropriate for this transformation workload?Data Ingestion and Transformation
  28. 78.A data engineering team manages a daily ETL pipeline that extracts data from a transactional database, transforms it using AWS Glue, and loads it into an Amazon S3 data lake. Over time, the volume of data has grown significantly, causing the Glue job to exceed its allocated runtime window. The team needs to reduce the job's execution time. They have already optimized the Glue script for efficiency. Which cost optimization strategy should they consider NEXT to improve runtime performance for the growing data volume?Data Operations and Monitoring
  29. 79.A data analytics company uses AWS Managed Workflows for Apache Airflow (MWAA) to orchestrate hundreds of data pipelines. Each pipeline is defined as a DAG (Directed Acyclic Graph) and often interacts with services like AWS S3, AWS Glue, and Amazon Redshift. The company needs to implement a centralized alerting mechanism for failed DAG runs that provides specific details about the failure, including the task name, error message, and DAG ID, to a dedicated Slack channel. Which solution provides the MOST efficient and scalable way to achieve this?Data Operations and Monitoring
  30. 80.A large manufacturing company needs to ingest real-time operational data from hundreds of industrial sensors located on its factory floor. The data includes temperature, pressure, and machine status, and must be processed immediately to detect anomalies and trigger alerts. The solution needs to handle fluctuating data volumes, ensure high durability, and integrate seamlessly with AWS Lambda for real-time processing. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  31. 81.A research institution collects genomic sequencing data from various instruments, generating files up to 100 GB each. This data needs to be securely transferred from on-premises storage to Amazon S3 for long-term archival and analysis. The transfer must be highly optimized for network throughput, fault-tolerant, and support scheduled, incremental transfers. Which AWS service is specifically designed for this type of data transfer?Data Ingestion and Transformation
  32. 82.A data engineering team uses AWS Step Functions to orchestrate a complex ETL workflow that runs hourly. The workflow processes data from Amazon S3, transforms it using AWS Glue, and loads it into Amazon Redshift. Lately, the team has noticed that some hourly runs occasionally fail due to transient network issues or temporary unavailability of downstream services. The team needs to implement a mechanism to automatically re-attempt failed steps within the workflow without restarting the entire execution. Which Step Functions feature should they leverage?Data Operations and Monitoring
  33. 83.A research institution collects genomic sequencing data from various instruments, generating files up to several terabytes in size. These files are initially stored on a network file system (NFS) in their on-premises data center. The institution needs to regularly move these files to an Amazon S3 Glacier Flexible Retrieval vault for long-term archival and cost-effective storage. The transfer process must be reliable, handle large files efficiently, and be scheduled to run weekly. Which AWS service should be used to automate this ingestion process?Data Ingestion and Transformation
  34. 84.A retail company collects point-of-sale (POS) data from thousands of stores. Each store sends small, infrequent batches of sales data (CSV files, ~10-20MB each) to a central location. This data needs to be ingested into an Amazon S3 staging bucket, then transformed and loaded into an Amazon Redshift data warehouse nightly. The transformation involves simple data cleansing, aggregation, and lookup against a product catalog in S3. The solution should be cost-effective for batch processing and easy to manage. Which AWS service is best suited for the transformation and loading (ETL) phase?Data Ingestion and Transformation
  35. 85.A gaming company collects telemetry data from millions of players, generating petabytes of data daily. This data needs to be processed in near real-time to detect fraudulent activities and identify emerging game trends. The processing requires complex transformations, aggregations, and pattern matching over sliding windows. The solution must be fully managed, highly scalable, and support custom code. Which AWS service is most suitable for this real-time data transformation?Data Ingestion and Transformation
  36. 86.A financial services company uses AWS Step Functions to orchestrate a complex data ingestion and processing pipeline. This pipeline involves multiple Lambda functions, Glue jobs, and S3 operations. Due to the sensitive nature of the data and strict compliance requirements, the company needs to ensure that all execution history, including input and output data for each step, is retained for a minimum of five years for auditing purposes. Standard Step Functions execution history retention is 90 days. Which is the MOST cost-effective and operationally efficient solution to meet this long-term retention requirement?Data Operations and Monitoring
  37. 87.A gaming company is developing a new mobile game and anticipates millions of concurrent players. They need to collect in-game telemetry data, such as player actions, scores, and session information, in real-time. This data must be immediately available for anomaly detection and live dashboards, and later archived to Amazon S3 for long-term analytics. The solution must handle unpredictable spikes in traffic and ensure high data durability. Which AWS service should be used to ingest this high-volume, real-time streaming data?Data Ingestion and Transformation
  38. 88.A global IoT company collects sensor data from millions of devices worldwide. The data is high-volume, low-latency, and needs to be ingested into AWS for real-time processing and durable storage. The company has a strict requirement for ordered delivery of events within each device's stream and needs to replay historical data for backfilling and reprocessing when new analytics models are deployed. Which AWS service is best suited to ingest and store this raw stream data?Data Ingestion and Transformation
  39. 89.A data engineering team operates a critical data pipeline that ingests data from various sources into Amazon S3, processes it with AWS Glue, and then loads it into Amazon Redshift. The pipeline runs hourly. The team has identified that the Redshift cluster is consistently over-provisioned during off-peak hours, leading to significant unnecessary costs. During peak hours, however, the cluster performs optimally. They need to implement a solution to automatically scale the Redshift cluster capacity up and down based on the actual workload, specifically for the data loading phase, to optimize costs without impacting peak performance. Which Redshift feature should they leverage?Data Operations and Monitoring
  40. 90.A data engineering team manages a critical data pipeline that processes sensitive customer information. The pipeline involves several AWS services, including Amazon S3 for storage, AWS Glue ETL jobs for transformations, and Amazon Redshift for analytical queries. The company has a strict compliance requirement to encrypt all data at rest and in transit. The team has already configured server-side encryption for S3 buckets and Amazon Redshift clusters. What is the most effective and secure way to ensure that data processed by AWS Glue ETL jobs is encrypted at rest and in transit within the Glue environment?Data Operations and Monitoring
  41. 91.A media company needs to process petabytes of video analytics data stored in Amazon S3. The data scientists require the ability to run ad-hoc SQL queries directly on this data without provisioning or managing any servers. The queries often involve complex joins and aggregations across multiple large datasets. The solution must be cost-effective, paying only for the data scanned. Which AWS service is best suited for this requirement?Data Ingestion and Transformation
  42. 92.A financial services company needs to process large batches of credit card transaction data (hundreds of gigabytes per batch) daily. These batches arrive as CSV files in an S3 bucket. The data must be validated, transformed (e.g., PII masking, currency conversion), and then aggregated before being loaded into a data warehouse for fraud detection and regulatory reporting. The solution needs to be robust, scalable, and support custom Python code for complex transformation logic. Which AWS service is most appropriate for this batch transformation process?Data Ingestion and Transformation
  43. 93.A data engineering team manages a daily ETL pipeline that processes terabytes of data using AWS Glue. The pipeline has been running successfully for months, but recently, the Glue jobs started failing intermittently with 'Container exited with a non-zero exit code' errors, specifically when processing larger datasets. These failures occur without any code changes and seem to correlate with peak load times. The team needs to diagnose and resolve these intermittent failures efficiently. Which action should they take FIRST?Data Operations and Monitoring
  44. 94.A large manufacturing company needs to ingest operational sensor data from thousands of industrial machines located globally. Each machine generates small (approx. 100-byte) telemetry records every second. The data must be ingested with extremely low latency, processed in real-time to detect anomalies, and then stored in an S3 data lake for historical analysis. The solution needs to be highly scalable to handle millions of records per second and provide strong durability guarantees. Which AWS service should be used for the initial ingestion layer?Data Ingestion and Transformation
  45. 95.A research institution collects high-resolution satellite imagery, with individual image files often exceeding 500 GB. These files are generated on-premises and need to be regularly transferred to Amazon S3 for long-term archival and processing by AWS analytics services. The institution has a 1 Gbps internet connection, but transferring these large files takes too long and frequently fails due to network interruptions. The solution must be reliable, secure, and minimize transfer time. Which AWS service should the institution use for data ingestion?Data Ingestion and Transformation
  46. 96.A global e-commerce company needs to track user behavior on its website and mobile applications. This generates billions of small (approx. 500-byte) JSON events daily, with unpredictable spikes in traffic. The data needs to be delivered to an Amazon S3 data lake for batch analytics and also to an Amazon Redshift data warehouse for reporting dashboards. The company wants a fully managed solution with minimal setup and operational overhead, and it needs basic data transformation (e.g., converting JSON to Parquet) before delivery. Which AWS service is most appropriate for this ingestion and delivery?Data Ingestion and Transformation
  47. 97.A financial services company processes sensitive customer transaction data using an AWS Glue ETL pipeline. The pipeline runs hourly and moves data from encrypted S3 buckets to an encrypted Amazon Redshift cluster. Due to compliance requirements, the company must ensure that all data at rest and in transit within the Glue environment (e.g., temporary files, logs, scripts) is encrypted using customer-managed keys (CMKs) from AWS Key Management Service (KMS). Which set of configurations must be applied to meet this strict security requirement?Data Operations and Monitoring
  48. 98.A global media company needs to ingest petabytes of video and audio files daily from various content providers. These files, often several gigabytes in size, arrive asynchronously and must be stored in Amazon S3 for subsequent processing and archiving. The company requires a highly scalable and cost-effective solution that can handle unpredictable peaks in data volume without manual intervention. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  49. 99.A financial services company needs to ingest real-time transaction data from various global branches into an Amazon S3 data lake for immediate analysis. The data arrives as continuous streams of JSON records, and the company requires minimal latency and guaranteed delivery. Which AWS service is best suited for this ingestion requirement?Data Ingestion and Transformation
  50. 100.A data analytics company processes large volumes of streaming data using Amazon Kinesis Data Streams and then delivers this data to Amazon S3 for archival and further processing. They have strict requirements for data integrity and ensuring that no data is lost or duplicated during the transfer from Kinesis to S3. They use an AWS Lambda function to consume records from Kinesis and write them to S3. Occasionally, due to network issues or S3 API throttling, the Lambda function fails to deliver a batch of records to S3. The team needs to implement a solution to prevent data loss and ensure exactly-once delivery semantics where possible, and at-least-once otherwise, for records from Kinesis to S3, with minimal custom code.Data Operations and Monitoring