AWS Certified Data Engineer – Associate practice questions

214 free questions with answers and explanations.

Practice test
  1. 101.A data engineering team operates an AWS Glue ETL job that processes large datasets daily. The job is configured to use a specific number of Data Processing Units (DPUs). Recently, the team noticed that the job's execution time has significantly increased, and CloudWatch metrics show high CPU utilization and memory pressure on the Glue worker nodes. They want to improve the job's performance and reduce its runtime without incurring excessive costs. What is the MOST effective first step to optimize the Glue job's performance in this scenario?Data Operations and Monitoring
  2. 102.A data engineering team manages a complex data pipeline using AWS Step Functions, which orchestrates various AWS Glue jobs and Lambda functions. The pipeline processes sensitive financial data, and there is a strict requirement to encrypt all temporary data generated during Glue job execution. This temporary data includes shuffle data, spill data, and intermediate results stored on disk or in S3. The team needs to implement a solution that ensures this temporary data is encrypted both at rest and in transit, with customer-managed keys (CMKs) from AWS Key Management Service (KMS), applied consistently across all Glue jobs in the pipeline. Which approach should they use?Data Operations and Monitoring
  3. 103.A global e-commerce company wants to centralize customer clickstream data from various regional websites. The data arrives as JSON objects, approximately 1KB each, with a peak ingestion rate of 50,000 records per second. The company needs to transform this data by filtering out bot traffic and enriching it with customer demographic information from a DynamoDB table before storing it in an Amazon S3 data lake in Parquet format. The solution must provide near real-time processing with low latency. Which AWS service is most appropriate for the *transformation* of this data?Data Ingestion and Transformation
  4. 104.A data analytics team is using Amazon S3 to store raw log files from various applications. They need to query this data directly in S3 using standard SQL without provisioning or managing any servers. The queries are ad-hoc and exploratory, and the team wants to minimize costs by paying only for the data scanned. Which AWS service should they use for this requirement?Data Ingestion and Transformation
  5. 105.A media company uses AWS DataSync to transfer large volumes of media files from an on-premises Network File System (NFS) to Amazon S3 for archival and processing. The transfers run daily during off-peak hours. Recently, the company noticed that some transfer tasks are failing intermittently, and they need to quickly identify the root cause of these failures. Which AWS service would be MOST effective for diagnosing and troubleshooting DataSync task failures?Data Operations and Monitoring
  6. 106.A global logistics company receives shipment tracking updates from thousands of partners in various formats (CSV, XML, JSON). These files are uploaded to an S3 bucket several times a day. The company needs to consolidate, normalize, and enrich this data with internal reference tables before storing it in a unified Parquet format in another S3 bucket for analytical purposes. The processing should be cost-effective and handle schema evolution. Which AWS service should be used for this transformation?Data Ingestion and Transformation
  7. 107.A global media company needs to ingest petabytes of video and audio files daily from various content providers. These files, often several gigabytes in size, arrive asynchronously and must be stored in Amazon S3 for subsequent processing and archiving. The company requires a highly scalable and cost-effective solution that can handle unpredictable peaks in data volume without manual intervention. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  8. 108.A financial services company needs to ingest real-time transaction data from various global branches into an Amazon S3 data lake for immediate analysis. The data arrives as continuous streams of JSON records, and the company requires minimal latency and guaranteed delivery. Which AWS service is best suited for this ingestion requirement?Data Ingestion and Transformation
  9. 109.A data analytics company processes large volumes of streaming data using Amazon Kinesis Data Streams and then delivers this data to Amazon S3 for archival and further processing. They have strict requirements for data integrity and ensuring that no data is lost or duplicated during the transfer from Kinesis to S3. They use an AWS Lambda function to consume records from Kinesis and write them to S3. Occasionally, due to network issues or S3 API throttling, the Lambda function fails to deliver a batch of records to S3. The team needs to implement a solution to prevent data loss and ensure exactly-once delivery semantics where possible, and at-least-once otherwise, for records from Kinesis to S3, with minimal custom code.Data Operations and Monitoring
  10. 110.A financial services company needs to process large volumes of historical transaction data stored in Amazon S3. The data is in CSV format, partitioned by date, and totals hundreds of terabytes. The company needs to perform complex aggregations, join data with external reference tables, and cleanse inconsistent records. The processing can run daily during off-peak hours, and cost efficiency is a major concern. Which AWS service is best suited for this transformation task?Data Ingestion and Transformation
  11. 111.A research institution collects high-resolution satellite imagery, with individual image files ranging from 500 MB to 5 GB. These files are generated on-premises and need to be regularly transferred to Amazon S3 for long-term archival and further processing by AWS Lambda functions. The transfers must be reliable, secure, and optimized for large file sizes over a WAN connection. Which AWS service would be most effective for this data ingestion task?Data Ingestion and Transformation
  12. 112.A data engineering team manages a data pipeline that uses AWS Step Functions to orchestrate a series of AWS Lambda functions. One particular Lambda function performs a critical data validation step that can sometimes take longer than expected due to external service dependencies, occasionally exceeding its configured timeout. When this happens, the Step Functions state machine fails, requiring manual reprocessing of the entire batch. The team wants to implement a robust solution that allows the Lambda function to complete its work, even if it exceeds its initial timeout, without failing the Step Functions workflow or requiring manual intervention, and ensuring the workflow progresses once the validation is complete. Which pattern should they adopt?Data Operations and Monitoring
  13. 113.A pharmaceutical company needs to ingest experimental results from laboratory instruments. These instruments generate data in various proprietary binary formats and store them on a local network file share (NFS). Data needs to be periodically transferred to Amazon S3, where custom Lambda functions will perform format conversion and initial processing. The solution must provide low-latency access to the data on-premises for a short period before it's moved to S3. Which AWS service provides a hybrid cloud storage solution that fits this requirement?Data Ingestion and Transformation
  14. 114.A data analytics company uses AWS Glue Data Catalog as a central metadata repository for its data lake, which resides in Amazon S3. They have several automated processes that frequently update table schemas, add new partitions, and drop old partitions in the Data Catalog. Recently, they've noticed that some downstream applications querying these tables are occasionally failing due to schema mismatches or missing partitions, even though the Glue jobs responsible for updates reported success. This indicates a potential eventual consistency issue with Data Catalog updates. What is the most robust solution to ensure that downstream applications consistently retrieve the latest metadata from the AWS Glue Data Catalog after updates?Data Operations and Monitoring
  15. 115.A data engineering team operates a critical data pipeline that ingests sensor data from IoT devices into Amazon S3, processes it with AWS Glue, and then stores the refined data in Amazon Redshift. The pipeline runs hourly. Recently, the team noticed that the Redshift cluster's CPU utilization is consistently high during the Glue job's write phase, leading to increased query latency for downstream analytics users. They need to optimize the Glue job's Redshift write performance to reduce cluster strain and improve overall pipeline efficiency. Which specific Glue job configuration or optimization should they implement?Data Operations and Monitoring
  16. 116.A company has a legacy on-premises database (Microsoft SQL Server) containing critical business data. They need to migrate this database to Amazon RDS for PostgreSQL and establish a continuous, low-latency replication link to keep the RDS instance up-to-date with changes in the on-premises database. The solution must handle ongoing schema changes and data type conversions automatically. Which AWS service should be used for this continuous data ingestion and transformation?Data Ingestion and Transformation
  17. 117.A data engineering team is building a new data lake on Amazon S3. They plan to use AWS Glue Data Catalog for metadata management and AWS Athena for ad-hoc querying. Data is ingested daily in CSV format. To optimize query performance and reduce Athena costs, the team needs to convert the data into an columnar, compressed format and partition it effectively. Which approach offers the BEST balance of cost-effectiveness, performance, and operational overhead?Data Operations and Monitoring
  18. 118.A global e-commerce company wants to analyze customer clickstream data in near real-time. The data is generated from their website and mobile applications, with peak volumes reaching several terabytes per hour. They need to enrich this raw data by joining it with customer profile information from a DynamoDB table and then store the enriched data in Amazon S3 for further analytics. Which AWS service is best suited for performing this near real-time data enrichment and transformation?Data Ingestion and Transformation
  19. 119.A data engineering team manages a critical real-time data pipeline that ingests high-velocity clickstream data into Amazon Kinesis Data Streams. Downstream, an AWS Lambda function processes these records, performs transformations, and stores them in Amazon DynamoDB. During peak traffic, the team observes that the Lambda function is frequently throttled, leading to increased Kinesis `IteratorAge` and potential data loss. They need to resolve the Lambda throttling issue to ensure continuous, real-time processing without data loss, while minimizing cost increases. Which combination of actions should they take?Data Operations and Monitoring
  20. 120.A data engineering team manages a critical ETL pipeline that processes sensitive customer data daily. The pipeline uses AWS Glue jobs, Amazon S3 for data storage, and AWS Lambda for orchestration. Recently, an audit revealed that while data at rest in S3 is encrypted, there is no explicit mechanism to ensure data in transit between Glue and S3, or Lambda and S3, is also encrypted. The team needs to implement a solution that encrypts all data in transit for this pipeline with minimal operational overhead, adhering to compliance requirements. Which AWS service or feature should they implement?Data Operations and Monitoring
  21. 121.A financial services company needs to migrate a 150 TB on-premises data warehouse (Microsoft SQL Server) to Amazon Redshift. The migration must be performed with minimal downtime, ensuring data consistency throughout the process, including ongoing replication of changes. The solution should be fully managed. Which AWS service is primarily designed for this type of database migration?Data Ingestion and Transformation
  22. 122.A global e-commerce company uses AWS Lambda functions to process real-time order data. These Lambda functions are invoked by Amazon Kinesis Data Streams. During peak sales events, the Lambda functions occasionally fail due to external API rate limits, causing data loss if not handled properly. The data engineering team needs a solution to gracefully handle these transient failures, ensuring that all failed records are reprocessed without manual intervention and without blocking the Kinesis stream. Which approach should they implement?Data Operations and Monitoring
  23. 123.A gaming company is developing a new mobile game and anticipates millions of concurrent players. The game generates high-velocity, low-latency telemetry data (player actions, scores, in-game events) that needs to be captured for real-time leaderboards, fraud detection, and analytics. The data needs to be available for processing within seconds of being generated. Which AWS service is best suited for ingesting this real-time event data?Data Ingestion and Transformation
  24. 124.A data engineering team manages a complex data pipeline involving multiple AWS services, including AWS Glue, Amazon Kinesis, and Amazon Redshift. They need to implement real-time monitoring and anomaly detection for key operational metrics (e.g., Kinesis PutRecord failures, Glue job error rates, Redshift query latencies). When an anomaly is detected, the system should automatically trigger a custom remediation action (e.g., a Lambda function to restart a Glue job or scale a Kinesis stream) and notify on-call engineers. Which combination of AWS services provides the MOST automated and integrated solution?Data Operations and Monitoring
  25. 125.A manufacturing company operates thousands of IoT sensors on its factory floor, generating high-volume, low-latency telemetry data (e.g., temperature, pressure, vibration) at a rate of millions of events per second. The company needs to ingest this data, apply simple filtering and aggregation in real-time, and then route the processed data to both an S3 data lake for historical analysis and Amazon DynamoDB for real-time dashboards. Which AWS service should be used for the ingestion and real-time processing of this IoT data?Data Ingestion and Transformation
  26. 126.A marketing team needs to collect customer clickstream data from a high-traffic website. The data, in JSON format, needs to be ingested in real-time and delivered to an Amazon S3 bucket for long-term storage and subsequent batch analysis. The solution must be fully managed, highly scalable, and minimize operational overhead. Which AWS service is the MOST appropriate for this ingestion task?Data Ingestion and Transformation
  27. 127.A global media company needs to process large volumes of video metadata (XML files) generated hourly from various content creation studios worldwide. These files are typically 10-50 MB each, and the company requires a serverless, scalable, and cost-effective solution to parse these XML files, extract specific attributes (e.g., title, duration, creation date), and store them in a structured format in Amazon DynamoDB for quick querying. Which combination of AWS services would BEST meet these requirements?Data Ingestion and Transformation
  28. 128.A financial services company needs to audit all access to sensitive customer data stored in Amazon S3. They need to capture every S3 object access event (read, write, delete) and send these events to a centralized logging system for compliance and security analysis. The solution must be highly available and scalable to handle millions of events per hour. Which AWS service should be enabled to capture these events, and which service should be used to deliver them for analysis?Data Ingestion and Transformation
  29. 129.A data analytics team needs to process semi-structured log data (JSON format) generated by various microservices and stored in an Amazon S3 bucket. They require a serverless solution to transform this data into a columnar format (Parquet) and partition it by date for optimized querying. The processing should be scheduled to run daily. Which AWS service is the MOST suitable for this transformation task?Data Ingestion and Transformation
  30. 130.A media company needs to process large volumes of video files (averaging 500 MB to 5 GB each) uploaded daily by content creators to an Amazon S3 bucket. The processing involves extracting metadata, generating thumbnails, and transcoding the videos into multiple formats. This is a batch process triggered by new file uploads. The solution must be highly scalable, cost-effective, and provide robust error handling and retry mechanisms. Which AWS service is most suitable for orchestrating and executing these transformations?Data Ingestion and Transformation
  31. 131.A data engineering team needs to build a highly available and fault-tolerant pipeline to process customer interaction data from various social media platforms. The data is ingested as raw JSON files into an Amazon S3 landing zone. They need to perform complex transformations, enrichments, and aggregations using Apache Spark, requiring a highly scalable and resilient processing environment. Which AWS service is the MOST appropriate for implementing this transformation solution?Data Ingestion and Transformation
  32. 132.A financial institution needs to process large volumes of historical market data, approximately 500 TB, stored in various on-premises network attached storage (NAS) devices. The data needs to be moved to Amazon S3 for long-term archival and subsequent analysis. The institution has limited internet bandwidth and strict security requirements, preventing direct internet transfers of such a large volume. Which AWS service is the MOST appropriate for this data ingestion task?Data Ingestion and Transformation
  33. 133.A data analytics team needs to build a robust and fault-tolerant pipeline to process customer interaction data from various social media platforms. The data arrives as semi-structured JSON, with varying schemas and nested structures. The processing involves schema inference, flattening nested data, enriching with internal customer IDs, and storing the transformed data in a data warehouse (Amazon Redshift). The pipeline must handle data volumes up to 10 TB daily and be able to recover from failures without data loss. Which AWS service is most appropriate for orchestrating and executing these complex transformations?Data Ingestion and Transformation
  34. 134.A large manufacturing company uses thousands of IoT sensors on its factory floor. These sensors generate telemetry data (temperature, pressure, vibration) at high frequency, sending small JSON messages. The company needs to ingest this data, apply basic filtering (e.g., discard data below a certain temperature threshold), and then route the filtered data to multiple destinations: an Amazon Kinesis Data Stream for real-time analytics and an Amazon S3 bucket for long-term archival. The solution must be fully managed and scalable.Data Ingestion and Transformation
  35. 135.A financial institution needs to process massive volumes of historical market data, approximately 500 TB, stored in an on-premises data center. The data is currently on Network Attached Storage (NAS) and needs to be moved to Amazon S3 for long-term archival and analytics. The institution has a 10 Gbps direct connection to AWS (AWS Direct Connect). Which AWS service is the most cost-effective and efficient for this one-time bulk migration?Data Ingestion and Transformation
  36. 136.A retail company needs to process daily sales reports generated by point-of-sale (POS) systems across thousands of stores. Each store uploads a CSV file containing transaction data to a central SFTP server hosted in AWS. The company requires a serverless solution to automatically detect new files on the SFTP server, move them to an S3 bucket, and then trigger an ETL job to transform the CSV data into a Parquet format, partition it by store ID and date, and store it in another S3 bucket for analytics. Which AWS services should be used to automate this end-to-end pipeline?Data Ingestion and Transformation
  37. 137.A global media company needs to process large volumes of video metadata (XML files) generated hourly by various content partners. These files are typically 10-50 MB each, and arrive via SFTP. The company requires a fully managed and scalable solution to ingest these files into Amazon S3, transform them into a structured format (Parquet), and store them in a data lake for analytics. What is the most appropriate architecture for this ingestion and transformation pipeline?Data Ingestion and Transformation
  38. 138.A research laboratory generates massive genomic datasets, with individual files often exceeding 100 GB. These files are stored on an on-premises Network Attached Storage (NAS) appliance. The lab needs to regularly transfer these files to Amazon S3 for long-term archival and analysis. The transfers must be reliable, resumable, and optimized for high bandwidth utilization over a dedicated network connection (AWS Direct Connect). What is the most efficient and cost-effective service for this data ingestion requirement?Data Ingestion and Transformation
  39. 139.A research institution is collecting sensor data from an array of environmental monitoring stations located in remote areas with intermittent internet connectivity. The data needs to be securely stored and eventually ingested into AWS for analysis. The stations can store up to 10 TB of data locally before needing to offload. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  40. 140.A healthcare provider needs to ingest patient data from various on-premises systems, including electronic health records (EHR) and laboratory information systems (LIS). These systems generate large, structured CSV and XML files (up to 50 GB each) nightly. The data needs to be moved to Amazon S3 for secure archiving and subsequent processing. The transfer must be highly secure, encrypted, and provide comprehensive logging and auditing capabilities for compliance reasons. Which AWS service is most suitable for this secure and automated data ingestion?Data Ingestion and Transformation
  41. 141.A global gaming company collects player interaction data (e.g., achievements, in-game purchases, session durations) from millions of concurrent players. This data needs to be ingested in real-time, aggregated into hourly summaries, and then loaded into an Amazon Redshift data warehouse for analytical reporting. The solution must be serverless and handle fluctuating data volumes efficiently. Which AWS service combination is best suited for the real-time ingestion and aggregation before loading to Redshift?Data Ingestion and Transformation
  42. 142.A large enterprise needs to ingest real-time operational metrics from thousands of on-premises servers into AWS for monitoring and analysis. The data consists of small, frequent updates (e.g., CPU utilization, memory usage) and must be delivered to Amazon S3 for long-term storage and Amazon CloudWatch for immediate visualization. The solution must be highly available, scalable, and minimize operational overhead. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  43. 143.A global media company receives large video files (up to 50 GB each) from content creators worldwide. These files often arrive via SFTP. The company needs to automatically process these files by extracting metadata, transcoding them into multiple formats, and storing the processed versions in Amazon S3. Which combination of AWS services provides a robust, scalable, and serverless solution for this workflow?Data Ingestion and Transformation
  44. 144.A financial institution needs to perform complex, ad-hoc queries on petabytes of historical transaction data stored in Amazon S3. The data is stored in various formats, including CSV, JSON, and Parquet. Analysts require a serverless solution that can query data directly in S3 without loading it into a database, and they need to pay only for the data scanned. Which AWS service is best suited for this requirement?Data Ingestion and Transformation
  45. 145.A global logistics company needs to track the real-time location of its delivery vehicles. Each vehicle sends GPS coordinates and status updates every 5 seconds. The data volume is expected to be extremely high, potentially millions of messages per second during peak times, with each message being small (around 200 bytes). The company needs to store this data for at least 7 days for immediate operational monitoring and analysis, and then archive it in Amazon S3 for long-term historical analysis. The solution must be highly scalable, durable, and provide ordered delivery of records within a partition. Which AWS service is most appropriate for ingesting this real-time, high-volume data stream?Data Ingestion and Transformation
  46. 146.A retail company collects daily sales reports from thousands of point-of-sale (POS) systems across various stores. Each report is a CSV file, typically 10-100 MB, and is uploaded to a centralized SFTP server hosted on-premises. The company needs to move these files to Amazon S3, combine them into a single Parquet file per day, and partition the data by date for efficient querying. This process must run daily after all reports are received. Which transformation service is most appropriate for this task?Data Ingestion and Transformation
  47. 147.A healthcare provider needs to migrate a 200 TB on-premises Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime for applications that rely on the database, and data consistency is paramount. The migration involves both schema conversion and continuous data replication. Which AWS service is the MOST appropriate for this migration?Data Ingestion and Transformation
  48. 148.A financial institution needs to ingest vast amounts of historical market data, approximately 500 TB, from on-premises network-attached storage (NAS) devices into Amazon S3 for long-term archival and future analytics. The data transfer needs to be highly secure, reliable, and minimize network bandwidth consumption. The institution has a 1 Gbps internet connection. Which AWS service is the MOST appropriate for this ingestion task?Data Ingestion and Transformation
  49. 149.A research team needs to process petabytes of scientific simulation data stored in Amazon S3. The data is in various formats, including CSV, JSON, and Parquet. They require a serverless, ad-hoc query engine that can directly query data in S3 without needing to load it into a database, supporting standard SQL for data exploration and analysis. Which AWS service is best suited for this requirement?Data Ingestion and Transformation
  50. 150.A global e-commerce company needs to process customer order data in near real-time. Each order record, averaging 1 KB in size, arrives concurrently from various regional microservices, with peak rates reaching 50,000 records per second. The company requires the data to be immediately available for fraud detection and personalized recommendations, and also stored in Amazon S3 for historical analysis. Which combination of AWS services should be used for this ingestion and initial processing?Data Ingestion and Transformation