AWS Certified Data Engineer – Associate practice questions
214 free questions with answers and explanations.
- 1.A data analytics team requires a highly performant, fully managed, and scalable time-series database for ingesting and querying billions of sensor data points per day from IoT devices. They need to analyze trends and anomalies in near real-time. Which AWS service is best suited for this requirement?Data Storage and Management
- 2.A data engineering team needs to implement a data catalog and governance solution for their existing data lake on Amazon S3. The solution must allow fine-grained access control to specific tables and columns, enable data discovery, and integrate with various AWS analytics services. They also want to simplify the process of setting up and managing data access permissions. Which AWS service is purpose-built for these requirements?Data Storage and Management
- 3.A data engineering team is designing a highly available and durable storage solution for frequently accessed, structured data that requires ACID transaction support. The data will be used by multiple applications for real-time analytics and operational reporting, and the team needs to scale read replicas independently. Which AWS service is best suited for this requirement?Data Storage and Management
- 4.A data engineering team is building a serverless data processing pipeline using AWS Lambda and Amazon S3. The pipeline processes billions of small JSON files (average 10KB each) daily. To improve the efficiency and cost-effectiveness of downstream analytics with Amazon Athena, the team wants to consolidate these small files into larger, optimized files and convert them to a columnar format. Which AWS service is best suited for orchestrating and executing this compaction and conversion process in a serverless manner?Data Storage and Management
- 5.A data engineering team is building a serverless data processing pipeline using AWS Lambda. They need to store temporary state and intermediate results for their Lambda functions, which can process up to 100GB of data per invocation. The storage needs to be low-latency, highly available, and accessible directly by the Lambda function without incurring network transfer costs or significant latency overhead for each access. Which storage option is most suitable for this scenario?Data Storage and Management
- 6.A data engineering team is migrating an on-premises data warehouse to Amazon Redshift. The on-premises warehouse uses a traditional ETL process that loads data nightly into a dimensional model. The new Redshift cluster will continue this nightly batch load. To ensure optimal query performance for analytical queries on large fact tables (billions of rows) with frequently joined dimension tables, what Redshift distribution style should be applied to the largest fact table?Data Storage and Management
- 7.A data engineer is designing a data lake on Amazon S3 for an e-commerce company. The data consists of customer transaction records stored as CSV files. To optimize query performance and reduce storage costs, the engineer decides to convert these CSV files into a columnar format. Which columnar data format offers the best balance of compression, query performance, and schema evolution capabilities for this use case?Data Storage and Management
- 8.A data engineering team manages a large data lake on Amazon S3. They use AWS Glue Data Catalog to store metadata for tables, which are partitioned by `year/month/day`. They observe that their Athena queries against these tables are performing poorly and scanning much more data than expected, even when filtering on specific dates. Upon investigation, they find that the Glue Data Catalog is not accurately reflecting the actual partitions on S3. What is the most efficient way to refresh the Glue Data Catalog to ensure partitions are correctly recognized for improved query performance?Data Storage and Management
- 9.A financial services company needs to store transactional data from various source systems. This data is critical, highly structured, and requires strict ACID (Atomicity, Consistency, Isolation, Durability) compliance. The data volume is expected to grow to several terabytes, and the application requires high-performance, low-latency queries and strong transactional integrity. Which AWS managed service is the most appropriate for this requirement?Data Storage and Management
- 10.A data architect is designing a highly available and durable storage solution for critical application backups. The backups are rarely accessed but must be retrievable within minutes in case of a disaster. The data volume is expected to be in the petabytes, and cost-effectiveness for long-term storage is a primary concern. Which Amazon S3 storage class offers the best balance of cost, durability, and retrieval time for this scenario?Data Storage and Management
- 11.A data engineering team is designing a new data lake on AWS. They need to store petabytes of raw, unstructured data, such as images and videos, with high durability and virtually unlimited scalability. The data will be accessed infrequently for analytics, but when accessed, low latency is critical. Which AWS storage service is most appropriate for this requirement?Data Storage and Management
- 12.A data engineer is designing a data archiving strategy for highly sensitive customer data that must be retained for 10 years for regulatory compliance. The data is accessed extremely rarely (e.g., once a year or less) and can tolerate retrieval times of several hours. The primary objective is to minimize storage costs while meeting durability and compliance requirements. Which S3 storage class should be chosen?Data Storage and Management
- 13.A data engineering team is working with a large dataset (hundreds of terabytes) stored in Amazon S3. They need to reduce the storage footprint and improve query performance for analytical queries run by Amazon Athena. The data is accessed frequently, and the team wants to use a compression format that offers a good balance between high compression ratio and fast decompression, without significantly increasing CPU overhead during query execution. Which compression format should they choose?Data Storage and Management
- 14.A data engineer is designing a data lake using Amazon S3. The data consists of large files (hundreds of MB to several GB each) that are frequently accessed by analytical queries. To optimize query performance and reduce scan times, the engineer decides to use a columnar storage format. Which columnar format is commonly used with S3-based data lakes and is known for its efficiency with analytical workloads?Data Storage and Management
- 15.A data engineering team is building a serverless data processing pipeline using AWS Lambda. The Lambda functions need to process large input files (up to 500 MB) and generate intermediate output files, requiring temporary storage that is accessible across multiple invocations within a short period and can be shared between different Lambda functions if needed. The storage solution must be highly performant for file I/O and automatically scale. Which AWS service should they use for this temporary, shared file storage?Data Storage and Management
- 16.A data engineer is tasked with migrating an on-premises data warehouse to Amazon Redshift. The source data contains sensitive customer information and must comply with strict regulatory requirements for data at rest and in transit encryption. The solution must ensure that encryption keys are managed centrally and can be rotated automatically. Which Redshift encryption configuration best meets these requirements?Data Storage and Management
- 17.A global e-commerce company stores customer order data in an Amazon S3 bucket. The data is initially accessed frequently for the first 30 days for order processing and fraud detection. After 30 days, access decreases significantly, but the data must be retained for compliance reasons for 7 years. Retrieval should still be possible within minutes. Which S3 Lifecycle policy transition strategy minimizes costs while meeting these requirements?Data Storage and Management
- 18.A data engineer is designing a data lake solution that ingests raw log files from various sources. These logs contain sensitive customer information and must be encrypted at rest and in transit. The solution needs to support querying the data directly using SQL-like queries without loading it into a database. Which combination of AWS services best meets these requirements?Data Storage and Management
- 19.A data engineering team is building a data pipeline that ingests continuous streams of clickstream data from a website. This data needs to be processed in real-time for immediate analytics and then stored for batch processing later. The solution must be able to handle fluctuating data volumes, scale automatically, and deliver data to multiple consumers while ensuring data durability. Which AWS service is best suited for ingesting and temporarily storing this real-time streaming data?Data Storage and Management
- 20.A data engineering team is designing a data lake using Amazon S3. The data consists of large files, and they want to improve query performance and reduce the amount of data scanned by analytical engines like Amazon Athena. They decide to organize their data in S3 using a directory structure that reflects the data's characteristics. Which strategy should they employ?Data Storage and Management
- 21.A data engineering team is designing a data warehouse using Amazon Redshift. They have a large fact table containing billions of records, which will be frequently joined with a smaller dimension table. To optimize query performance for these joins, the team needs to choose an appropriate distribution style for the fact table. Which distribution style is most suitable for this scenario?Data Storage and Management
- 22.A data engineering team needs to store unstructured data, such as images and videos, generated by a mobile application. This data will be accessed frequently for analytics and machine learning training, and requires high durability and availability. Which AWS storage service is most appropriate for this use case?Data Storage and Management
- 23.A data engineer is designing a data lake using Amazon S3. The data consists of large files (hundreds of MBs to several GBs each) that are frequently accessed by analytical queries. The team wants to reduce storage costs and improve query performance by minimizing I/O operations. Which compression format should the data engineer recommend for these files, considering both compression ratio and query performance?Data Storage and Management
- 24.A data engineer is working with a large dataset (hundreds of terabytes) stored in Amazon S3, consisting of millions of small CSV files. This data is frequently queried by Amazon Athena, leading to high query costs and slow performance. To optimize this, the engineer decides to convert the data to Parquet format and implement a compaction strategy. Which compression codec, when combined with Parquet, would offer the best balance of compression ratio and query performance for columnar analytical queries?Data Storage and Management
- 25.A data engineer needs to store a vast amount of historical sensor data from thousands of IoT devices. This data arrives continuously, is time-series in nature, and requires near real-time ingestion and query capabilities for operational monitoring and anomaly detection. The solution must be highly scalable, fully managed, and cost-effective for petabytes of data. Which AWS service is purpose-built for this use case?Data Storage and Management
- 26.A financial services company needs to store highly sensitive customer transaction data in an Amazon S3 data lake. This data must be encrypted at rest and in transit to meet stringent compliance requirements. They prefer to have full control over the encryption keys and require integration with AWS CloudTrail for auditing key usage. Which encryption method should the data engineer implement for the S3 objects?Data Storage and Management
- 27.A data engineering team is building a new data lake on Amazon S3. They need to implement a comprehensive security and governance solution that allows fine-grained access control to specific tables and columns within the data lake, based on user roles and data classifications. The solution must also integrate with AWS Glue Data Catalog and enable auditing of data access. Which AWS service is purpose-built to centralize and simplify these security and governance tasks for data lakes?Data Storage and Management
- 28.A data engineer is designing a data processing pipeline that involves ingesting large volumes of streaming data (e.g., IoT sensor readings, clickstreams) from various sources. This data needs to be temporarily stored, processed in real-time, and then loaded into a data lake for further analysis. The solution must be highly scalable, durable, and capable of handling fluctuating data ingestion rates without data loss. Which AWS service is best suited for reliably collecting and buffering this streaming data?Data Storage and Management
- 29.A data engineering team is building a data lake and needs a centralized metadata repository for all their datasets stored in Amazon S3, Amazon RDS, and Amazon Redshift. They require a service that can automatically discover schemas, track data lineage, and integrate with various AWS analytics services. Which AWS service should they use?Data Storage and Management
- 30.A data engineering team is designing a new data lake on AWS. They need to store petabytes of unstructured data, such as images and videos, that will be accessed infrequently after an initial upload period. The solution must be highly durable, cost-effective, and provide options for eventual archival. Which AWS storage service is most appropriate for this requirement?Data Storage and Management
- 31.A data engineering team is migrating an on-premises data warehouse to Amazon Redshift. The source system has a customer table with a `customer_id` column that is frequently used in JOIN operations with other large fact tables. To optimize query performance in Redshift, particularly for queries involving this `customer_id`, which distribution style should be applied to the customer table?Data Storage and Management
- 32.A data engineering team manages a large data lake on Amazon S3. They use AWS Glue Data Catalog to store metadata, but frequently add new partitions (folders) to their S3 data based on new data arriving. After adding new data and partitions, queries using Amazon Athena do not immediately reflect the new data. Which AWS Glue Data Catalog operation should the team perform to ensure Athena can query the newly added partitions?Data Storage and Management
- 33.A data engineer needs to implement data partitioning for a large dataset stored in Amazon S3. The data is generated daily and queried primarily by date range, specifically `year`, `month`, and `day`. The goal is to optimize query performance and reduce the amount of data scanned by analytics engines like Amazon Athena. Which partitioning strategy should be applied?Data Storage and Management
- 34.A data analytics team uses Amazon Redshift for their data warehousing needs. They frequently run queries that aggregate data across various tables, and some of these queries involve joining a large fact table with a relatively small dimension table (e.g., a few hundred rows). To optimize query performance and minimize data movement during these specific join operations, which distribution style should be applied to the small dimension table?Data Storage and Management
- 35.A data engineering team is building a new data lake on Amazon S3. They need to implement a robust data catalog that supports schema evolution, data quality checks, and integrates seamlessly with various analytics services like Amazon Athena and Amazon Redshift Spectrum. The solution must also provide fine-grained access control to specific tables and columns within the data lake. Which AWS service provides these capabilities?Data Storage and Management
- 36.A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time streaming data from various sources and store it in Amazon S3. To ensure data quality and schema enforcement for downstream analytics, the engineer needs a mechanism to validate incoming data against predefined schemas and manage metadata. Which AWS service combination should be used?Data Storage and Management
- 37.A data engineer is designing a highly available and durable storage solution for critical business documents. These documents are infrequently accessed (once a quarter) but must be retrieved within seconds when needed. The budget is a significant constraint, and the solution needs to minimize storage costs while meeting the retrieval performance. Which S3 storage class is most suitable?Data Storage and Management
- 38.A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time data from various sources, and multiple analytics teams will use different tools (e.g., Amazon Athena, Amazon Redshift Spectrum, AWS Glue ETL, Amazon EMR) to query and process this data. The engineer needs a centralized repository for table definitions, schema versions, and partition information that all these services can access consistently. Which AWS service provides this foundational metadata management capability?Data Storage and Management
- 39.A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time streaming data from IoT devices, batch data from on-premises systems, and semi-structured logs from web applications. The team needs a centralized metadata repository that can automatically discover schemas, track data lineage, and be queried by various analytics services like Amazon Athena and Amazon Redshift Spectrum. Which AWS service should the data engineer use to meet these requirements?Data Storage and Management
- 40.A data engineer is designing a data lake using Amazon S3. The data consists of large files (several GB each) that are appended daily, but historical data is rarely updated. The team needs to optimize for query performance and cost efficiency when using services like Amazon Athena. Which partitioning strategy should the data engineer recommend?Data Storage and Management
- 41.A global manufacturing company collects sensor data from thousands of IoT devices. This data is time-series in nature, high-volume (petabytes annually), and requires near real-time ingestion and analysis. The data will be used for operational dashboards, anomaly detection, and machine learning models. Which AWS database service is purpose-built to handle these requirements efficiently?Data Storage and Management
- 42.A data engineering team is building a new data lake on AWS. They need to store petabytes of raw, unstructured data from various sources, including application logs, social media feeds, and sensor data. The data needs to be highly available, durable, and cost-effective for long-term storage, with occasional access for analytics. Which AWS service is the most appropriate for this primary storage layer?Data Storage and Management
- 43.A data engineering team is building a real-time analytics pipeline for customer clickstream data. The data arrives in high volume and velocity and needs to be ingested, processed, and made available for querying within seconds. The solution must be fully managed and scale automatically to handle fluctuating traffic. Which AWS service is best suited for ingesting this continuous stream of data?Data Storage and Management
- 44.A financial services company needs to store transactional data from various source systems. This data is highly structured, requires ACID compliance, and must support complex SQL queries for reporting and analytics. The data volume is expected to grow to several terabytes within a year, and high availability is paramount. Which AWS service is the most suitable choice?Data Storage and Management
- 45.A large e-commerce company uses Amazon S3 to store operational data, including customer orders and product catalog information. This data is constantly being updated and new versions are frequently created. The company needs to retain all versions of an object for compliance and to allow for recovery from accidental deletions or overwrites. Which S3 feature should be enabled on the bucket?Data Storage and Management
- 46.A data analytics team uses Amazon Redshift for their data warehousing needs. They frequently run complex analytical queries that involve large joins and aggregations on tables with billions of rows. The current query performance is slow, and the team suspects that data distribution is a bottleneck. Which Redshift table design strategy should the data engineer explore to improve query performance significantly?Data Storage and Management
- 47.A media company stores large video files in Amazon S3. These files are accessed frequently for the first 30 days after upload, then accessed once every 90 days for content review, and finally archived for regulatory compliance for 5 years with very rare access. The company wants to minimize storage costs while ensuring appropriate access performance. Which S3 storage class transition policy should be implemented?Data Storage and Management
- 48.A gaming company collects telemetry data from millions of players, generating petabytes of event data daily. This data is ingested into Amazon Kinesis Data Streams. The company needs to perform real-time aggregations (e.g., calculating average session duration per game, total active players) and detect anomalies in player behavior. The results of these aggregations and anomaly detections should be continuously published to a dashboard for operational monitoring. Which AWS service is best suited for this real-time stream processing and analytics?Data Ingestion and Transformation
- 49.A media company needs to process petabytes of video analytics data stored in Amazon S3. The data scientists require the ability to run ad-hoc SQL queries directly on this data without provisioning or managing any servers. The queries often involve complex joins and aggregations across multiple large datasets. The solution must be cost-effective, paying only for the data scanned. Which AWS service is best suited for this requirement?Data Ingestion and Transformation
- 50.A financial services company needs to process large batches of credit card transaction data (hundreds of gigabytes per batch) daily. These batches arrive as CSV files in an S3 bucket. The data must be validated, transformed (e.g., PII masking, currency conversion), and then aggregated before being loaded into a data warehouse for fraud detection and regulatory reporting. The solution needs to be robust, scalable, and support custom Python code for complex transformation logic. Which AWS service is most appropriate for this batch transformation process?Data Ingestion and Transformation