AWS Certified Data Engineer – Associate practice questions

214 free questions with answers and explanations.

Practice test
  1. 151.A media company needs to process large volumes of video metadata (XML files) generated hourly from its content management system. These files are stored in an on-premises SFTP server. The company wants to automate the ingestion of these files into Amazon S3, parse the XML, extract specific metadata fields, and store the extracted data in a structured format (Parquet) in a separate S3 bucket for analytics. The process should be fully managed and serverless. Which combination of AWS services should be used?Data Ingestion and Transformation
  2. 152.A biotechnology company generates large genomic sequencing files (tens of gigabytes each) from its on-premises lab equipment. These files need to be securely and efficiently transferred to Amazon S3 for archival and processing. The transfers occur periodically, and the company wants to optimize network utilization and ensure data integrity during transit. Which AWS service is best suited for this task?Data Ingestion and Transformation
  3. 153.A large manufacturing company uses thousands of IoT sensors on its factory floor. These sensors generate telemetry data (temperature, pressure, vibration) at a high frequency (tens of thousands of messages per second), each message being small (a few hundred bytes). The company needs to ingest this data, apply simple filtering rules (e.g., discard messages below a certain temperature threshold), and then route the filtered data to Amazon S3 for archival and Amazon DynamoDB for real-time dashboards. Which AWS service is best suited for this real-time ingestion and rule-based routing?Data Ingestion and Transformation
  4. 154.A global e-commerce company needs to track user behavior on its website and mobile applications. The data consists of page views, clicks, and search queries, generated by millions of users worldwide. The total data volume is expected to be several terabytes per day, with high variability in traffic. This data needs to be delivered reliably and cost-effectively to Amazon S3 for long-term storage and to Amazon Redshift for analytical dashboards. The solution should minimize operational overhead and handle automatic scaling. Which AWS service is most appropriate for ingesting this data?Data Ingestion and Transformation
  5. 155.A global e-commerce company needs to process customer orders in real-time. Each order record contains various attributes like customer ID, product details, quantity, and payment information. The company anticipates up to 10,000 orders per second during peak sales events. The data needs to be ingested, validated, and transformed before being stored in an analytical data store. Which AWS service combination should be used to achieve this real-time ingestion and transformation with high throughput and low latency?Data Ingestion and Transformation
  6. 156.A data engineer is tasked with securing an Amazon Redshift cluster. The company policy requires that all data access from applications must use temporary credentials, and these credentials should be automatically rotated. Additionally, access to specific tables within Redshift must be restricted based on the application's role. Which set of AWS services and features should the data engineer use?Data Governance and Security
  7. 157.A global e-commerce company operates its data analytics platform on AWS. They need to ensure that customer data stored in Amazon S3 is retained for a minimum of 7 years for compliance, but also automatically deleted after 10 years to manage storage costs and comply with data minimization principles. The solution should be cost-effective and automated. Which S3 lifecycle policy configuration meets these requirements?Data Governance and Security
  8. 158.A media company stores large volumes of video assets in an S3 bucket. Due to contractual agreements, certain video categories must be permanently deleted exactly 5 years after their creation date, regardless of whether they have been accessed. Other categories have no such deletion requirement. How can the data engineering team automate this selective and precise deletion of specific video categories?Data Governance and Security
  9. 159.A data analytics platform uses Amazon Athena to query data stored in S3. The company needs to implement fine-grained access control such that different teams can only query specific rows of data based on their department ID, which is a column in the dataset. This must be enforced at query time without duplicating data or modifying the underlying S3 objects. Which AWS service or feature should be used?Data Governance and Security
  10. 160.A global pharmaceutical company is building a data lake on AWS S3 to store clinical trial data. This data is subject to strict regulatory requirements, including GDPR and HIPAA. The company needs to implement a robust data retention policy where data must be retained for a minimum of 15 years, and then automatically moved to the lowest-cost archival storage class, and finally deleted after 25 years. The data is rarely accessed after 5 years. Given these requirements, which S3 lifecycle policy configuration is most cost-effective and compliant?Data Governance and Security
  11. 161.A financial institution is building a data warehouse on Amazon Redshift. Due to strict regulatory compliance requirements, all access to the data must be logged and audited, including successful and failed login attempts, as well as all SQL queries executed. Which AWS services should be configured to meet these auditing requirements for Redshift?Data Governance and Security
  12. 162.A data platform team needs to implement a data retention policy for their Amazon S3 data lake. They have varying retention requirements for different datasets, some requiring 3 years, others 5 years, and some 10 years, before permanent deletion. All data is initially stored in S3 Standard. The goal is to automate transitions to a cost-effective infrequent access tier after 30 days and then to an archival tier after 1 year, before final deletion. How should the data engineer configure S3 lifecycle policies to efficiently manage these diverse retention periods and transitions?Data Governance and Security
  13. 163.A healthcare provider is migrating patient data to an Amazon S3 data lake. The data contains Protected Health Information (PHI) and must comply with HIPAA regulations. The security team requires that all PHI be masked for non-production environments and for analytics users who do not require direct access to sensitive identifiers. Which AWS service or feature can be used to implement dynamic data masking for sensitive columns without altering the source data?Data Governance and Security
  14. 164.A data engineering team is designing a new data lake on AWS S3. The team needs to ensure that all data at rest is encrypted, but they also want to minimize operational overhead and avoid managing encryption keys. Which S3 encryption option should they choose?Data Governance and Security
  15. 165.A data analytics team is building a new application that will store sensitive customer data in an Amazon DynamoDB table. To meet compliance requirements, the data must be encrypted at rest. The security team insists on using customer-managed keys for encryption to maintain full control over the encryption key lifecycle. Which encryption option should the team choose for DynamoDB?Data Governance and Security
  16. 166.A data engineering team is building a real-time analytics pipeline using Amazon Kinesis Data Streams. The data contains sensitive customer information that must be encrypted in transit from the producers to the Kinesis stream. Additionally, the data must be encrypted at rest within the Kinesis stream. Which combination of encryption methods should be implemented?Data Governance and Security
  17. 167.A data engineer is designing an access control strategy for a new data lake built on Amazon S3 and AWS Lake Formation. The requirement is to grant different data analysts access to specific columns within a table, while redacting sensitive Personally Identifiable Information (PII) columns for a subset of analysts. The underlying data in S3 must remain unchanged. Which Lake Formation feature should be used to achieve this fine-grained, dynamic control?Data Governance and Security
  18. 168.A financial institution is migrating its on-premises data warehouse to Amazon Redshift. Regulatory requirements state that all data at rest within the data warehouse must be encrypted using FIPS 140-2 validated cryptographic modules, and the encryption keys must be managed by the customer with full auditability of key usage. Which Redshift encryption option should the data engineer choose to meet these requirements?Data Governance and Security
  19. 169.A data analytics company needs to archive historical log data for compliance purposes. The data is stored in Amazon S3 and must be retained for 15 years, with infrequent access expected (once or twice a year). The primary concern is minimizing storage costs while meeting the long-term retention requirement. Which S3 storage class is the most cost-effective choice for this scenario?Data Governance and Security
  20. 170.A data engineer is designing a data lake solution on AWS S3 for a company that handles sensitive customer data. The company's security policy mandates that all data written to S3 must be encrypted using keys that are stored in a FIPS 140-2 Level 3 validated hardware security module (HSM). Which AWS service allows the company to meet this stringent key management requirement for S3 encryption?Data Governance and Security
  21. 171.A data engineering team is designing a new real-time analytics pipeline using Amazon Kinesis Data Streams. The data being ingested contains payment card information (PCI) and other personally identifiable information (PII). Due to PCI DSS compliance, all data in transit through the Kinesis stream must be encrypted. The company requires that the encryption keys be managed centrally by AWS KMS, with the ability to audit key usage and rotate keys annually. Which Kinesis Data Streams encryption option should the team choose?Data Governance and Security
  22. 172.A financial institution is building a new real-time analytics pipeline using Amazon Kinesis Data Streams. Due to strict regulatory compliance, all data ingested into Kinesis must be encrypted at rest and in transit. The solution must also ensure that the encryption keys are managed by the financial institution and rotated automatically. Which Kinesis Data Streams encryption configuration meets these requirements?Data Governance and Security
  23. 173.A healthcare organization stores patient health information (PHI) in an Amazon S3 data lake. Due to HIPAA compliance, all access to this data must be logged, and these logs must be retained for seven years. The organization also needs to ensure that access attempts, both successful and failed, are captured. Which AWS service should the data engineer configure to meet these logging and retention requirements for S3 data access?Data Governance and Security
  24. 174.A large e-commerce platform uses Amazon Redshift for its analytics warehouse. Regulatory requirements dictate that all data in the Redshift cluster must be encrypted at rest. The security team also requires that the encryption keys be rotated annually and that key usage be auditable. Which encryption configuration should the data engineer choose for the Redshift cluster?Data Governance and Security
  25. 175.A financial services company is building a new data lake on AWS S3 to store highly sensitive customer financial records. Regulatory compliance mandates that all data at rest must be encrypted using customer-managed encryption keys (CMKs) from AWS Key Management Service (AWS KMS) and that these keys must be fully managed and controlled by the customer. Which S3 encryption method should the data engineer implement to meet these strict requirements?Data Governance and Security
  26. 176.A global company maintains a data lake on AWS S3, storing various datasets. Due to GDPR and other privacy regulations, the company must ensure that any data containing personal information is automatically identified and masked or redacted before being used for analytics by non-authorized personnel. Which AWS service can help automate the discovery of sensitive data and integrate with data governance workflows for masking?Data Governance and Security
  27. 177.An analytics team uses Amazon Athena to query data stored in an S3 data lake. The data contains personally identifiable information (PII) that must be protected. The security team requires that access to specific sensitive columns be restricted based on the user's department, meaning users from the 'Marketing' department should not see 'Salary' data, but 'HR' users should. How can a data engineer implement this fine-grained column-level access control for Athena users querying S3 data?Data Governance and Security
  28. 178.A financial services company stores highly sensitive customer transaction data in an Amazon S3 bucket. Compliance regulations mandate that all data at rest must be encrypted using encryption keys that the company controls and manages independently. Which S3 encryption option should the data engineer implement to meet this requirement?Data Governance and Security
  29. 179.A financial institution uses Amazon DynamoDB to store transaction data. Due to strict audit requirements, every access attempt to the DynamoDB tables, including successful reads, writes, and administrative actions, must be logged and immutable. These logs must be available for forensic analysis for at least five years. Which AWS service should be integrated with DynamoDB to meet these logging and immutability requirements?Data Governance and Security
  30. 180.A data engineering team is building a solution to manage sensitive customer data across various S3 buckets. The company's compliance policy requires that all PII (Personally Identifiable Information) and PCI (Payment Card Industry) data stored in S3 must be automatically discovered, classified, and reported. This process needs to be continuous and provide alerts for non-compliant data. Which AWS service is specifically designed to meet these discovery, classification, and reporting requirements for sensitive data in S3?Data Governance and Security
  31. 181.A global pharmaceutical company is building a data lake on AWS S3 to store sensitive clinical trial data. The company needs to enforce strict access control, ensuring that data scientists can only query specific columns of a dataset (e.g., patient ID, drug dosage) and only rows pertaining to trials they are authorized for. The solution must integrate seamlessly with Amazon Athena, which is used for querying the data lake. Which AWS service is best suited to implement this fine-grained access control for S3 data queried by Athena?Data Governance and Security
  32. 182.A pharmaceutical company is auditing its data governance practices. They need to ensure that all data access events, including who accessed which data, when, and from where, are logged and immutable for seven years to comply with regulatory requirements. The data is stored across S3, Redshift, and DynamoDB. Which AWS service is primarily responsible for capturing these audit logs?Data Governance and Security
  33. 183.A media company stores petabytes of video content in an S3 data lake. This content is accessed by various internal teams (e.g., editorial, marketing, legal) with different access levels. The company wants to implement a solution that allows each team to only view the content relevant to their roles, including specific video clips or segments, without duplicating data. Which AWS service and feature combination should the data engineer use?Data Governance and Security
  34. 184.A data engineering team is building a new application that will process sensitive customer data. To meet compliance requirements, all access attempts to the data must be logged, including successful and failed attempts, and these logs must be immutable for seven years. Which AWS service should the team use to capture and ensure the integrity of these access logs?Data Governance and Security
  35. 185.A global e-commerce platform uses Amazon Redshift for its analytical data warehouse. Due to GDPR regulations, specific customer attributes (e.g., email addresses, phone numbers) must be masked for users who do not have explicit permission to view them, while still allowing other users (e.g., customer support) to see the full, unmasked data. The solution must be performant and not require significant changes to existing ETL processes. How should a data engineer implement this requirement in Redshift?Data Governance and Security
  36. 186.A global research institution uses Amazon S3 to store large datasets from various projects. They need to implement a data retention policy where project data is retained for a minimum of 5 years, and then automatically deleted. However, some specific highly sensitive datasets within these projects need to be retained for 10 years. What is the most efficient way to manage this dual retention requirement using S3 features?Data Governance and Security
  37. 187.A large enterprise uses multiple AWS accounts and needs to centralize the auditing of all S3 bucket policy changes, IAM role modifications, and network access control list (NACL) updates across its entire AWS environment. The audit logs must be immutable and retained for 7 years for compliance. The security team also requires the ability to quickly search and analyze these logs. Which AWS service combination should be used to meet these requirements?Data Governance and Security
  38. 188.A global financial services company is building a new data lake on Amazon S3 to store highly sensitive customer transaction data. Regulatory compliance dictates that all data at rest must be encrypted using FIPS 140-2 validated cryptographic modules, and the customer must have exclusive control over the encryption keys. Which AWS encryption option best meets these requirements?Data Governance and Security
  39. 189.A data engineering team is migrating a legacy on-premises data warehouse to Amazon Redshift. The legacy system used a custom data masking solution for sensitive columns (e.g., Social Security Numbers, credit card numbers) to prevent unauthorized users from viewing actual data while still allowing analytics. The team needs to replicate this masking functionality in Redshift, ensuring that only authorized users can see unmasked data. Which approach should the data engineer implement?Data Governance and Security
  40. 190.A compliance officer needs to ensure that all changes to S3 bucket policies, IAM roles, and encryption keys used for a data lake are logged and immutable for auditing purposes. This includes tracking who made the change, when, and the exact details of the change. Which AWS service configuration is essential for capturing this type of governance and security-related event data?Data Governance and Security
  41. 191.A healthcare organization is building a data lake on AWS S3 to store patient records. Due to strict HIPAA compliance requirements, all access to the data lake must be authenticated and authorized, with access policies centrally managed and auditable. Which AWS service is best suited for managing fine-grained data access permissions to the S3 data lake?Data Governance and Security
  42. 192.A global media company uses AWS S3 to store petabytes of media assets. Due to licensing agreements, certain assets must be deleted exactly 30 days after their last access, while others must be archived to a cost-effective storage class 90 days after creation and then permanently deleted after 5 years. The company needs an automated and cost-efficient solution to manage these diverse retention and deletion policies. Which AWS service should the data engineer utilize?Data Governance and Security
  43. 193.A data platform team needs to implement a data retention policy for their Amazon S3 data lake. The policy states that all log data should be moved to a cost-effective archival storage class after 30 days and then permanently deleted after 5 years. However, specific log files related to security incidents must be retained indefinitely. The team wants an automated solution that minimizes operational overhead. How should the data engineer design this retention strategy?Data Governance and Security
  44. 194.A data analytics team is building a new application that will store sensitive customer transaction data in an Amazon S3 data lake. The company's security policy dictates that encryption keys for this data must be stored in a FIPS 140-2 Level 3 validated hardware security module (HSM) and that the customer must have exclusive control over the cryptographic operations. The data engineer needs to select an S3 encryption method that meets these stringent requirements. Which option provides the highest level of key control and hardware security for S3 data?Data Governance and Security
  45. 195.A data analytics company needs to archive historical log data for compliance purposes. The data is rarely accessed (less than once a year), but when needed, a retrieval time of 12-48 hours is acceptable. The primary concern is minimizing storage costs while ensuring data durability and compliance with a 10-year retention policy. Which Amazon S3 storage class is the most cost-effective solution for this scenario?Data Governance and Security
  46. 196.A research institution collects high-resolution satellite imagery, with individual image files often exceeding 50 GB. These files are generated on-premises and need to be securely and efficiently transferred to Amazon S3 for long-term archival and subsequent processing. The institution has limited internet bandwidth (100 Mbps upload speed) at its remote data collection sites. Which AWS service is best suited for ingesting this data?Data Ingestion and Transformation
  47. 197.A global gaming company collects player interaction data (e.g., achievements, in-game purchases, chat messages) from its mobile game. This data is generated by millions of players concurrently, resulting in high-volume, low-latency streams. The company needs to perform real-time analytics on this data to detect fraudulent activities and provide personalized player experiences. The solution must be fully managed, scalable, and support SQL-based queries on the incoming streams. Which combination of AWS services is the MOST appropriate for this real-time analytics use case?Data Ingestion and Transformation
  48. 198.A data analytics team needs to process semi-structured log data (JSON format) generated by various applications. These logs are stored hourly in an Amazon S3 bucket. The team wants to perform ad-hoc queries directly on this data without loading it into a traditional database or setting up complex ETL pipelines. They also need to easily discover the schema of the JSON files for querying purposes. Which AWS service is the MOST cost-effective and flexible solution for this scenario?Data Ingestion and Transformation
  49. 199.A global e-commerce company needs to process customer order data in near real-time. Each order record needs to be validated, enriched with customer demographic information from a DynamoDB table, and then stored in an Amazon S3 data lake in Parquet format. The solution must be serverless, cost-effective, and scale automatically with fluctuating order volumes. Which AWS service is the MOST appropriate for this transformation?Data Ingestion and Transformation
  50. 200.A media company needs to process large volumes of video metadata (XML files) generated hourly. These files are stored in an Amazon S3 bucket and require parsing, validation against a schema, and transformation into a normalized JSON format before being loaded into a data warehouse for analytics. The solution must be cost-effective, serverless, and handle varying volumes of data efficiently. Which combination of AWS services is MOST suitable for this transformation?Data Ingestion and Transformation