1. A data engineering team manages a large data lake on Amazon S3. They use AWS Glue Data Catalog to store metadata, but frequently add new partitions (folders) to their S3 data based on new data arriving. After adding new data and partitions, queries using Amazon Athena do not immediately reflect the new data. Which AWS Glue Data Catalog operation should the team perform to ensure Athena can query the newly added partitions?
Data Storage and Management
A.Execute `MSCK REPAIR TABLE` in Amazon Athena.
B.Update the table schema manually in the Glue Data Catalog.
C.Create a new Glue Crawler to re-scan the entire S3 path.
D.Run a Glue ETL job to transform the data.
Show answerAnswer
A. Execute `MSCK REPAIR TABLE` in Amazon Athena.
When new partitions are manually added to an S3 location that corresponds to an AWS Glue Data Catalog table, the `MSCK REPAIR TABLE` command in Athena (or a Glue Crawler) is necessary to update the Glue Data Catalog with the metadata for these new partitions. This ensures Athena can discover and query the new data.
2. A data engineering team is building a serverless data processing pipeline using AWS Lambda. The Lambda functions need to process large input files (up to 500 MB) and generate intermediate output files, requiring temporary storage that is accessible across multiple invocations within a short period and can be shared between different Lambda functions if needed. The storage solution must be highly performant for file I/O and automatically scale. Which AWS service should they use for this temporary, shared file storage?
Data Storage and Management
A.AWS Lambda /tmp directory
B.Amazon S3
C.Amazon EBS
D.Amazon EFS
Show answerAnswer
D. Amazon EFS
Amazon EFS (Elastic File System) is a scalable, elastic, cloud-native NFS file system that can be mounted to multiple Lambda functions concurrently. It provides shared, persistent storage with high throughput and low latency, making it ideal for scenarios where Lambda functions need to access and share large files or maintain state across invocations beyond the ephemeral /tmp directory.
3. A data engineer needs to implement data partitioning for a large dataset stored in Amazon S3. The data is generated daily and queried primarily by date range, specifically `year`, `month`, and `day`. The goal is to optimize query performance and reduce the amount of data scanned by analytics engines like Amazon Athena. Which partitioning strategy should be applied?
Data Storage and Management
A.Partition by `customer_id`
B.Store all data in a single, large partition
C.Partition by `region` and `product_category`
D.Partition by `year/month/day`
Show answerAnswer
D. Partition by `year/month/day`
Partitioning data by `year/month/day` aligns directly with the primary query pattern (date range). This strategy allows analytics engines to prune unnecessary S3 prefixes, scanning only the relevant data, which significantly improves query performance and reduces costs.
4. A data analytics team uses Amazon Redshift for their data warehousing needs. They frequently run queries that aggregate data across various tables, and some of these queries involve joining a large fact table with a relatively small dimension table (e.g., a few hundred rows). To optimize query performance and minimize data movement during these specific join operations, which distribution style should be applied to the small dimension table?
Data Storage and Management
A.DISTSTYLE AUTO
B.DISTSTYLE ALL
C.DISTSTYLE KEY
D.DISTSTYLE EVEN
Show answerAnswer
B. DISTSTYLE ALL
For small dimension tables that are frequently joined with large fact tables, `DISTSTYLE ALL` is the most efficient choice. It copies the entire dimension table to every compute node, allowing Redshift to perform local joins without requiring data movement across the network. This significantly improves query performance for such joins.
5. A data engineer is tasked with migrating an on-premises data warehouse to Amazon Redshift. The source data contains sensitive customer information and must comply with strict regulatory requirements for data at rest and in transit encryption. The solution must ensure that encryption keys are managed centrally and can be rotated automatically. Which Redshift encryption configuration best meets these requirements?
Data Storage and Management
A.Configure Redshift to use a hardware security module (HSM) for key management and disable SSL.
B.Use client-side encryption for all data before loading into Redshift and disable Redshift encryption.
C.Enable encryption at rest using AWS-managed keys and rely on default Redshift connection encryption.
D.Enable encryption at rest with AWS Key Management Service (KMS) and enforce SSL for all client connections.
Show answerAnswer
D. Enable encryption at rest with AWS Key Management Service (KMS) and enforce SSL for all client connections.
KMS provides centralized key management, automatic rotation, and integration with Redshift for at-rest encryption. Enforcing SSL ensures data in transit is encrypted, meeting both requirements.
6. A data engineer is designing a data lake on Amazon S3 for an e-commerce company. The data consists of customer transaction records stored as CSV files. To optimize query performance and reduce storage costs, the engineer decides to convert these CSV files into a columnar format. Which columnar data format offers the best balance of compression, query performance, and schema evolution capabilities for this use case?
Data Storage and Management
A.JSON
B.XML
C.Avro
D.Parquet
Show answerAnswer
D. Parquet
Apache Parquet is a columnar storage format optimized for analytical queries. It provides efficient data compression, improves query performance by allowing column pruning, and supports schema evolution, making it ideal for data lake scenarios.
7. A data engineering team is building a new data lake on Amazon S3. They need to implement a robust data catalog that supports schema evolution, data quality checks, and integrates seamlessly with various analytics services like Amazon Athena and Amazon Redshift Spectrum. The solution must also provide fine-grained access control to specific tables and columns within the data lake. Which AWS service provides these capabilities?
Data Storage and Management
A.AWS Lake Formation
B.Amazon OpenSearch Service
C.AWS Data Exchange
D.Amazon QuickSight
Show answerAnswer
A. AWS Lake Formation
AWS Lake Formation is a fully managed service that helps build, secure, and manage data lakes. It extends the AWS Glue Data Catalog with fine-grained access control, data quality blueprints, and integrates with analytics services, fulfilling all requirements.
8. A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time streaming data from various sources and store it in Amazon S3. To ensure data quality and schema enforcement for downstream analytics, the engineer needs a mechanism to validate incoming data against predefined schemas and manage metadata. Which AWS service combination should be used?
Data Storage and Management
A.Amazon DynamoDB for metadata management, AWS Lambda for schema validation, and Amazon Kinesis for streaming.
B.Amazon EFS for storage, AWS Systems Manager Parameter Store for schema, and Amazon MSK for streaming.
C.Amazon Redshift for storage, AWS Lake Formation for permissions, and AWS Data Pipeline for schema enforcement.
D.Amazon S3 for storage, AWS Glue Data Catalog for metadata management, and AWS Glue crawlers for schema discovery.
Show answerAnswer
D. Amazon S3 for storage, AWS Glue Data Catalog for metadata management, and AWS Glue crawlers for schema discovery.
AWS Glue Data Catalog acts as a central metadata repository, and AWS Glue Crawlers automatically infer schemas from data in S3, storing them in the catalog. This enables schema validation and governance for data lake analytics.
9. A global e-commerce company stores customer order data in an Amazon S3 bucket. The data is initially accessed frequently for the first 30 days for order processing and fraud detection. After 30 days, access decreases significantly, but the data must be retained for compliance reasons for 7 years. Retrieval should still be possible within minutes. Which S3 Lifecycle policy transition strategy minimizes costs while meeting these requirements?
Data Storage and Management
A.Transition to S3 One Zone-IA after 30 days, then delete after 7 years.
B.Transition to S3 Intelligent-Tiering after 30 days, then to S3 Glacier Instant Retrieval after 90 days.
C.Transition to S3 Standard-IA after 30 days, then to S3 Glacier Flexible Retrieval after 90 days.
D.Transition to S3 Glacier after 30 days, then to S3 Glacier Deep Archive after 90 days.
Show answerAnswer
C. Transition to S3 Standard-IA after 30 days, then to S3 Glacier Flexible Retrieval after 90 days.
S3 Standard-IA is cost-effective for data accessed less frequently but requiring rapid access. After 90 days, S3 Glacier Flexible Retrieval offers even lower costs for archival with retrieval times of minutes to hours, meeting the 7-year retention and 'minutes' retrieval requirement (Standard retrieval is within 3-5 hours, Expedited within 1-5 minutes).
10. A data engineer is designing a highly available and durable storage solution for critical business documents. These documents are infrequently accessed (once a quarter) but must be retrieved within seconds when needed. The budget is a significant constraint, and the solution needs to minimize storage costs while meeting the retrieval performance. Which S3 storage class is most suitable?
Data Storage and Management
A.Amazon S3 Glacier Deep Archive
B.Amazon S3 Standard
C.Amazon S3 Glacier Instant Retrieval
D.Amazon S3 Intelligent-Tiering
Show answerAnswer
C. Amazon S3 Glacier Instant Retrieval
Amazon S3 Glacier Instant Retrieval is designed for archived data that needs millisecond retrieval. It offers lower storage costs than S3 Standard or Standard-IA, making it ideal for infrequently accessed data that requires rapid access, balancing cost and performance for this specific use case.
11. A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time data from various sources, and multiple analytics teams will use different tools (e.g., Amazon Athena, Amazon Redshift Spectrum, AWS Glue ETL, Amazon EMR) to query and process this data. The engineer needs a centralized repository for table definitions, schema versions, and partition information that all these services can access consistently. Which AWS service provides this foundational metadata management capability?
Data Storage and Management
A.Amazon S3
B.AWS Glue Data Catalog
C.Amazon DynamoDB
D.AWS CloudFormation
Show answerAnswer
B. AWS Glue Data Catalog
The AWS Glue Data Catalog is a centralized, persistent metadata repository for all your data assets, regardless of where they are stored. It allows various AWS analytics services to discover, query, and process data consistently by providing table definitions, schema versions, and partition information. This is foundational for a data lake where multiple tools need a unified view of the data.
12. A data engineer is designing a data lake solution that ingests raw log files from various sources. These logs contain sensitive customer information and must be encrypted at rest and in transit. The solution needs to support querying the data directly using SQL-like queries without loading it into a database. Which combination of AWS services best meets these requirements?
Data Storage and Management
A.AWS DataSync for ingestion, Amazon S3 for storage with SSE-KMS encryption, and Amazon Athena for querying.
B.AWS Glue for ETL, Amazon DynamoDB for storage, and AWS QuickSight for reporting.
C.Amazon Kinesis Data Firehose for ingestion, Amazon RDS for storage, and Amazon Redshift for querying.
D.AWS Glue for ETL, Amazon S3 for storage with SSE-S3 encryption, and Amazon Athena for querying.
Show answerAnswer
A. AWS DataSync for ingestion, Amazon S3 for storage with SSE-KMS encryption, and Amazon Athena for querying.
AWS DataSync can securely ingest data. Amazon S3 with SSE-KMS provides encryption at rest with KMS key management. Amazon Athena allows direct SQL querying of data in S3, making it suitable for serverless analytics on log files.
13. A data engineering team manages a large data lake on Amazon S3. They use AWS Glue Data Catalog to store metadata for tables, which are partitioned by `year/month/day`. They observe that their Athena queries against these tables are performing poorly and scanning much more data than expected, even when filtering on specific dates. Upon investigation, they find that the Glue Data Catalog is not accurately reflecting the actual partitions on S3. What is the most efficient way to refresh the Glue Data Catalog to ensure partitions are correctly recognized for improved query performance?
Data Storage and Management
A.Use `MSCK REPAIR TABLE` in Athena or Presto to update the Glue Data Catalog for the table.
B.Implement a Lambda function triggered by S3 events to call `aws glue batch-create-partition` for new folders.
C.Manually add each new partition using the Glue console or AWS CLI `aws glue create-partition` command.
D.Run an AWS Glue Crawler on the S3 path of the table to discover new partitions.
Show answerAnswer
A. Use `MSCK REPAIR TABLE` in Athena or Presto to update the Glue Data Catalog for the table.
`MSCK REPAIR TABLE` is a command available in Athena and Presto that scans the S3 location of a table and adds any newly discovered partitions to the AWS Glue Data Catalog. This is often the most efficient way to refresh partitions for existing tables, especially when dealing with many new partitions, without needing to run a full Glue Crawler or develop custom Lambda functions.
14. A data engineering team is designing a highly available and durable storage solution for frequently accessed, structured data that requires ACID transaction support. The data will be used by multiple applications for real-time analytics and operational reporting, and the team needs to scale read replicas independently. Which AWS service is best suited for this requirement?
Data Storage and Management
A.Amazon RDS for PostgreSQL
B.Amazon Aurora
C.Amazon S3 Glacier Deep Archive
D.Amazon DynamoDB
Show answerAnswer
B. Amazon Aurora
Amazon Aurora is a MySQL and PostgreSQL-compatible relational database built for the cloud, combining the performance and availability of traditional enterprise databases with the simplicity and cost-effectiveness of open-source databases. It offers high availability, ACID compliance, and the ability to scale read replicas independently, making it ideal for the described use case.
15. A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time streaming data from IoT devices, batch data from on-premises systems, and semi-structured logs from web applications. The team needs a centralized metadata repository that can automatically discover schemas, track data lineage, and be queried by various analytics services like Amazon Athena and Amazon Redshift Spectrum. Which AWS service should the data engineer use to meet these requirements?
Data Storage and Management
A.Amazon DynamoDB
B.Amazon S3
C.AWS Glue Data Catalog
D.AWS Lake Formation
Show answerAnswer
C. AWS Glue Data Catalog
AWS Glue Data Catalog is a fully managed metadata repository that allows for schema discovery, data lineage tracking, and integration with various AWS analytics services. It serves as a central catalog for all data assets in a data lake.
16. A data engineering team is migrating an on-premises data warehouse to Amazon Redshift. The source system has a customer table with a `customer_id` column that is frequently used in JOIN operations with other large fact tables. To optimize query performance in Redshift, particularly for queries involving this `customer_id`, which distribution style should be applied to the customer table?
Data Storage and Management
A.DISTSTYLE EVEN
B.DISTSTYLE ALL
C.DISTSTYLE KEY
D.DISTSTYLE AUTO
Show answerAnswer
C. DISTSTYLE KEY
When a table is frequently joined with other large tables on a specific column, using DISTSTYLE KEY on that column (the join key) ensures that rows with matching join key values are stored on the same compute node. This minimizes data movement across the network during joins, significantly improving query performance.
17. A data engineer is designing a data lake using Amazon S3. The data consists of large files (several GB each) that are appended daily, but historical data is rarely updated. The team needs to optimize for query performance and cost efficiency when using services like Amazon Athena. Which partitioning strategy should the data engineer recommend?
Data Storage and Management
A.Partition by file size (e.g., /data/small/, /data/medium/, /data/large/)
B.No partitioning, store all files in a single bucket prefix
C.Partition by year and month (e.g., /data/year=YYYY/month=MM/)
D.Partition by data type (e.g., /data/csv/, /data/json/, /data/parquet/)
Show answerAnswer
C. Partition by year and month (e.g., /data/year=YYYY/month=MM/)
Partitioning by year and month aligns with time-series data access patterns, allowing Athena to scan only relevant subsets of data, significantly improving query performance and reducing costs by minimizing data scanned.
18. A data engineering team is building a data pipeline that ingests continuous streams of clickstream data from a website. This data needs to be processed in real-time for immediate analytics and then stored for batch processing later. The solution must be able to handle fluctuating data volumes, scale automatically, and deliver data to multiple consumers while ensuring data durability. Which AWS service is best suited for ingesting and temporarily storing this real-time streaming data?
Data Storage and Management
A.Amazon DynamoDB
B.Amazon S3
C.Amazon Kinesis Data Streams
D.Amazon SQS
Show answerAnswer
C. Amazon Kinesis Data Streams
Amazon Kinesis Data Streams is a highly scalable and durable real-time data streaming service. It is designed to continuously capture and store gigabytes per second of data, making it ideal for ingesting clickstream data, handling fluctuating volumes, and delivering data to multiple real-time and batch consumers.
19. A global manufacturing company collects sensor data from thousands of IoT devices. This data is time-series in nature, high-volume (petabytes annually), and requires near real-time ingestion and analysis. The data will be used for operational dashboards, anomaly detection, and machine learning models. Which AWS database service is purpose-built to handle these requirements efficiently?
Data Storage and Management
A.Amazon RDS for MySQL
B.Amazon Timestream
C.Amazon DynamoDB
D.Amazon Neptune
Show answerAnswer
B. Amazon Timestream
Amazon Timestream is a purpose-built time-series database that efficiently handles high-volume, time-stamped data, optimizing for ingestion, storage, and querying of time-series data at scale for applications like IoT analytics.
20. A data engineering team is building a new data lake on AWS. They need to store petabytes of raw, unstructured data from various sources, including application logs, social media feeds, and sensor data. The data needs to be highly available, durable, and cost-effective for long-term storage, with occasional access for analytics. Which AWS service is the most appropriate for this primary storage layer?
Data Storage and Management
A.Amazon EBS
B.Amazon S3
C.Amazon RDS
D.Amazon EFS
Show answerAnswer
B. Amazon S3
Amazon S3 is the foundational storage service for data lakes, offering virtually unlimited scalability, high durability, and cost-effectiveness for storing unstructured data. It supports various access patterns, from frequent to archival, making it ideal for raw data storage.
21. A data engineering team is designing a data lake using Amazon S3. The data consists of large files, and they want to improve query performance and reduce the amount of data scanned by analytical engines like Amazon Athena. They decide to organize their data in S3 using a directory structure that reflects the data's characteristics. Which strategy should they employ?
Data Storage and Management
A.Encrypt all S3 objects using server-side encryption with S3-managed keys (SSE-S3).
B.Store all data in a single flat directory.
C.Partition the data by relevant columns such as `year`, `month`, and `day`.
D.Use S3 bucket versioning for all data files.
Show answerAnswer
C. Partition the data by relevant columns such as `year`, `month`, and `day`.
Partitioning data in S3 by relevant columns (e.g., year, month, day) creates a hierarchical directory structure. Analytical engines like Athena can then use these partitions to prune data, scanning only the necessary subsets of data, which significantly improves query performance and reduces costs.
22. A financial services company needs to store transactional data from various source systems. This data is critical, highly structured, and requires strict ACID (Atomicity, Consistency, Isolation, Durability) compliance. The data volume is expected to grow to several terabytes, and the application requires high-performance, low-latency queries and strong transactional integrity. Which AWS managed service is the most appropriate for this requirement?
Data Storage and Management
A.Amazon S3
B.Amazon Aurora (PostgreSQL-compatible)
C.Amazon DynamoDB
D.Amazon Redshift
Show answerAnswer
B. Amazon Aurora (PostgreSQL-compatible)
Amazon Aurora is a highly performant, fully managed relational database service that offers strong ACID compliance, high availability, and scalability, making it ideal for critical transactional workloads that require strong data integrity and low-latency access to structured data. The PostgreSQL-compatible edition provides familiar SQL capabilities.
23. A data engineering team is building a real-time analytics pipeline for customer clickstream data. The data arrives in high volume and velocity and needs to be ingested, processed, and made available for querying within seconds. The solution must be fully managed and scale automatically to handle fluctuating traffic. Which AWS service is best suited for ingesting this continuous stream of data?
Data Storage and Management
A.AWS Database Migration Service (DMS)
B.AWS Transfer Family
C.Amazon SQS
D.Amazon Kinesis Data Streams
Show answerAnswer
D. Amazon Kinesis Data Streams
Amazon Kinesis Data Streams is a fully managed, scalable service for real-time processing of large streams of data records. It can continuously capture and store GBs of data per second from hundreds of thousands of sources, making it ideal for high-volume, real-time ingestion.
24. A financial services company needs to store transactional data from various source systems. This data is highly structured, requires ACID compliance, and must support complex SQL queries for reporting and analytics. The data volume is expected to grow to several terabytes within a year, and high availability is paramount. Which AWS service is the most suitable choice?
Data Storage and Management
A.Amazon Aurora PostgreSQL-Compatible Edition
B.Amazon DynamoDB
C.Amazon S3 Glacier Deep Archive
D.Amazon Redshift
Show answerAnswer
A. Amazon Aurora PostgreSQL-Compatible Edition
Amazon Aurora PostgreSQL-Compatible Edition provides a highly available, ACID-compliant relational database that supports complex SQL queries, making it ideal for structured transactional data with high availability requirements.
25. A data engineering team is designing a data warehouse using Amazon Redshift. They have a large fact table containing billions of records, which will be frequently joined with a smaller dimension table. To optimize query performance for these joins, the team needs to choose an appropriate distribution style for the fact table. Which distribution style is most suitable for this scenario?
Data Storage and Management
A.ALL distribution
B.EVEN distribution
C.KEY distribution
D.AUTO distribution
Show answerAnswer
C. KEY distribution
KEY distribution is ideal when a large fact table is frequently joined with a smaller dimension table on a common key. Distributing data based on the join key ensures that matching rows are co-located on the same compute nodes, minimizing data transfer during query execution.
A DDL command in Amazon Athena (and Hive) used to update the AWS Glue Data Catalog with new partitions that have been manually added to the S3 location of a table.
Discovers and adds new partitions to the Glue Data Catalog
Ensures Athena can query newly added data
Alternative to running a Glue Crawler for partition updates
Amazon EFS provides scalable, shared, and persistent file storage that can be mounted by AWS Lambda functions, enabling them to process large files and share data across invocations.
S3 data partitioning organizes data in S3 buckets by creating a folder structure based on one or more column values (e.g., `s3://bucket/key=value/`). This allows analytics engines to prune data, scanning only relevant subsets, which significantly improves query performance and reduces costs.
Organizes data into logical folders
Improves query performance by reducing scanned data
A Redshift table distribution style that copies the entire table to every compute node, typically used for small dimension tables to optimize join performance.
Copies full table to all compute nodes
Eliminates data movement for joins with large fact tables
Best for small dimension tables (few hundred MBs/thousands of rows)
Amazon Redshift can encrypt data at rest using AWS KMS for key management and enforce SSL/TLS for data in transit, ensuring comprehensive protection for sensitive data.
Data at rest encrypted by default
Use KMS for customer-managed keys, rotation, and auditing
Apache Parquet is a free and open-source columnar storage format designed for efficient data storage and retrieval. It is optimized for analytical workloads, offering high compression ratios, improved query performance, and support for complex nested data structures and schema evolution.
AWS Lake Formation is a fully managed service that helps you build, secure, and manage data lakes. It simplifies the process of building data lakes by centralizing security, governance, and cataloging, offering fine-grained access control to data.
Simplifies building and securing data lakes
Extends AWS Glue Data Catalog for governance
Provides fine-grained access control (table/column-level)
AWS Glue Data Catalog is a persistent metadata store for your data assets, serving as a central repository for table definitions, schema versions, and other metadata for data lakes and data warehousing.
S3 Lifecycle policies define rules to automatically transition objects to different S3 storage classes or expire them, based on age or other criteria. This helps optimize storage costs by moving data to more cost-effective classes as access patterns change.
Amazon S3 Glacier Instant Retrieval is an archive storage class that delivers the lowest cost storage for data that is rarely accessed but requires millisecond retrieval. It's designed for use cases that demand immediate access to archived data, such as medical images, news media assets, or backup data.
Amazon Athena is an interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It's serverless, so there's no infrastructure to manage, and you pay only for the queries you run.
A MySQL and PostgreSQL-compatible relational database built for the cloud, combining the performance and availability of traditional enterprise databases with the simplicity and cost-effectiveness of open-source databases.
A Redshift table distribution style that distributes rows based on the hash of values in a specified column, commonly used to co-locate data for efficient joins.
Optimizes join performance
Co-locates matching join keys on the same compute node
A real-time data streaming service that can continuously capture and store large streams of data records for real-time analytics and subsequent batch processing.
Amazon Timestream is a fast, scalable, and serverless time-series database service for IoT and operational applications, designed to store and analyze time-series data efficiently at petabyte scale.
Amazon S3 (Simple Storage Service) is the primary storage layer for data lakes on AWS, providing highly scalable, durable, and cost-effective object storage.
Supports virtually unlimited data storage.
Offers 11 nines of durability.
Various storage classes for different access patterns and costs.
Amazon Redshift KEY distribution ensures that rows with the same value in the specified distribution column are stored on the same compute node. This is crucial for optimizing join performance between large fact tables and dimension tables.
Distributes data based on a column's value
Optimizes join performance by co-locating data
Used for large fact tables joined with dimension tables
Amazon S3 Versioning allows you to keep multiple versions of an object in the same bucket, providing protection against accidental overwrites, deletions, and enabling easy recovery of previous object states.
Amazon S3 (Simple Storage Service) is an object storage service offering industry-leading scalability, data availability, security, and performance. It is ideal for storing unstructured data like images, videos, and backups.
AWS Glue ETL jobs are serverless Apache Spark-based jobs that enable data engineers to extract, transform, and load data at scale, automatically handling provisioning, setup, and scaling of compute resources.
S3 Lifecycle policies define rules to automatically transition objects to different S3 storage classes or expire them after a specified time, optimizing storage costs based on access patterns.
A fast compression/decompression algorithm optimized for speed rather than maximum compression, commonly used in big data ecosystems for analytical workloads.
Optimized for speed (fast compression/decompression)
Amazon Simple Storage Service (S3) Standard is an object storage service offering high durability, availability, and performance for frequently accessed data.
Server-Side Encryption for Amazon S3 objects using keys managed by AWS Key Management Service (KMS), providing customer control and auditing of encryption keys.
Questions are original practice items written to match the published exam objectives. Step2Study is not affiliated with or endorsed by any certification body.