A media company uses a content management system (CMS) that stores large media files (up to 5 TB each) in Amazon S3. The CMS generates daily reports that require scanning the metadata of all objects in a specific S3 bucket, which can contain millions of objects. The current process uses S3 ListObjectsV2 API calls, which are slow and expensive for large buckets. The company wants a more efficient and cost-effective way to retrieve and query this object metadata. Which solution should a solutions architect recommend?
- AConfigure S3 Batch Operations to retrieve metadata for all objects concurrently.
- BUse Amazon S3 Inventory reports to export object metadata to a flat file, then query it with Amazon Athena.
- CImplement S3 Select to directly query the metadata within the S3 objects.
- DEnable S3 Event Notifications to trigger a Lambda function that stores metadata in a separate DynamoDB table.
Show answer & explanationAnswer & explanation
Correct answer: B. Use Amazon S3 Inventory reports to export object metadata to a flat file, then query it with Amazon Athena.
S3 Inventory reports provide a daily or weekly flat-file list of objects and their metadata for an S3 bucket, which is efficient for large-scale analysis. Storing these reports in S3 and querying them with Amazon Athena (serverless SQL) is a highly cost-effective and scalable way to analyze object metadata without making millions of individual API calls.
Why the other options are wrong
- A. S3 Batch Operations can perform operations on a list of S3 objects, but it's for performing actions (like copying, tagging, or restoring) on existing objects, not for efficiently querying metadata of millions of objects. It would still rely on generating a manifest, which is similar to the problem S3 Inventory solves more directly.
- C. S3 Select allows you to retrieve a subset of data from an S3 object (e.g., CSV, JSON, Parquet) by running SQL queries directly on the object. It's for querying *inside* an object, not for querying the metadata *of* millions of objects in a bucket.
- D. While this approach can work, setting up S3 Event Notifications for millions of objects (especially existing ones) and maintaining a DynamoDB table for metadata can become complex and potentially expensive. S3 Inventory is purpose-built for large-scale metadata collection.
Amazon S3 Inventory
Amazon S3 Inventory provides a flat-file list of your objects and their corresponding metadata (e.g., storage class, encryption status) for an S3 bucket or prefix. It's used for auditing and analysis of objects at scale.
- Generates daily or weekly reports of object metadata.
- Output stored as CSV, ORC, or Parquet files.
- Efficient for analyzing billions of objects.
- Can be queried using Amazon Athena or other analytics tools.
Memory trick: S3 Inventory is the 'Librarian' for your S3 objects, giving you a full catalog to search with Athena.