AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium

A data engineering team is building a new feature store to manage features for various machine learning models across their organization. They need a solution that can serve both online inference requests (low-latency, real-time access) and offline training jobs (high-throughput, batch access). Additionally, the solution must provide versioning of features and allow for easy discovery and reuse. Which AWS service or feature is best suited for this requirement?

  1. AAmazon SageMaker Feature Store.
  2. BAmazon Redshift for offline features and Redis for online features.
  3. CAWS Glue Data Catalog with Amazon Athena.
  4. DAmazon DynamoDB for online features and Amazon S3 for offline features.
Show answer & explanation

Correct answer: A. Amazon SageMaker Feature Store.

Amazon SageMaker Feature Store is purpose-built for managing, storing, and serving machine learning features for both online inference and offline training. It offers low-latency access for real-time predictions and high-throughput access for batch processing, along with capabilities for feature versioning, discovery, and reuse across multiple models and teams.

Why the other options are wrong

  • B. This combination, like option A, requires custom integration and lacks the integrated management, versioning, and discovery capabilities of a dedicated feature store.
  • C. AWS Glue Data Catalog and Athena are excellent for data discovery and querying but do not provide the low-latency online serving capabilities or native feature versioning of a feature store.
  • D. While technically possible, this approach requires significant custom integration, synchronization, and management overhead compared to a dedicated feature store solution.

SageMaker Feature Store

A fully managed service that provides a centralized repository for creating, storing, and serving machine learning features for both online inference and offline training.

  • Supports both online (low-latency) and offline (high-throughput) access.
  • Enables feature versioning, discovery, and reuse.
  • Improves MLOps efficiency and consistency across models.

Memory trick: Features Unified, SageMaker's Store, ML Data No More Chore.

More Machine Learning Implementation and Operations questions