AWS Certified Machine Learning – SpecialtyData EngineeringEasy

A data scientist is working with a large dataset of customer reviews, which contains free-form text. Before feeding this data into a natural language processing (NLP) model, the text needs to be cleaned by removing special characters, converting to lowercase, tokenizing, and removing common stopwords. This preprocessing step must be easily repeatable and scalable. Which AWS service is best suited for performing these common text data preparation tasks in a managed and scalable way?

  1. AAWS Glue
  2. BAWS Textract
  3. CAmazon Comprehend
  4. DAmazon Personalize
Show answer & explanation

Correct answer: A. AWS Glue

AWS Glue is a serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development. It supports custom Python or Scala scripts (e.g., using NLTK or SpaCy) to perform complex text transformations like tokenization, stopword removal, and normalization on large datasets in a scalable manner.

Why the other options are wrong

  • B. AWS Textract is a service that automatically extracts text and data from scanned documents. It's for OCR, not for cleaning or transforming free-form text data.
  • C. Amazon Comprehend is an NLP service that provides pre-trained models for tasks like sentiment analysis, entity recognition, and topic modeling. It's for analyzing text, not for custom, granular text preprocessing before an NLP model.
  • D. Amazon Personalize is a machine learning service that makes it easy for developers to create personalized recommendations for their customers. It consumes data but does not perform general text preprocessing.

AWS Glue for Text Preprocessing

AWS Glue is a serverless ETL service that can be used to perform various data preparation tasks, including text cleaning and feature engineering, for machine learning workloads.

  • Serverless and scalable.
  • Supports custom Python/Scala scripts (e.g., with NLTK, SpaCy).
  • Ideal for large datasets stored in S3.
  • Automates schema discovery and job orchestration.
  • Integrates with other AWS services for ML pipelines.

Memory trick: Glue cleans text for ML, making it shiny and new.

More Data Engineering questions