CompTIA Data+ (DA0-002)Data MiningMedium

A data analyst is performing an exploratory data analysis on a customer feedback dataset. They observe that the 'Feedback_Text' column is often filled with short, informal phrases and sometimes includes emoticons and repetitive punctuation (e.g., 'Great!!!', 'So happy :)'). Before applying sentiment analysis, these elements need to be handled. Which data cleansing technique specifically targets removing or standardizing such informal text features?

  1. AStemming
  2. BTokenization
  3. CText normalization
  4. DNumerical encoding
Show answer & explanation

Correct answer: C. Text normalization

Text normalization is the overarching technique that encompasses various steps like lowercasing, removing punctuation, special characters, and standardizing informal elements such as emoticons or repetitive characters. This ensures text data is consistent and clean for subsequent natural language processing tasks like sentiment analysis.

Why the other options are wrong

  • A. Stemming reduces words to their root form (e.g., 'running' to 'run'), which is different from handling emoticons or repetitive punctuation.
  • B. Tokenization breaks text into words or phrases, but doesn't remove or standardize informal features.
  • D. Numerical encoding converts text to numbers, but doesn't clean informal text features.

Text Normalization

A crucial step in natural language processing (NLP) that transforms raw text into a standardized and clean format, making it more suitable for analysis. It includes tasks like lowercasing, removing punctuation, special characters, and standardizing informal language.

  • Aims to reduce variability in text data.
  • Essential for tasks like sentiment analysis, topic modeling, and text classification.
  • Often involves regular expressions for pattern-based cleaning.

Memory trick: Cleaning text for NLP is like polishing a gem: normalization makes it shine.

More Data Mining questions