AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsHard
A team is building a machine learning model to detect anomalies in network traffic, identifying unusual patterns that might indicate a cyberattack. They have a massive dataset of normal network traffic but very few examples of actual cyberattacks. Which type of learning approach is most suitable for this scenario, given the imbalanced and rare nature of the 'attack' class?
- AReinforcement Learning
- BUnsupervised Learning
- CSupervised Learning
- DSemi-Supervised Learning
Show answer & explanationAnswer & explanation
Correct answer: B. Unsupervised Learning
Given a massive dataset of 'normal' traffic and very few 'attack' examples, unsupervised learning is highly suitable for anomaly detection. Algorithms like K-Means, Isolation Forest, or Autoencoders can learn the patterns of normal data and then flag any data points that deviate significantly from these learned patterns as anomalies, without needing explicit labels for the rare attack class during training.
Why the other options are wrong
- A. Reinforcement learning involves agents learning through trial and error in an environment, which is not directly applicable to identifying rare patterns in existing data.
- C. Supervised learning requires a substantial amount of labeled data for both normal and anomaly classes, which is not available for the 'attack' class here.
- D. Semi-supervised learning uses a small amount of labeled data and a large amount of unlabeled data, but for rare anomaly detection, unsupervised methods that only learn 'normal' are often more robust.
Unsupervised Learning for Anomaly Detection
A machine learning approach where models learn patterns from unlabeled data, typically identifying anomalies as data points that deviate significantly from the learned 'normal' patterns.
- Does not require labeled anomaly data for training.
- Learns the underlying structure of the majority (normal) class.
- Effective for rare events or when labeling is difficult/expensive.
- Examples: Clustering, Density Estimation, Autoencoders.
Memory trick: ML 'learns' from data, either 'with' or 'without' a teacher.