AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisHard

A data scientist is analyzing sensor data from industrial machinery. The data contains several time-series features, and they suspect that some sensors occasionally report erroneous, extremely high values that are physically impossible and occur as sudden, isolated spikes. These spikes are not part of any normal operational variance and can severely distort subsequent analysis. Which advanced outlier detection technique is best suited to identify these specific types of outliers in time-series data?

  1. AMoving average-based anomaly detection, comparing each point to its local moving average and standard deviation.
  2. BInterquartile Range (IQR) rule, applied to the entire time series.
  3. CIsolation Forest, due to its effectiveness with high-dimensional data and non-parametric nature.
  4. DZ-score method, applied to each data point individually.
Show answer & explanation

Correct answer: A. Moving average-based anomaly detection, comparing each point to its local moving average and standard deviation.

For time-series data with sudden, isolated spikes, a moving average-based anomaly detection method is highly effective. It establishes a 'normal' range based on recent past values (local context) and identifies points that deviate significantly from this local norm, thus capturing transient spikes that global methods like Z-score or IQR might miss or misinterpret.

Why the other options are wrong

  • B. Incorrect. The IQR rule is robust to global skewness but doesn't inherently account for the temporal context of time-series data, making it less effective for identifying transient, localized spikes compared to methods that consider sequence.
  • C. Incorrect. Isolation Forest is powerful for multivariate outlier detection and high-dimensional data but might not be the most targeted or efficient approach for simple, isolated spikes in a single time-series feature, where local temporal context is key.
  • D. Incorrect. The Z-score method is sensitive to the global mean and standard deviation, which can be heavily influenced by the very spikes one is trying to detect, leading to masking or reduced sensitivity for local anomalies.

Time-Series Anomaly Detection (Moving Average)

A technique for identifying anomalous data points in time-series data by comparing each point to a statistically derived 'normal' range (e.g., mean ± std dev) calculated from a preceding moving window of data.

  • Considers temporal context (local patterns).
  • Effective for detecting sudden spikes or drops.
  • Adapts to changing baseline trends over time.

Memory trick: Moving average sees the local flow, spikes above it must go.

More Exploratory Data Analysis questions