A company is developing a machine learning model to detect anomalies in sensor data from critical industrial equipment. False negatives (failing to detect an actual anomaly) are extremely costly, potentially leading to equipment failure. However, a high rate of false positives (triggering alerts for normal operation) is also undesirable as it leads to alert fatigue and unnecessary maintenance. Which model evaluation metric should the team primarily focus on optimizing while also keeping an eye on the secondary metric?
- APrimary: Accuracy, Secondary: F1-score
- BPrimary: Precision, Secondary: Recall
- CPrimary: F1-score, Secondary: Accuracy
- DPrimary: Recall, Secondary: Precision
Show answer & explanationAnswer & explanation
Correct answer: D. Primary: Recall, Secondary: Precision
Minimizing false negatives ('failing to detect an actual anomaly') directly corresponds to maximizing Recall. This is the primary objective due to the high cost of missed anomalies. However, a high rate of false positives ('alert fatigue') means Precision should also be monitored and kept at an acceptable level. Therefore, Recall is primary, and Precision is secondary.
Why the other options are wrong
- A. Accuracy can be misleading in highly imbalanced anomaly detection scenarios, and F1-score is a balance, not prioritizing one over the other.
- B. Prioritizing Precision would mean minimizing false positives, which is secondary here; missing actual anomalies (low Recall) is more costly.
- C. F1-score balances Precision and Recall, which might not be optimal when one metric (Recall) has a significantly higher cost associated with its negative outcome. Accuracy is often unsuitable for imbalanced anomaly detection.
Recall vs. Precision Trade-off
Recall measures the proportion of actual positives correctly identified (minimizing false negatives), while Precision measures the proportion of predicted positives that are actually correct (minimizing false positives). There is often a trade-off between these two metrics.
- High Recall is crucial when false negatives are expensive (e.g., fraud, disease detection).
- High Precision is crucial when false positives are expensive (e.g., spam detection, recommending wrong products).
- F1-score is a harmonic mean used when both are equally important.
- The business context dictates which metric to prioritize.
Memory trick: When 'COSTS ARE HIGH', choose your 'METRIC WISELY' to match the most critical error.