A data engineering team is troubleshooting a Dataflow job that processes real-time sensor data. Users are reporting that some data points appear to be missing from the final aggregated results in BigQuery, particularly during periods of high data volume. The Dataflow job itself reports no errors and appears to be running, but the output is incomplete. The team suspects that some data elements might be getting dropped before they even reach the Dataflow pipeline or are being lost in transit. Which initial step should the team take to investigate potential data loss at the ingestion layer?
- ACheck Dataflow job metrics for worker CPU utilization and memory usage.
- BReview BigQuery `INFORMATION_SCHEMA` for table write errors.
- CAnalyze Cloud Logging for Dataflow worker logs for `ArrayIndexOutOfBoundsException`.
- DExamine Pub/Sub subscription metrics for `oldest_unacked_message_age` and `num_unacked_messages`.
Show answer & explanationAnswer & explanation
Correct answer: D. Examine Pub/Sub subscription metrics for `oldest_unacked_message_age` and `num_unacked_messages`.
If data is missing from the output and the Dataflow job itself reports no errors, the problem often lies in the ingestion layer. For Pub/Sub-based ingestion, checking `oldest_unacked_message_age` (how long the oldest message is waiting to be acknowledged) and `num_unacked_messages` (count of messages not yet acknowledged) in Cloud Monitoring for the Pub/Sub subscription is critical. High values for these metrics indicate that data is accumulating in Pub/Sub and not being consumed by Dataflow quickly enough, potentially leading to message expiration and loss.
Why the other options are wrong
- A. Dataflow worker metrics are for pipeline processing performance; if data isn't reaching the pipeline, these won't show ingestion issues.
- B. BigQuery `INFORMATION_SCHEMA` or write errors would indicate problems *after* Dataflow processing, not data loss *before* it reaches the pipeline.
- C. Dataflow worker logs for specific exceptions indicate issues *within* the pipeline, but the problem description suggests data is missing *before* or *during* ingestion into Dataflow.
Pub/Sub Ingestion Monitoring
Monitoring key Pub/Sub subscription metrics to detect if messages are being consumed effectively by downstream services like Dataflow, preventing data loss.
- `oldest_unacked_message_age` indicates message backlog
- `num_unacked_messages` shows pending messages
- High values suggest consumer lag or issues
Memory trick: Follow the data flow; check ingestion first for missing pieces.