Professional Data EngineerBuilding and operationalizing data processing systemsHard
A data engineer is designing a real-time analytics pipeline using Cloud Pub/Sub and Dataflow. The pipeline needs to process millions of messages per second, and the incoming data stream has highly variable throughput. To ensure message delivery and prevent data loss, the Pub/Sub topic configuration needs to be optimized for high volume and bursty traffic. Which Pub/Sub feature should be enabled to handle this scenario effectively?
- ASchema enforcement with Avro format
- BMessage retention duration set to 1 day
- CMessage ordering enabled on the subscription
- DAutomatic scaling of throughput capacity
Show answer & explanationAnswer & explanation
Correct answer: D. Automatic scaling of throughput capacity
Cloud Pub/Sub automatically scales its throughput capacity to handle varying message volumes, including massive bursts, without manual intervention or configuration, ensuring messages are ingested reliably and preventing data loss.
Why the other options are wrong
- A. Schema enforcement ensures data quality but doesn't directly address the challenge of high volume and bursty traffic ingestion.
- B. Message retention duration affects how long messages are kept, not how efficiently they are ingested during high bursts.
- C. Message ordering ensures messages are delivered in the order they were published but can sometimes slightly reduce throughput, not enhance it for bursts.
Cloud Pub/Sub Automatic Scaling
Cloud Pub/Sub automatically scales its throughput capacity to accommodate varying message volumes, from low rates to millions of messages per second, without requiring users to pre-provision or manage resources.
- Managed service, no servers to provision
- Automatically adjusts to traffic spikes
- Ensures high availability and durability
Memory trick: Pub/Sub: The elastic highway for messages.