Microsoft Certified: Azure Developer Associate (AZ-204)Develop for Azure storageHard

A company stores customer order data in Azure Cosmos DB. Each order document contains a `customerId` and an `orderId`. The application frequently queries orders by a specific `customerId` and also needs to retrieve individual orders very quickly using both `customerId` and `orderId`. The development team needs to choose an optimal partitioning strategy to ensure low latency and cost-effective operations.

  1. APartition by `/customerId`.
  2. BPartition by `orderId`.
  3. CPartition by `/orderId` and use `customerId` as a secondary index.
  4. DPartition by a composite key of `/customerId` and `/orderId`.
Show answer & explanation

Correct answer: A. Partition by `/customerId`.

Partitioning by `/customerId` is optimal for this scenario. It ensures that all orders for a single customer reside on the same logical partition, making queries by `customerId` efficient (single-partition queries). When retrieving individual orders by `customerId` and `orderId`, specifying `customerId` as the partition key in the query allows Cosmos DB to quickly locate the correct logical partition, then efficiently find the `orderId` within that partition.

Why the other options are wrong

  • B. Partitioning by `orderId` would create many small partitions, and queries by `customerId` would be cross-partition, leading to higher RU costs and latency.
  • C. Cosmos DB does not support secondary indexes in the traditional sense; the partition key is primary. Using `customerId` as a secondary index isn't a valid partitioning strategy.
  • D. A composite partition key of `/customerId` and `/orderId` would be too granular for the `customerId` queries, potentially creating many single-item partitions and making `customerId`-only queries cross-partition and expensive.

Cosmos DB Partitioning Strategy (Single-Key)

Choosing an effective partition key in Azure Cosmos DB is crucial for performance and cost. A good partition key distributes data evenly and supports frequent query patterns by allowing single-partition queries, where all relevant data for a query resides within one logical partition.

  • Distributes data across logical and physical partitions.
  • Impacts query performance and cost (single- vs. cross-partition).
  • Should have high cardinality and spread request units evenly.

Memory trick: Partition by the popular path for performance.

More Develop for Azure storage questions