A financial application needs to store transaction details in Azure Cosmos DB. Each transaction has a unique TransactionId, a CustomerId, and a TransactionDate. Queries will frequently retrieve all transactions for a specific CustomerId, sometimes filtered by TransactionDate. Point reads by TransactionId are also common. To optimize for these query patterns, which partitioning strategy should be implemented?
- APartition by TransactionDate
- BPartition by TransactionId
- CPartition by CustomerId
- DPartition by a composite key of CustomerId and TransactionDate
Show answer & explanationAnswer & explanation
Correct answer: C. Partition by CustomerId
Partitioning by CustomerId allows for efficient retrieval of all transactions for a specific customer, as all relevant data for a customer will reside in the same logical partition. Filtering by TransactionDate within that customer's partition is also efficient. While TransactionId is unique, partitioning by it would lead to many small partitions and inefficient range queries for a customer. Composite keys can lead to hot partitions if not carefully designed for query patterns.
Why the other options are wrong
- A. Partitioning by TransactionDate would group all transactions from all customers on a specific date, making it inefficient to retrieve all transactions for a single customer across multiple dates.
- B. Partitioning by TransactionId would lead to a very high number of partitions, making range queries for a CustomerId inefficient as data would be scattered.
- D. A composite key of CustomerId and TransactionDate could be viable but might lead to hot partitions if a single customer has many transactions on a specific date, or make range queries across dates for a customer less optimal than just using CustomerId.
Cosmos DB Partitioning Strategy
Partitioning in Azure Cosmos DB distributes data across logical and physical partitions based on a partition key. Choosing an effective partition key is crucial for query performance, scalability, and cost optimization, as it determines how data is distributed and how requests are routed.
- Partition key determines logical partitions.
- Data within a logical partition is physically stored together.
- Queries targeting the partition key are most efficient.
- Aim for high cardinality and even distribution of data.
- Avoid 'hot partitions' where one partition receives disproportionate requests.
Memory trick: Partition for Primary Query Patterns.