Microsoft Azure Data FundamentalsDescribe an analytics workload on AzureMedium
A retail company is analyzing customer purchase patterns. They have several terabytes of historical sales data stored in Azure Data Lake Storage Gen2 in Parquet format. They need to query this data ad-hoc to identify trends without needing to load it into a dedicated database or provision persistent compute resources. Which Azure Synapse Analytics component is best suited for this task?
- AData Explorer pool
- BServerless SQL pool
- CSpark pool
- DDedicated SQL pool
Show answer & explanationAnswer & explanation
Correct answer: B. Serverless SQL pool
Azure Synapse Analytics serverless SQL pool allows you to query data directly in your data lake using standard T-SQL syntax without provisioning or managing any dedicated resources. It's ideal for ad-hoc queries, data exploration, and logical data warehousing over large volumes of data in various formats like Parquet.
Why the other options are wrong
- A. Data Explorer pools (Kusto) are for log and telemetry analytics, not general-purpose data lake querying of Parquet files.
- C. Spark pools are for big data processing using Apache Spark, requiring cluster management and are not specifically designed for ad-hoc SQL querying of data lake files.
- D. Dedicated SQL pools require provisioning and management of compute resources and are best for pre-loaded, structured data warehouses.
Azure Synapse Serverless SQL Pool
A serverless query service in Azure Synapse Analytics that enables you to query data directly from Azure Data Lake Storage Gen2 using T-SQL.
- No infrastructure to set up or manage.
- Pay-per-query model based on data processed.
- Supports various file formats like Parquet, CSV, JSON.
Memory trick: Serverless lets you query the lake without a boat.