Microsoft Azure Data FundamentalsDescribe an analytics workload on AzureMedium
A data analytics team needs to perform ad-hoc queries on petabytes of data stored in Azure Data Lake Storage Gen2. The data is in various formats, including Parquet, CSV, and JSON. They need a service that allows them to query this data without moving or transforming it, and they want to pay only for the queries they run. Which Azure service should they use?
- AAzure Databricks
- BAzure Synapse Analytics dedicated SQL pool
- CAzure SQL Database
- DAzure Synapse Analytics serverless SQL pool
Show answer & explanationAnswer & explanation
Correct answer: D. Azure Synapse Analytics serverless SQL pool
Azure Synapse Analytics serverless SQL pool allows querying data directly in Azure Data Lake Storage Gen2 using T-SQL, without provisioning resources, and charges only for the data processed per query, fitting the ad-hoc, pay-per-query requirement.
Why the other options are wrong
- A. Azure Databricks is a unified analytics platform for more complex data engineering and machine learning, often more expensive for simple ad-hoc queries.
- B. Dedicated SQL pool requires provisioning and continuous billing, not ideal for ad-hoc, pay-per-query scenarios.
- C. Azure SQL Database is a relational database and requires data to be loaded into it, not for querying data directly in a data lake.
Azure Synapse Serverless SQL Pool
A serverless query service in Azure Synapse Analytics that enables users to analyze data directly in Azure Data Lake Storage Gen2 using T-SQL.
- Queries data without moving or transforming it.
- Pay-per-query model based on data processed.
- Supports various file formats like Parquet, CSV, and JSON.
Memory trick: Serverless queries explore the data lake.