Microsoft Azure Data FundamentalsDescribe an analytics workload on AzureMedium

A data analytics team needs to perform ad-hoc queries on petabytes of data stored in Azure Data Lake Storage Gen2. The data is in various formats, including Parquet, CSV, and JSON. They need a service that allows them to query this data without moving or transforming it, and they want to pay only for the queries they run. Which Azure service should they use?

  1. AAzure Databricks
  2. BAzure Synapse Analytics dedicated SQL pool
  3. CAzure SQL Database
  4. DAzure Synapse Analytics serverless SQL pool
Show answer & explanation

Correct answer: D. Azure Synapse Analytics serverless SQL pool

Azure Synapse Analytics serverless SQL pool allows querying data directly in Azure Data Lake Storage Gen2 using T-SQL, without provisioning resources, and charges only for the data processed per query, fitting the ad-hoc, pay-per-query requirement.

Why the other options are wrong

  • A. Azure Databricks is a unified analytics platform for more complex data engineering and machine learning, often more expensive for simple ad-hoc queries.
  • B. Dedicated SQL pool requires provisioning and continuous billing, not ideal for ad-hoc, pay-per-query scenarios.
  • C. Azure SQL Database is a relational database and requires data to be loaded into it, not for querying data directly in a data lake.

Azure Synapse Serverless SQL Pool

A serverless query service in Azure Synapse Analytics that enables users to analyze data directly in Azure Data Lake Storage Gen2 using T-SQL.

  • Queries data without moving or transforming it.
  • Pay-per-query model based on data processed.
  • Supports various file formats like Parquet, CSV, and JSON.

Memory trick: Serverless queries explore the data lake.

More Describe an analytics workload on Azure questions