Microsoft Azure Data FundamentalsDescribe an analytics workload on AzureMedium
A financial institution is building an analytics solution to detect fraudulent transactions. They have petabytes of historical transaction data stored in Azure Data Lake Storage Gen2. Data scientists need to run complex machine learning models, which require distributed processing frameworks like Apache Spark, on this large dataset. Which Azure service is most suitable for this workload?
- AAzure Data Explorer
- BAzure Analysis Services
- CAzure Databricks
- DAzure SQL Database
Show answer & explanationAnswer & explanation
Correct answer: C. Azure Databricks
Azure Databricks is an Apache Spark-based analytics platform optimized for big data processing and machine learning workloads, making it ideal for data scientists working with petabytes of data in a data lake.
Why the other options are wrong
- A. Azure Data Explorer is optimized for high-performance ingestion and querying of telemetry and time-series data, not for general-purpose distributed ML with Spark.
- B. Azure Analysis Services is an analytical data engine for semantic models and reporting, not for distributed machine learning model training.
- D. Azure SQL Database is a relational database and not designed for petabyte-scale distributed processing with Spark for machine learning.
Azure Databricks
An Apache Spark-based analytics platform for data science, machine learning, and big data engineering.
- Optimized for large-scale data processing and ML.
- Supports Python, Scala, R, Java, and SQL.
- Integrates deeply with Azure services like Data Lake Storage.
Memory trick: Databricks builds big brains on big data.