A data engineer is designing a new semantic model in Microsoft Fabric. The model will consume data from a large Azure Synapse Analytics dedicated SQL pool. The reporting requirements dictate that users need to perform ad-hoc analysis on the full dataset, which is several terabytes in size, with near real-time latency. However, some aggregate reports only require daily snapshots of summarized data. Which storage mode should the data engineer primarily choose for this semantic model to meet these requirements efficiently?
- AImport mode
- BDual mode
- CDirectQuery mode
- DDirect Lake mode
Show answer & explanationAnswer & explanation
Correct answer: C. DirectQuery mode
DirectQuery mode is suitable for large datasets where near real-time data is required and the underlying data source can handle the query load. While it doesn't cache data, it directly queries the source, fulfilling the 'ad-hoc analysis on full dataset with near real-time latency' requirement. Import mode would not handle the multi-terabyte size efficiently for ad-hoc queries, and Direct Lake is optimized for Delta Lake tables, not dedicated SQL pools.
Why the other options are wrong
- A. Import mode caches data, which would be inefficient for a multi-terabyte dataset and would not provide near real-time latency without frequent, resource-intensive refreshes.
- B. Dual mode is a combination of Import and DirectQuery, but the primary requirement for ad-hoc analysis on the full multi-terabyte dataset with near real-time latency points to DirectQuery as the foundational choice.
- D. Direct Lake mode is specifically designed for querying Delta Lake tables in OneLake directly, not Azure Synapse Analytics dedicated SQL pools.
DirectQuery Mode
DirectQuery mode in Microsoft Fabric semantic models allows direct querying of the underlying data source without importing data into the model. This is ideal for very large datasets and scenarios requiring near real-time data.
- Data is not cached in the semantic model.
- Queries are translated into native queries for the source system.
- Good for large datasets and near real-time requirements.
- Performance depends heavily on the underlying data source.
Memory trick: Fabric's storage modes are a direct path to data or an imported treasure chest.