Microsoft Certified: Fabric Analytics Engineer Associate flashcards
162 free flashcards. Tap a card to flip it.
PySpark Date Part Extraction
Flip cardThe process of deriving specific components (like year, month, day) from a date or timestamp column in a PySpark DataFrame using optimized built-in functions.
- Use `pyspark.sql.functions` for date/time functions.
- `year()` and `month()` are common functions.
- `withColumn()` is used to add or update columns.
- Avoid RDD operations for DataFrame transformations when possible.
Memory trick: Functions Extract Dates, Columns Get New Names, All in PySpark's Flow.
PySpark agg() function
Flip cardThe PySpark `.agg()` function is used with grouped DataFrames to apply one or more aggregate functions to the grouped data, producing summary statistics.
- Follows a `groupBy()` operation.
- Takes aggregate functions (e.g., `avg`, `sum`, `count`).
- Returns a DataFrame with aggregated results.
Memory trick: Group your data, then Aggregate the stats.
Delta Lake COPY INTO
Flip cardThe `COPY INTO` command provides an idempotent and fault-tolerant way to load data from external sources (like CSV, Parquet, JSON files) into Delta tables, supporting schema inference and evolution.
- Idempotent: can be run multiple times without duplicates.
- Supports various file formats.
- Handles schema evolution with `mergeSchema` option.
- Optimized for large-scale data ingestion.
Memory trick: COPY INTO brings files to Delta, smart and safe.
Micro-batching
Flip cardMicro-batching is a data processing technique where continuous streams of data are broken down into small, time-based batches, which are then processed as traditional batches, providing near real-time analytics capabilities.
- Balances latency and throughput for streaming data.
- Processes data in small, frequent intervals (e.g., seconds).
- Often used in Spark Streaming or similar frameworks.
Memory trick: Small batches, fast insights, micro-batching ignites!
Spark SQL DENSE_RANK()
Flip cardThe `DENSE_RANK()` window function assigns a rank to each row within its partition, with identical values receiving the same rank, and subsequent ranks being consecutive (no gaps).
- Requires an `OVER` clause with `PARTITION BY` and `ORDER BY`.
- Assigns ranks starting from 1.
- Does not produce gaps in the ranking sequence when ties occur.
- Useful for identifying top N items where ties should share the same rank and the next rank should be immediately after.
Memory trick: Dense rank is for when you want shared gold, no skipping numbers!
SQL WHERE Clause with AND/OR
Flip cardThe SQL WHERE clause is used to filter records based on specified conditions, allowing multiple conditions to be combined using logical operators like AND, OR, and NOT.
- Filters individual rows before any grouping or aggregation.
- `AND` requires all combined conditions to be true.
- `OR` requires at least one combined condition to be true.
- Commonly used with comparison operators, date functions, and mathematical operators.
Memory trick: Where All Conditions Meet, And Or Not, Filters the Rows Neatly.
Line Chart Use Case
Flip cardA line chart is a type of chart that displays information as a series of data points called 'markers' connected by straight line segments. It is primarily used to visualize trends over time.
- Excellent for showing continuous data.
- Clearly illustrates trends, acceleration, deceleration, and volatility.
- Time series data is a primary use case.
Memory trick: Lines for time, Bars for categories, Pies for parts.
Column-Level Security (CLS)
Flip cardColumn-Level Security (CLS) in data platforms like Microsoft Fabric restricts access to specific columns in a table based on the user's role or permissions, ensuring sensitive data is only visible to authorized individuals.
- Protects sensitive data at a granular level.
- Integrates with SQL endpoint and other query engines.
- Complements RLS for comprehensive data protection.
Memory trick: CLS guards your columns, RLS guards your rows.
Data Strategy and Requirements Definition
Flip cardThe initial phase in data solution planning where business objectives are translated into data needs, including identifying data sources, defining data quality standards, and establishing data retention policies.
- Crucial for successful data projects.
- Involves business and technical stakeholders.
- Sets the foundation for data ingestion and modeling.
Memory trick: First define your data, then build your Fabric.
Spark SQL Query Logical Order
Flip cardThe conceptual sequence in which Spark SQL processes clauses in a query, which dictates where filtering, grouping, and ordering operations occur.
- FROM specifies the data source.
- WHERE filters individual rows before grouping.
- GROUP BY aggregates rows into groups.
- HAVING filters groups after aggregation.
Memory trick: From Where Groups Have Selected Orders Limited.
Dual Storage Mode
Flip cardDual Storage Mode in Microsoft Fabric semantic models allows a table to operate in both Import and DirectQuery modes simultaneously, providing the benefits of both.
- Cached for fast queries (Import).
- Queries live data for freshness (DirectQuery).
- Optimizes performance for different query types.
Memory trick: Import is Fast, Direct is Fresh, Dual is Both.
Date Dimension Table
Flip cardA dedicated table in a semantic model that contains all date-related attributes (year, quarter, month, day, day of week, etc.) and is used to filter and group data based on time.
- Provides a single source for date-related analysis.
- Essential for complex time intelligence functions (e.g., YTD, MoM).
- Relates to fact tables on their date columns.
Memory trick: A single date table rules all time analyses.
One-to-Many Relationship
Flip cardA one-to-many relationship in a semantic model connects a table where each value in a key column is unique (the 'one' side) to another table where those key values can appear multiple times (the 'many' side).
- Most common relationship type in star schemas.
- Filters propagate from the 'one' side to the 'many' side by default.
- Ensures accurate aggregation and filtering from dimension to fact tables.
- Key column on 'one' side must be unique.
Memory trick: Cardinality: How Many on Each Side?
Key Columns in Semantic Models
Flip cardA column designated as a 'key' in a semantic model (e.g., in Power BI Desktop or Tabular Editor) indicates to the modeling engine that its values are unique and non-nullable. This is critical for defining relationships and optimizing query performance.
- Enforces uniqueness and non-nullability for the specified column.
- Essential for defining valid relationships between tables.
- Helps the modeling engine optimize data storage and query execution.
Memory trick: Key column set, uniqueness met, data integrity, you bet!
DirectQuery Mode
Flip cardDirectQuery mode in Microsoft Fabric semantic models allows direct querying of the underlying data source without importing data into the model. This is ideal for very large datasets and scenarios requiring near real-time data.
- Data is not cached in the semantic model.
- Queries are translated into native queries for the source system.
- Good for large datasets and near real-time requirements.
- Performance depends heavily on the underlying data source.
Memory trick: Fabric's storage modes are a direct path to data or an imported treasure chest.
Calculation Groups
Flip cardA modeling feature in Microsoft Fabric and Power BI that reduces redundant measures by allowing calculation items (e.g., time intelligence, currency conversion) to be applied to existing measures.
- Simplifies model management by reducing the number of explicit measures.
- Enhances reusability of calculation logic.
- Applied to measures at query time.
Memory trick: Group your calcs, simplify your model, make it sparkle.
Automatic Aggregations
Flip cardAutomatic aggregations in Microsoft Fabric semantic models optimize query performance by automatically creating and managing in-memory aggregate tables. Queries that can be answered by these aggregates are resolved faster, while drill-down queries seamlessly use the underlying detailed data.
- Improves query performance for aggregated data.
- Automatically generated and optimized by Fabric/Power BI.
- Seamlessly combines aggregated and detailed data access.
- Reduces load on the underlying data source.
Memory trick: Optimize queries by aggregating automatically, then drill down when needed, magically.
Incremental Refresh
Flip cardA refresh policy for large Import mode semantic models that optimizes refresh times by processing only new or changed data partitions, rather than a full reload of the entire table.
- Significantly reduces refresh duration for large datasets.
- Requires defining RangeStart and RangeEnd parameters in Power Query.
- Allows different refresh frequencies for historical vs. recent data.
Memory trick: Incremental refresh, only recent data gets fresh.
Direct Lake Mode
Flip cardDirect Lake mode in Microsoft Fabric semantic models allows direct querying of Delta Lake tables in OneLake without importing or duplicating data. It combines the performance benefits of Import mode with the real-time capabilities of DirectQuery for Delta Lake sources.
- Queries Delta Lake tables directly from OneLake.
- No data movement or duplication.
- Offers Import-like performance with DirectQuery-like freshness.
- Optimized for large-scale analytics on Delta Lake.
Memory trick: For Delta Lake, go Direct to the Lake for speed and freshness.
Mark as Date Table
Flip cardMarking a table as a 'Date Table' in a Microsoft Fabric semantic model is a critical step to enable and correctly utilize DAX time intelligence functions (e.g., TOTALYTD, SAMEPERIODLASTYEAR). It tells the DAX engine which table contains the definitive date column for time-based calculations.
- Essential for DAX time intelligence functions to work.
- Requires a single, continuous date column with unique values.
- Configured in the model view by right-clicking the table.
- Ensures correct context for time-based calculations.
Memory trick: To make time intelligence smart, mark the date table's heart.
Continuous Date Table
Flip cardA prerequisite for DAX time-intelligence functions, requiring a date dimension table that contains every single date within its defined range, with no gaps.
- Essential for accurate time-intelligence calculations.
- Must have a unique date column with a date or datetime data type.
- Typically generated using M query or DAX (CALENDARAUTO/CALENDAR).
Memory trick: Date table must be complete, marked, and linked.
Manual Aggregations
Flip cardManual aggregations in Microsoft Fabric semantic models involve creating pre-summarized tables (often in Import mode) that are used by the query engine to accelerate aggregated queries on large DirectQuery or Dual mode fact tables, while still allowing access to granular data.
- Optimizes performance for large fact tables.
- Combines benefits of Import (for aggregates) and DirectQuery (for detail).
- Requires explicit definition of summary tables and mapping rules.
- Query engine automatically rewrites queries to use aggregations.
Memory trick: Aggregates for Speed, Direct for Detail!
CROSSFILTER DAX Function
Flip cardThe CROSSFILTER DAX function modifies the filter direction of a specified relationship within a calculation, enabling correct filtering across many-to-many relationships or overriding default filter behavior.
- Used in DAX expressions to control relationship filter behavior.
- Crucial for many-to-many relationships, especially with bridge tables.
- Can change filter direction to Both, OneWay, or None.
Memory trick: Cross the Filter, Don't Get Lost in the Many-to-Many!
Disconnected Slicer Table with TREATAS
Flip cardA DAX pattern used in semantic models where a non-related table (disconnected slicer) is used for filtering, and DAX functions like `TREATAS` or `FILTER` with `ALL` are used in measures to dynamically apply the selection from the slicer to the model's data.
- Enables dynamic segmentation and filtering based on custom logic.
- The slicer table is not joined to other tables in the model.
- Provides flexibility for scenarios like 'New vs. Repeat Customers' without altering the base model structure.
Memory trick: Disconnected slicer, TREATAS connects, new and repeat, a DAX perfect reflex.
XMLA Endpoint
Flip cardAn industry-standard protocol used for communication with Analysis Services engines, enabling administrative tasks, data refresh, and metadata management for semantic models.
- Provides read/write access to semantic models.
- Allows deployment of .bim files (Tabular models).
- Used by tools like SSMS, Tabular Editor, and Azure Data Studio.
Memory trick: XMLA for models, APIs for automation, PowerShell for scripts, UI for basics.
Model Relationships
Flip cardConnections between tables in a semantic model that define how data from different tables are related, enabling proper filtering and aggregation across the model.
- Crucial for correct data aggregation and filtering.
- Defined by cardinality (one-to-many, many-to-many, etc.) and cross-filter direction.
- Can be active or inactive.
Memory trick: Relationships link, Measures calculate, Tables hold data, Columns describe.
Event-Driven Refresh
Flip cardA semantic model refresh strategy where the refresh process is initiated automatically in response to specific events, such as source data changes, rather than on a fixed schedule.
- Achieved using Power BI REST APIs.
- Often integrated with Azure Functions, Logic Apps, or webhooks.
- Ensures data freshness upon source updates, reduces unnecessary refreshes.
Memory trick: Scheduled for routine, API for events, Incremental for big data.
Cross-filter Direction (Both)
Flip cardThe 'Both' cross-filter direction in a semantic model relationship allows filters to propagate from both sides of the relationship, enabling bi-directional filtering.
- Filters flow from 'one' to 'many' and 'many' to 'one'.
- Useful for enabling filtering from fact tables to dimension tables.
- Can impact performance and ambiguity; use judiciously.
Memory trick: One-Way is Simple, Both-Ways is Flexible.
Cross-filter Direction
Flip cardCross-filter direction in a semantic model relationship dictates how filters propagate between related tables, either unidirectionally (Single) or bidirectionally (Both), impacting report interactions and calculation results.
- Controls filter flow between tables.
- Single (One-Way): filter propagates from 'one' side to 'many' side.
- Both (Two-Way): filter propagates in both directions, use cautiously.
Memory trick: Filter Flow: Choose Your Path, Single or Both!
DAX Time Intelligence
Flip cardDAX Time Intelligence functions facilitate calculations across various time periods (e.g., YTD, QTD, MTD, moving averages) within a semantic model, requiring a marked date table for proper functionality.
- Requires a dedicated and marked Date table.
- Simplifies complex date-based comparisons and aggregations.
- Functions like TOTALYTD, DATESBETWEEN, SAMEPERIODLASTYEAR.
Memory trick: DAX Functions: Do You Know Your Time, Text, and Logic?
Real-time Refresh
Flip cardA refresh strategy that provides the freshest possible data by leveraging DirectQuery or Direct Lake modes, where queries are executed directly against the source data, minimizing latency.
- Achieved primarily through DirectQuery or Direct Lake storage modes.
- Data is always current, reflecting source changes immediately.
- Suitable for scenarios requiring minimal data latency, but can impact query performance.
Memory trick: Real-time data: no wait, just query straight.
Bi-directional Filtering (Many-to-Many through Fact)
Flip cardEnabling bi-directional filtering on one-to-many relationships connected by a fact table implicitly creates a many-to-many relationship between the two dimension tables, allowing filters to flow between them.
- Filters flow from both sides through the fact table.
- Allows dimensions to filter each other via a common fact.
- Can impact performance; use with understanding of data model.
Memory trick: Dimensions Meet in the Middle, Fact is the Bridge.
Relationship Cross-filter Direction
Flip cardThe cross-filter direction property in a semantic model relationship defines the direction in which filters propagate between related tables, being either unidirectional ('Single') or bidirectional ('Both').
- Crucial for controlling how filters affect related tables.
- Default for one-to-many is 'Single' (from 'one' to 'many').
- Setting to 'Both' allows filters to flow in both directions.
- Can impact query results if not set correctly for desired analysis.
Memory trick: Relationships: Cardinality, Direction, and Activity!
Mixed Mode Storage
Flip cardA semantic model storage mode in Microsoft Fabric that combines Import and DirectQuery modes, allowing different tables to use different storage types.
- Offers a balance between performance and real-time data access.
- Tables can be configured individually as Import or DirectQuery.
- Useful for scenarios requiring both high performance and fresh data.
Memory trick: Mix and match for the best data dance.
Combined RLS and OLS
Flip cardCombining Row-Level Security (RLS) and Object-Level Security (OLS) in a Microsoft Fabric semantic model allows for comprehensive data access control. RLS filters rows based on user identity, while OLS hides entire tables, columns, or measures, providing both horizontal and vertical security.
- RLS: Filters rows (horizontal security).
- OLS: Hides tables, columns, or measures (vertical security).
- Used together for granular and complete data access control.
- OLS is configured using Tabular Editor, RLS via the Fabric UI or Tabular Editor.
Memory trick: For full security, filter rows with RLS, hide objects with OLS.
Cross-filter Direction (One-to-Many)
Flip cardThe direction in which filters propagate through a relationship in a semantic model, typically from the 'one' side of a one-to-many relationship to the 'many' side.
- Determines how filters from one table affect another.
- For one-to-many, usually from 'one' to 'many' for correct aggregation.
- Incorrect direction can lead to wrong calculations or errors.
Memory trick: Filter flows from the source of truth to the detailed records.
Mixed Storage Mode
Flip cardA semantic model storage mode in Microsoft Fabric that combines Import and DirectQuery modes, allowing different tables or partitions to use different storage types.
- Offers flexibility for performance and data freshness.
- Allows some tables to be imported for speed, others to be DirectQuery for real-time.
- Balances performance, data freshness, and resource consumption.
Memory trick: Mix and match for the perfect data recipe.
Measure Performance Optimization (Aggregations)
Flip cardFor semantic models with large fact tables in Import mode, optimizing measures that perform common aggregations (especially time intelligence) is often achieved by implementing aggregations (manual or automatic) to pre-calculate results.
- Aggregations significantly reduce the number of rows processed at query time.
- Automatic Aggregations are managed by Fabric, simplifying implementation.
- Especially effective for measures involving SUM, COUNT, AVG over large datasets.
Memory trick: Millions of rows, time intelligence slow, aggregates make the data flow.
Many-to-Many Relationship
Flip cardA type of relationship in a data model where a record in Table A can relate to multiple records in Table B, and a record in Table B can also relate to multiple records in Table A. Often implemented using a bridging (or junction) table.
- Used when direct one-to-many relationships are not sufficient.
- Requires an intermediate bridging table to resolve in relational models.
- Supported directly in Fabric semantic models, but bridging tables are still best practice.
Memory trick: Many-to-many, a bridge you'll need, for complex data, it's the right deed.
Manual Aggregation Tables
Flip cardManual aggregation tables are pre-calculated summary tables created within a Microsoft Fabric semantic model to significantly improve query performance for aggregated data by redirecting queries from detailed tables.
- Engineers define aggregations explicitly.
- Queries are automatically rewritten to use aggregations.
- Reduces query time for large datasets.
Memory trick: Accelerate Queries: Aggregate, Cache, or Stream!
Semantic Model Aggregations
Flip cardPre-calculated and stored summarized tables within a semantic model that are used to speed up queries by returning results from aggregated data when possible.
- Improves query performance for large datasets.
- Can be configured as Import, DirectQuery, or Dual storage.
- Requires careful design to cover common query patterns.
Memory trick: Aggregations speed up summaries, columns optimize storage, partitions refresh faster.
DAX for RLS (Context)
Flip cardData Analysis Expressions (DAX) used within security roles in a semantic model to dynamically filter data based on the authenticated user's identity or attributes.
- Evaluated in the context of the current user.
- Commonly uses functions like USERNAME(), USERPRINCIPALNAME(), and LOOKUPVALUE.
- Filters data at the row level, ensuring users only see authorized data.
Memory trick: Lookup the user's identity to filter their view.
OLS and RLS Combination
Flip cardUsing both Object-Level Security (OLS) to hide entire objects (tables/columns) and Row-Level Security (RLS) to filter rows within visible objects, providing comprehensive data access control.
- OLS controls visibility of objects (tables, columns, measures).
- RLS controls visibility of rows within tables.
- Often used together for granular security requirements.
Memory trick: Objects hidden, rows filtered, all secure and covered.
Many-to-Many Relationships
Flip cardA type of relationship between two tables where records in one table can relate to multiple records in the other table, and vice-versa. It is implemented using an intermediary 'bridge' table.
- Requires a bridge table to resolve in semantic models.
- Bridge table connects two dimension tables with one-to-many relationships.
- Enables correct filtering and aggregation across both tables.
Memory trick: When many meet many, build a bridge between them.
Direct Lake with Import (Mixed Mode)
Flip cardA semantic model configuration that combines Direct Lake mode for tables sourced from OneLake (for near real-time, high performance) and Import mode for other tables (like small dimensions from SQL databases) to achieve optimal performance and data freshness across diverse sources.
- Leverages Direct Lake benefits for OneLake data.
- Uses Import mode for other data for maximum performance.
- Enables a 'mixed' approach within a single semantic model.
- Ideal for composite models with OneLake and traditional DB sources.
Memory trick: Direct Lake for the lake, Import for the static stake.
DAX Formula Optimization
Flip cardThe process of refactoring or rewriting Data Analysis Expressions (DAX) formulas to improve their execution speed and reduce resource consumption in a semantic model.
- Crucial for complex measures and large datasets.
- Involves simplifying logic, choosing efficient functions, and avoiding performance anti-patterns.
- Can significantly reduce query times and improve user experience.
Memory trick: Tune your DAX like a finely-tuned engine for speed.
Source-Side Filtering for Security
Flip cardSource-side filtering for security involves applying filters directly in the data source query or data transformation (e.g., Power Query) before data is loaded into the semantic model. This minimizes data transfer, enhances performance, and enforces security at the earliest possible stage.
- Filters data at the source, reducing data loaded into the model.
- Improves performance by minimizing data transfer.
- Enforces security 'at the data source level'.
- Can be implemented via source query parameters or Power Query filters.
Memory trick: Filter at the source, secure the flow, before the data's in the show.
CALCULATE Function
Flip cardThe most important function in DAX, which evaluates an expression in a context modified by new filters or removal of existing filters.
- Changes filter context to evaluate an expression.
- Essential for time intelligence, comparisons, and dynamic filtering.
- Can accept multiple filter arguments (tables, columns, boolean expressions).
Memory trick: Calculate to control the context of your data.
Import Mode
Flip cardA data storage mode in Power BI/Fabric where data is loaded and cached directly into the semantic model, providing high performance and supporting advanced features.
- Data is cached in memory.
- Offers fastest query performance.
- Supports full Power BI/Fabric capabilities, including incremental refresh.
Memory trick: Import is in, Direct is out, Dual is both, Live is linked.
Dynamic RLS
Flip cardA Row-Level Security implementation where filtering criteria are determined at query time based on the querying user's identity or attributes.
- Uses DAX expressions to filter rows.
- Often relies on USERNAME() or USERPRINCIPALNAME() functions.
- Scalable for many users with varying access patterns.
Memory trick: Row-level filters data, Object-level hides objects, Workspace controls access, Gateway secures connections.
Spark for Large-Scale ETL
Flip cardSpark Notebooks in Microsoft Fabric, utilizing PySpark or Scala, are the preferred tool for performing complex, large-scale ETL (Extract, Transform, Load) operations on multi-terabyte datasets, offering distributed processing, custom logic, and high performance.
- Handles multi-terabyte datasets efficiently.
- Supports custom code (PySpark, Scala).
- Leverages distributed computing for performance.
- Ideal for complex cleansing, enrichment, and aggregation.
Memory trick: Spark handles complex ETL at scale.
Spark DataFrame API Optimization
Flip cardThe practice of structuring PySpark code to leverage Spark's Catalyst Optimizer and native functions for maximum performance, especially by avoiding Python UDFs.
- Python UDFs can be performance bottlenecks.
- Prioritize built-in PySpark SQL functions and DataFrame API operations.
- Allows Spark to perform whole-stage code generation and other optimizations.
- Minimizes serialization/deserialization overhead.
Memory trick: Use Spark's native tools, don't force a Python peg into a Spark hole.
Copy Data Wildcard Path
Flip cardThe 'Wildcard file path' setting in a Data Pipeline's Copy Data activity allows dynamic selection of source files based on patterns, often incorporating dynamic expressions.
- Enables filtering files by name, date, or other patterns.
- Supports system variables like 'utcNow()' for dynamic dates.
- Efficient for ingesting a subset of files from a folder.
Memory trick: Pipeline flows, date-filtered, like a river finding its daily path.
Dataflows Gen2 for API Ingestion
Flip cardDataflows Gen2 leverage Power Query to connect to and ingest data from various sources, including web APIs and semi-structured formats like JSON, enabling visual transformations.
- Uses Power Query Editor for visual ETL.
- Supports a wide range of connectors, including Web API.
- Ideal for ingesting semi-structured data and applying basic to intermediate transformations.
- Output can be loaded directly into a Lakehouse or Data Warehouse.
Memory trick: Web API data flows visually to the Lakehouse.
Fabric Eventstream
Flip cardMicrosoft Fabric Eventstream is a real-time analytics component for ingesting, transforming, and routing high-volume streaming data with low latency.
- Designed for real-time data ingestion (e.g., IoT, logs).
- Supports various streaming sources and destinations.
- Enables real-time analytics and anomaly detection.
Memory trick: Data flows into Fabric, either a calm lake or a rushing stream.
Power Query Combine Files
Flip cardA feature in Power Query that simplifies combining multiple files with the same schema from a folder into a single dataset.
- Automates the creation of a transformation function.
- Applies transformations to all files consistently.
- Efficient for folder-based data sources.
Memory trick: Folder of files becomes one table with a single click.
Data Pipeline Parameterized ADLS Gen2 Ingestion
Flip cardThe ability to dynamically specify source file paths in Azure Data Lake Storage Gen2 within a Microsoft Fabric Data Pipeline using parameters, allowing for flexible and reusable ingestion.
- Uses pipeline parameters to pass dynamic values (e.g., dates, folders).
- Dataset path is defined with parameters.
- Enables ingestion from varying paths without pipeline modification.
- Essential for incremental or date-based loads.
Memory trick: Parameters are like GPS coordinates, guiding your pipeline to the right data each time.
Dataflows Gen2
Flip cardA low-code data integration tool in Microsoft Fabric for ingesting, transforming, and preparing data from various sources.
- Uses Power Query for transformations.
- Supports a wide range of data sources.
- Can output directly to Lakehouse tables.
Memory trick: Flowing data, visually transformed, ready for the Lakehouse.
Copy Data Fault Tolerance
Flip cardSettings within the Copy Data activity that define how errors are handled during data transfer, ensuring resilience and data integrity.
- Can skip incompatible rows.
- Can log inconsistent data.
- Prevents pipeline failure due to minor data issues.
Memory trick: Fault tolerance is the pipeline's safety net.