Microsoft Certified: Fabric Analytics Engineer Associate flashcards
162 free flashcards. Tap a card to flip it.
Fabric Capacity Region
Flip cardThe geographical region where a Microsoft Fabric capacity is provisioned, which dictates the physical location of all data processing and storage for workspaces and items assigned to that capacity.
- Primary control for data residency in Fabric.
- All associated workspaces and items inherit this region.
- Crucial for compliance with local regulations.
Memory trick: The Fabric Capacity's Region is the 'Data's Homeland' – where it lives.
Fabric Capacity Metrics App
Flip cardThe Fabric Capacity Metrics app is a pre-built Power BI solution that provides administrators with detailed insights into the performance, resource utilization, and workload distribution of their Microsoft Fabric capacities.
- Monitors CPU, memory, and other resource usage.
- Shows consumption by different Fabric workloads (Spark, SQL, etc.).
- Helps identify bottlenecks and optimize capacity settings.
- Available from the Fabric Admin portal.
Memory trick: To check your Fabric engine's health, open the 'Capacity Metrics' dashboard.
Fabric Customer-Managed Keys (CMK)
Flip cardCustomer-Managed Keys (CMK) in Microsoft Fabric allow organizations to use their own encryption keys from Azure Key Vault to encrypt data at rest within Fabric capacities, enhancing security and compliance.
- Configured at the Fabric capacity level.
- Uses Azure Key Vault for key management.
- Encrypts all data at rest within the associated capacity.
Memory trick: CMK for Fabric is like giving your 'Master Key' from Azure Key Vault to unlock your Fabric 'Data Vault'.
Fabric Tenant Setting: Workspace Creation
Flip cardThis Microsoft Fabric tenant setting controls which users or security groups are permitted to create new workspaces, allowing for centralized governance of workspace provisioning.
- Governs who can create new Fabric workspaces.
- Found in the Fabric Admin Portal under Tenant Settings.
- Can be set for the entire organization or specific security groups.
Memory trick: Think of the 'Create Workspaces' setting as the 'Master Key' to the Fabric building.
Fabric Git Integration
Flip cardMicrosoft Fabric's Git integration allows developers to connect workspaces to Azure DevOps Git repositories, enabling version control, collaborative development, and change tracking for Fabric items.
- Supports collaborative development.
- Enables versioning and change tracking.
- Facilitates rollback to previous stable states.
- Crucial for CI/CD pipelines.
Memory trick: Git is your time machine for Fabric items, letting you go back to stable versions.
Azure Policy for Fabric
Flip cardAzure Policy allows organizations to create, assign, and manage policies that enforce rules and effects over their Azure resources, including Microsoft Fabric items, ensuring compliance with organizational standards.
- Enforces organizational standards and governance.
- Can validate naming conventions, resource types, locations.
- Applies at resource creation/update time to prevent non-compliance.
Memory trick: Azure Policy is like the 'Name Tag Enforcer' for your Fabric items.
Fabric Git & Pull Requests
Flip cardIntegrating Microsoft Fabric with Git (e.g., Azure DevOps) and utilizing Pull Requests enables robust version control, collaborative code development, code reviews, and enforced merging policies for Fabric artifacts.
- Connects Fabric workspaces to Git repositories.
- Supports branching, committing, and merging.
- Pull Requests facilitate code reviews and approval workflows.
- Essential for CI/CD and quality assurance.
Memory trick: Git and Pull Requests are like your code's quality control gateway and collaborative whiteboard.
Fabric Deployment Pipelines
Flip cardDeployment pipelines in Microsoft Fabric streamline the development lifecycle by enabling content creators to develop, test, and deploy content across different environments (stages) with built-in governance.
- Supports multiple stages (e.g., Dev, Test, Prod).
- Enables consistent content deployment.
- Provides version control and deployment rules.
- Facilitates approval workflows.
Memory trick: Pipelines take your Fabric items on a journey from dev to production, stage by stage.
Lakehouse RLS and OLS
Flip cardRole-level security (RLS) and object-level security (OLS) in a Fabric Lakehouse allow for granular control over who can see which rows and columns of data, respectively, within tables.
- RLS filters rows based on user identity.
- OLS restricts access to specific columns or tables.
- Applied at the data model or query level.
- Overrides broader workspace permissions for data access.
Memory trick: To secure your Lakehouse data, use RLS for rows and OLS for objects.
Microsoft 365 Audit Log (for Fabric)
Flip cardThe Microsoft 365 Audit log is the centralized repository for auditing user and administrative activities across Microsoft 365 services, including all actions performed within Microsoft Fabric workspaces and items.
- Records administrative and user activities.
- Crucial for security audits and compliance.
- Can be searched via the Microsoft Purview compliance portal.
Memory trick: Think of the Microsoft 365 Audit Log as Fabric's 'Activity Diary' – everything is written down.
Fabric Capacity per Team
Flip cardAssigning dedicated Fabric Capacity SKUs to individual teams or projects allows for resource isolation, performance guarantees, and granular cost allocation (chargeback) within an organization.
- Ensures performance isolation between teams.
- Enables accurate cost tracking for each team's compute.
- Provides dedicated compute resources for specific workloads.
- Requires purchasing separate Fabric Capacity SKUs.
Memory trick: Give each team their own 'house' (workspace) and their own 'power meter' (capacity SKU).
Object-Level Security (OLS)
Flip cardObject-Level Security (OLS) is a data security feature that controls access to database objects like tables, columns, or views, restricting who can even see or reference these objects.
- Restricts access to entire columns or tables.
- Users without permission cannot see the protected object.
- Complements RLS for comprehensive data security.
Memory trick: OLS is like having a 'Secret Drawer' for columns; if you don't have the key, you can't even see it's there.
Data Loss Prevention (DLP)
Flip cardDLP policies are security measures that prevent sensitive data from leaving an organization's control or being shared with unauthorized entities, often by monitoring, identifying, and blocking data transfers.
- Prevents unauthorized sharing of sensitive data.
- Works by identifying and monitoring sensitive information.
- Can be configured to block, warn, or audit sharing activities.
Memory trick: DLP is your Digital Leakage Protector, stopping sensitive data from escaping.
Information Protection Policies in Fabric
Flip cardThese policies allow organizations to classify and label sensitive data within Microsoft Fabric items, ensuring compliance and data governance.
- Integrates with Microsoft Purview.
- Enforces sensitivity labels and data classification.
- Helps meet compliance requirements.
Memory trick: Protect your Fabric data with policies, just like a secure vault.
Fabric Inherent Encryption
Flip cardMicrosoft Fabric provides inherent, baseline encryption for all data at rest and in transit, utilizing Microsoft-managed keys for storage and industry-standard protocols like TLS for network communication.
- Data at rest is encrypted by default (Microsoft-managed keys).
- Data in transit is encrypted by default (TLS 1.2+).
- No user configuration required for baseline encryption.
- CMK is an optional, advanced feature.
Memory trick: Fabric encrypts your data automatically, like a secure courier and a locked vault.
Fabric Item Lifecycle Policies
Flip cardItem lifecycle policies in Microsoft Fabric automate the management of workspace items, enabling actions like deletion based on conditions such as age or inactivity.
- Automates item management in workspaces.
- Can be configured based on last modified date or creation date.
- Helps with workspace hygiene and cost management.
Memory trick: Item Lifecycle Policies are like a digital janitor, tidying up old Fabric items automatically.
Fabric Capacity Workload Settings
Flip cardFabric Capacity Workload Settings allow administrators to control the maximum percentage of a Fabric capacity's resources that can be utilized by specific workload types (e.g., Data Engineering, Data Warehousing).
- Allocates compute resources within a capacity.
- Prevents resource contention between different workloads.
- Configurable in the Fabric Admin Portal for each capacity.
Memory trick: Capacity Workload Settings are like 'Traffic Cops' for your Fabric compute, directing and limiting flow.
Fabric Audit Logs
Flip cardAudit logs in Microsoft Fabric record user and system activities, crucial for security, compliance, and troubleshooting. They are typically accessed via the Microsoft Purview compliance portal.
- Records actions like item creation, modification, access.
- Essential for regulatory compliance (e.g., GDPR, HIPAA).
- Retainable for extended periods as per policy.
Memory trick: For compliance, think 'Purview Audit Logs' – it's your digital paper trail.
Fabric Capacity SKUs
Flip cardFabric Capacity SKUs (Stock Keeping Units) define the dedicated computational resources purchased for Microsoft Fabric, enabling resource isolation, performance guarantees, and granular administration.
- Represent dedicated compute for Fabric workloads.
- Available in various sizes (e.g., F2, F4, F8, F64).
- Can be assigned to workspaces.
- Allow for region selection and workload settings.
Memory trick: Think of Fabric Capacity SKUs as your dedicated 'Fabric Factories' – each with its own machines and location.
Fabric Tenant Settings
Flip cardTenant settings in the Microsoft Fabric Admin portal are global configurations that apply across all workspaces and users within an organization's Fabric tenant.
- Control organization-wide policies and defaults.
- Managed by Fabric administrators.
- Affect aspects like security, data governance, and feature availability.
Memory trick: Tenant settings are like the company handbook, applying to everyone.
Row-Level Security (RLS)
Flip cardRow-Level Security (RLS) is a data security feature that restricts access to individual rows in a database table based on the execution context of the user, typically their identity or role.
- Filters data at the row level.
- Ensures users only see authorized data subsets.
- Implemented using security predicates or filters.
Memory trick: RLS is like a bouncer at a club, letting only certain 'rows' of people in.
Power BI Matrix Visual
Flip cardThe Power BI Matrix visual is a grid-like chart that supports displaying data across multiple dimensions with hierarchical row and column groupings, enabling drill-down capabilities and aggregated totals.
- Ideal for hierarchical and cross-tabulated data.
- Supports drill-down/up functionality.
- Allows for multiple row and column fields.
- Displays subtotals and grand totals.
Memory trick: Matrix for multi-level data, Card for one number, Table for raw, Gauge for goals.
Spark SQL FIRST_VALUE()
Flip cardA Spark SQL window function that returns the value of the specified expression from the first row within its window frame, as defined by the PARTITION BY and ORDER BY clauses.
- Retrieves the value from the first row of a window.
- Requires `PARTITION BY` and `ORDER BY`.
- Useful for finding initial states, first events, or starting values.
- More direct than `ROW_NUMBER()=1` for just getting the value.
Memory trick: First Value, First Sight, in Each Partition's Ordered Line.
Delta Lake ACID Transactions
Flip cardDelta Lake provides ACID (Atomicity, Consistency, Isolation, Durability) properties to data operations, ensuring reliability for data writes and reads, especially critical for streaming and concurrent workloads.
- Atomicity: All or nothing changes.
- Consistency: Data always in a valid state.
- Isolation: Concurrent operations don't interfere.
- Durability: Committed changes are permanent.
Memory trick: ACID ensures your data is always valid and sound!
Delta Lake MERGE INTO
Flip cardThe Delta Lake MERGE INTO command allows for conditional insertion, updating, and deletion of records in a Delta table based on a source table or DataFrame, providing robust upsert capabilities.
- Performs upsert (update or insert) operations.
- Can also include delete functionality.
- Ensures atomicity for complex data synchronization tasks.
Memory trick: Merge your data, new or old, a consistent story to be told!
OneLake Shortcuts
Flip cardOneLake shortcuts enable virtualized data access by creating references to data files or folders located elsewhere within OneLake or external cloud storage, promoting data reuse and avoiding duplication.
- Acts as a pointer, not a copy of data.
- Supports cross-Lakehouse and external data referencing.
- Facilitates logical data organization and sharing.
Memory trick: Shortcuts link, don't duplicate, for a single data state!
Fabric Data Pipelines
Flip cardData Pipelines in Microsoft Fabric provide a cloud-native, serverless orchestration service for creating, scheduling, and managing data integration workflows (ETL/ELT).
- Supports various data sources and destinations.
- Enables data movement and transformation activities.
- Offers scheduling, monitoring, and error handling.
Memory trick: Pipelines orchestrate data flow, from source to Lakehouse.
Delta Lake OPTIMIZE
Flip cardThe Delta Lake OPTIMIZE command improves query performance on Delta tables by compacting small files into larger, more efficient files.
- Reduces metadata overhead for query engines.
- Can be run with optional ZORDER by columns for further performance gains.
- Does not alter the data content, only its physical storage.
Memory trick: Small files slow? Optimize and go!
Spark SQL BROADCAST Hint
Flip cardThe `BROADCAST` hint in Spark SQL forces the specified table (usually the smaller one) to be broadcast to all worker nodes during a join operation, thereby optimizing performance by avoiding a shuffle of the larger table.
- Used for joining a small table with a large table.
- Table size threshold for broadcasting is configurable (default is 10MB).
- Significantly reduces network I/O and improves join speed.
- Can lead to OutOfMemory (OOM) errors if the broadcasted table is too large.
Memory trick: Broadcast small tables, shuffle large ones, merge if you have to!
Spark SQL Running Total
Flip cardA running total, also known as a cumulative sum, is an aggregate calculation that adds each new value to the sum of previous values within a defined group and order.
- Uses a window function with `SUM()`.
- Requires `PARTITION BY` to group data (e.g., per product).
- Requires `ORDER BY` to define the accumulation sequence (e.g., by date).
- Uses `ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW` to define the accumulating window frame.
Memory trick: Partition to group, order to sequence, frame to accumulate!
Spark SQL Top N Query
Flip cardTo find the top N records based on an aggregate value in Spark SQL, you typically group the data, calculate the aggregate, order the results in descending order of the aggregate, and then use the `LIMIT` clause to restrict the number of rows.
- Sequence: GROUP BY -> Aggregate -> ORDER BY DESC -> LIMIT N.
- `LIMIT` is standard in Spark SQL for top N.
- Crucial for ranking and identifying leading entities.
Memory trick: Group, Sum, Order, Limit to find the best.
PySpark Filtering Aggregated Data
Flip cardIn PySpark, filtering based on an aggregated value (equivalent to SQL's `HAVING` clause) is achieved by first performing the `groupBy()` and `agg()` operations, and then applying a `filter()` transformation on the resulting DataFrame using the aggregated column.
- Aggregation (`groupBy().agg()`) must precede the filter.
- The `filter()` method is used for both pre- and post-aggregation filtering.
- Refer to the aggregated column by its alias within the `filter()` clause.
- There is no direct `.having()` method in PySpark DataFrames.
Memory trick: Group, then aggregate, then filter the aggregates!
Medallion Architecture: Silver Layer
Flip cardThe Silver layer in the Medallion Architecture stores cleansed, conformed, and semi-processed data, where data quality rules, schema enforcement, and basic transformations are applied.
- Data is structured and standardized.
- Focus on data quality and consistency.
- Serves as the foundation for the Gold layer.
Memory trick: Bronze is raw, Silver is clean, Gold is keen!
SQL WHERE Clause
Flip cardThe SQL WHERE clause is used to extract only those records that fulfill a specified condition.
- Filters rows before grouping.
- Can combine multiple conditions using AND, OR, NOT.
- Essential for precise data retrieval.
Memory trick: Select From Where Group Having Order — So F**king What, Get Hired On!
Spark SQL LAG() Function
Flip cardThe `LAG()` window function returns the value of an expression from a row that precedes the current row by a specified offset within its partition.
- Requires an `OVER` clause with `PARTITION BY` and `ORDER BY`.
- Syntax: `LAG(expression, offset, default_value)`.
- Offset specifies how many rows back to look (default is 1).
- Default value is returned if the offset goes beyond the partition boundary (default is NULL).
Memory trick: Lag looks back, Lead looks forward, First/Last grab the ends!
Spark SQL DAYOFWEEK()
Flip cardThe `DAYOFWEEK()` function in Spark SQL extracts the day of the week from a date or timestamp expression, returning an integer value.
- Returns an integer from 1 to 7.
- 1 represents Sunday.
- 7 represents Saturday.
- Commonly used for filtering or grouping by day of the week.
Memory trick: Day of week, 1 is Sunday, 7 is Saturday, remember the start and end!
Spark SQL Date Part Extraction
Flip cardThe `EXTRACT(part FROM date)` function in Spark SQL is used to retrieve specific components (e.g., year, month, day) from a date or timestamp expression.
- Standard SQL syntax for date part extraction.
- Supported by Spark SQL queries in Fabric Lakehouse.
- Useful for filtering and grouping by date components.
- Returns an integer for the specified date part.
Memory trick: Extract the parts, Truncate for periods, Format for display.
PySpark rangeBetween() for Time Windows
Flip cardA PySpark window frame clause (`WindowSpec.rangeBetween()`) used to define a window based on a time interval relative to the current row's value in the ordering column, rather than a fixed number of rows.
- Defines window using value/time offsets, not row offsets.
- Requires an `orderBy()` clause on a numeric or timestamp column.
- Commonly used with `F.expr('-INTERVAL X DAYS')` for time-based windows.
- Ideal for rolling averages/sums over specific periods (e.g., 'last 30 days').
Memory trick: Range Between Time, for each Partition's Average.
Spark SQL COUNT(*)
Flip cardThe `COUNT(*)` aggregate function in Spark SQL returns the total number of rows in a table or a specified group, including rows that contain NULL values in any column.
- Counts all rows, regardless of NULL values.
- Often the most efficient way to get a total row count.
- Equivalent to `COUNT(1)` in most SQL dialects.
- Does not require specifying a column name.
Memory trick: Count star for all; count column for non-nulls only!
SQL COUNT DISTINCT
Flip cardThe `COUNT(DISTINCT column_name)` aggregate function in SQL returns the number of unique, non-null values in a specified column.
- Counts only unique values.
- Ignores NULL values.
- Essential for unique item counts (e.g., unique users, products).
Memory trick: Count all, count not null, count distinct.
Data Strategy & Requirements Definition
Flip cardData strategy and requirements definition is the foundational phase of any data analytics project, focusing on understanding business objectives, data needs, consumption patterns, quality, and governance requirements.
- Aligns technical efforts with business value.
- Involves extensive stakeholder engagement.
- Covers data consumption, freshness, security, and quality expectations.
Memory trick: Plan first, then build, so your data dreams are fulfilled!
Azure Data Box
Flip cardAzure Data Box is a portfolio of physical and virtual appliances used for secure, offline data transfer to and from Azure storage, optimized for large datasets (terabytes to petabytes) to minimize network costs and transfer times.
- Physical appliance for large-scale data transfer.
- Minimizes network egress costs.
- Secure and reliable transfer process.
- Supports various data types and Azure storage destinations.
Memory trick: Data Box moves your terabytes, avoiding network pains.
Spark SQL LAG Function
Flip cardThe `LAG(column, offset, default)` window function in Spark SQL retrieves the value of a column from a row that is a specified `offset` number of rows before the current row within its partition.
- Used for comparing current row with previous rows.
- Requires `OVER (PARTITION BY ... ORDER BY ...)` clause.
- Essential for time-series analysis and pattern detection.
- Can specify a default value if no prior row exists.
Memory trick: Lag to look back, Lead to look forward, Windows for rows.
Spark SQL COUNT(DISTINCT)
Flip cardAn aggregate function in Spark SQL that counts the number of unique, non-NULL values in a specified column within a group or the entire result set.
- Counts only unique values.
- Ignores NULL values by default.
- Often used with `GROUP BY` or on the entire result set.
- Essential for measuring unique entities (e.g., unique users, distinct products).
Memory trick: Count Distinct, Filtered by Action and Recent Date.
Micro-batching for Streaming Ingestion
Flip cardMicro-batching is an ingestion pattern for streaming data where small, continuous data streams are collected into larger, time-based batches before being written to storage, mitigating the 'small file problem' and improving write/read efficiency.
- Reduces the number of files generated.
- Optimizes for distributed file systems.
- Balances latency with throughput efficiency.
Memory trick: Batch your streams, keep files big, and queries fast.
PySpark when().otherwise()
Flip cardA PySpark SQL function that implements conditional logic similar to SQL's CASE WHEN statement, allowing different values to be assigned to a column based on specified conditions.
- Used with `pyspark.sql.functions.when`.
- Allows chaining multiple `when()` conditions.
- Ends with `otherwise()` for a default value.
- Excellent for creating derived columns with complex logic.
Memory trick: When this is true, then that; Otherwise, default to the rest.
Spark SQL HAVING Clause
Flip cardThe Spark SQL HAVING clause is used to filter the groups of rows returned by a GROUP BY clause, based on aggregate conditions.
- Applied after GROUP BY and aggregate functions.
- Filters groups, not individual rows.
- Often includes aggregate functions in its condition.
Memory trick: Group and then Have a condition.
Delta Lake OPTIMIZE and ZORDER
Flip card`OPTIMIZE` consolidates small files into larger ones, improving read performance. `ZORDER` is an optional clause for `OPTIMIZE` that physically co-locates related data based on specified columns, accelerating data skipping for queries.
- OPTIMIZE reduces small file problem.
- ZORDER improves data skipping for query predicates.
- Best used on frequently queried columns or high-cardinality columns for filtering.
Memory trick: Optimize and ZORDER your Delta tables for lightning-fast queries.
PySpark Rolling Average Window
Flip cardA rolling average (or moving average) in PySpark is calculated using a window function that partitions data, orders it, and defines a frame (e.g., 7 preceding rows) over which an aggregate function (like `avg`) is applied.
- Requires `Window.partitionBy()`, `Window.orderBy()`.
- Uses `Window.rowsBetween()` or `Window.rangeBetween()` for the frame.
- Calculates aggregate (e.g., `avg`, `sum`) over the defined frame.
- Ideal for trend analysis over a moving period.
Memory trick: Partition, Order, Frame, Aggregate to get your window results.
Handling Null Values in Data Ingestion
Flip cardStrategies for addressing null values during data ingestion to ensure data quality and meet target schema constraints, often involving imputation, default values, or removing records based on business rules.
- Impacts data integrity and analysis.
- Requires business context to choose the best strategy.
- Can involve replacement, removal, or separate handling.
Memory trick: Don't lose data to nulls, replace or generate a new ID.
Histogram Use Case
Flip cardA histogram is a graphical representation of the distribution of numerical data. It is an estimate of the probability distribution of a continuous variable.
- Used to visualize the distribution of a single numerical variable.
- Data is grouped into 'bins' or ranges.
- The height of each bar represents the frequency or count of data points in that bin.
- Helps identify central tendency, spread, and shape of data distribution.
Memory trick: Histograms show how numbers are spread, like a bar chart for continuous data!
Fabric Data Pipelines for Ingestion
Flip cardMicrosoft Fabric Data Pipelines provide a scalable, serverless solution for orchestrating data movement and transformation activities, enabling automated ingestion into Lakehouses.
- Can be scheduled to run at specific intervals.
- Supports various data sources and destinations.
- Offers activities like 'Copy data' for efficient large-scale data transfer.
Memory trick: Automate data flow, watch it grow!
Scatter Chart for Correlation
Flip cardA Power BI visual type that plots individual data points on a two-dimensional plane, where each axis represents a numerical variable, used to display the relationship or correlation between the two variables.
- Shows relationship between two numerical variables.
- Each point represents an observation.
- Helps identify patterns, clusters, and outliers.
- Can be enhanced with trend lines or regression lines.
Memory trick: Scatter points for relationships, each dot tells a story.
Spark SQL SHUFFLE_HASH Join Hint
Flip cardA Spark SQL hint that explicitly tells the Catalyst optimizer to use a shuffle hash join strategy for the specified table(s) in a join operation.
- Forces Spark to use shuffle hash join.
- Beneficial for joining large tables.
- Requires shuffling data for both tables based on join keys.
- One table's partitions should ideally fit in memory after shuffling.
Memory trick: Hints Guide Spark's Joins: Hash for Shuffles, Broadcast for Small.
PySpark LAG() Window Function
Flip cardA PySpark window function that retrieves the value of an expression from a row that precedes the current row by a specified offset within its partition, ordered by a given column.
- Used to compare current row data with previous row data.
- Requires `Window.partitionBy()` and `orderBy()` for definition.
- Returns `NULL` by default if no preceding row is found.
- Essential for time-series analysis like calculating deltas.
Memory trick: Lag back one step, for each device's ordered path.
Column Chart for Distribution
Flip cardA Power BI visual type that uses vertical bars to represent the frequency or count of items within discrete categories or numerical bins, ideal for showing data distribution.
- Each bar represents a category or bin.
- Height of the bar indicates frequency or count.
- Effective for comparing magnitudes across distinct groups.
- Often used as a histogram for binned numerical data.
Memory trick: Bins need Columns to show their Counts, clearly and distinctly.
Spark SQL APPROX_PERCENTILE
Flip cardThe `APPROX_PERCENTILE(col, percentage)` function in Spark SQL calculates an approximate percentile of a numeric column, offering better performance for large datasets compared to exact methods.
- Provides an approximate result, not exact.
- More performant for large-scale data.
- Takes a column and a percentile value (0.0 to 1.0).
- Useful for exploratory analysis and metrics.
Memory trick: Approximate is fast, exact can be slow.
Incremental Load Ingestion
Flip cardIncremental load ingestion is a data loading strategy that processes only the data that has changed or been added since the last successful ingestion cycle, using mechanisms like timestamp columns or change data capture.
- Reduces data volume transferred and processed.
- Optimizes ingestion performance and resource usage.
- Requires a tracking mechanism (e.g., timestamp, sequence ID) in the source.
Memory trick: Full for all, Incremental for a small call!
Data Pipeline Copy Data Activity
Flip cardThe 'Copy data' activity in Microsoft Fabric Data Pipelines facilitates efficient and robust data ingestion from diverse sources to destinations, supporting various file formats, schema inference, and fault tolerance.
- Supports a wide range of connectors (e.g., SFTP, ADLS Gen2).
- Handles common file formats like CSV, Parquet.
- Includes options for schema inference and fault tolerance (skip/redirect malformed rows).
Memory trick: Copy data, handle errors, infer schema, no terrors!
PySpark F.hour()
Flip cardThe `F.hour()` function in PySpark (from `pyspark.sql.functions`) extracts the hour component (0-23) from a timestamp, date, or a string that can be parsed as such, returning it as an integer.
- Part of `pyspark.sql.functions` module.
- Takes a Column object (timestamp, date, or string) as input.
- Returns an integer representing the hour (0-23).
- Automatically handles string to timestamp conversion if format is standard.
Memory trick: Functions for parts, format for strings, substring if you're desperate!