Microsoft Certified: Fabric Analytics Engineer Associate practice questions

216 free questions with answers and explanations.

Practice test
  1. 1.A company is migrating its entire data estate to Microsoft Fabric. As part of their data governance strategy, they require that all data accessed by Fabric items (e.g., Lakehouses, Data Warehouses, Dataflows) must be encrypted both at rest and in transit. Which security feature is inherently provided by Microsoft Fabric to meet these encryption requirements without additional configuration by the user?Govern and administer Fabric (10-15%)
  2. 2.A data analytics department has multiple teams, each working on distinct projects within Microsoft Fabric. They want to implement a solution where each team has its own isolated environment for development and testing, ensuring that changes made by one team do not accidentally affect another. Furthermore, they need a clear separation of billing for the compute resources consumed by each team. How should the Fabric administrator structure their environment to meet these requirements efficiently?Govern and administer Fabric (10-15%)
  3. 3.A data analytics team needs to integrate a Microsoft Fabric workspace with their existing version control system, Azure DevOps Git, to manage code for notebooks and data pipelines. They want to ensure that all changes are tracked, and a proper branching and merging workflow is followed before deploying to production. Which Fabric feature enables this integration?Govern and administer Fabric (10-15%)
  4. 4.A data engineering team uses a Microsoft Fabric workspace for developing and testing new data pipelines. They frequently create and delete various items like Lakehouses, Notebooks, and Data Pipelines. To maintain an organized and clean workspace, the team wants to automatically remove items that have not been modified for 90 days. Which Fabric feature allows an administrator to configure this automated cleanup?Govern and administer Fabric (10-15%)
  5. 5.An organization is using Microsoft Fabric for its data analytics needs and wants to implement a robust monitoring solution. They need to track the performance and resource consumption of their Fabric capacities, identify bottlenecks, and optimize workload distribution across different engines (e.g., Spark, SQL, Power BI). Which built-in Fabric tool is specifically designed for these monitoring and optimization tasks?Govern and administer Fabric (10-15%)
  6. 6.A global manufacturing company uses Microsoft Fabric across several regions. They have a central data engineering team responsible for creating and maintaining core data models and pipelines, which are then deployed to regional workspaces. These regional workspaces are managed by local teams. The central team needs to ensure that changes to the core items are propagated consistently and reliably to all regional workspaces, with proper version control and approval workflows. Which Fabric feature is best suited for this scenario?Govern and administer Fabric (10-15%)
  7. 7.A Fabric administrator needs to ensure that all data engineers working on a project adhere to a specific set of coding standards for Python notebooks and Spark definitions. They also want to facilitate code reviews and ensure that only approved code changes are merged into the main development branch. Which Fabric feature, combined with an external service, is best suited to enforce these practices?Govern and administer Fabric (10-15%)
  8. 8.A data engineering team at a large enterprise is developing several data pipelines in Microsoft Fabric. They need to ensure that all workspaces adhere to specific organizational standards for data classification and sensitivity labels. Which Fabric feature should they leverage to enforce these standards consistently across all new and existing workspaces?Govern and administer Fabric (10-15%)
  9. 9.A data platform administrator is setting up a new Microsoft Fabric environment. The organization requires that all new workspaces default to a specific set of data retention policies and sensitivity labels. They also want to restrict the types of data sources that can be connected to Fabric items globally. Which administrative setting in the Fabric Admin portal should the administrator configure to achieve these organization-wide defaults and restrictions?Govern and administer Fabric (10-15%)
  10. 10.A financial services company uses Microsoft Fabric for its analytics platform. Due to strict regulatory compliance, all data processing activities must be logged and audited for at least seven years. The company needs to monitor all user and system activities within Fabric, including item creation, modification, and access. Which Fabric monitoring capability should the administrator configure to meet this requirement?Govern and administer Fabric (10-15%)
  11. 11.A data platform team is managing a critical production Lakehouse in Microsoft Fabric. They need to implement a security measure that ensures data within the Lakehouse can only be accessed by specific service principals and users, even if they have broader permissions at the workspace level. This access control must be granular, applying directly to tables and files within the Lakehouse, and should respect column-level and row-level security definitions. Which security model should the team leverage?Govern and administer Fabric (10-15%)
  12. 12.A data administrator needs to restrict the ability of users to create new workspaces in Microsoft Fabric. The goal is to centralize workspace creation through a dedicated team to ensure compliance with naming conventions, capacity assignments, and data governance policies. Which tenant setting should the administrator configure to achieve this?Govern and administer Fabric (10-15%)
  13. 13.A data platform administrator is setting up a new Microsoft Fabric environment for a company that prioritizes data residency. They need to ensure that all data processed and stored within Fabric remains within a specific geographical region to comply with local regulations. Which setting or configuration is paramount to achieve this data residency requirement for all new Fabric workspaces and items?Govern and administer Fabric (10-15%)
  14. 14.A data analytics team has built several critical reports and dashboards in a Microsoft Fabric workspace. They want to ensure that these items remain available even if the underlying data sources or connected services experience temporary outages. They need a mechanism to quickly revert to a previous, stable state of these reports and dashboards if issues arise after a new deployment. Which Fabric feature should they primarily rely on for this recovery capability?Govern and administer Fabric (10-15%)
  15. 15.A company is migrating its data analytics workloads to Microsoft Fabric. They have a strict policy that all sensitive data must be encrypted at rest using customer-managed keys (CMK) for enhanced security and compliance. How can the Fabric administrator ensure that all data stored in Fabric workspaces associated with a specific F-SKU capacity is encrypted using CMK?Govern and administer Fabric (10-15%)
  16. 16.A Microsoft Fabric administrator needs to ensure that all data engineers adhere to specific naming conventions for new Lakehouses created within their workspaces to maintain consistency and facilitate easier management. The administrator wants to enforce this policy automatically when a new Lakehouse is provisioned. Which approach should the administrator implement?Govern and administer Fabric (10-15%)
  17. 17.A global company uses Microsoft Fabric for its analytics platform. Due to strict data governance requirements, they need to track all administrative and user activities within Fabric, such as workspace creation, item deletion, and permission changes. This information is crucial for security audits and compliance. Where should the administrator look to find a comprehensive record of these activities?Govern and administer Fabric (10-15%)
  18. 18.A data analytics team is developing several critical reports and dashboards within a Microsoft Fabric workspace. They need to ensure that specific sensitive datasets, which are used in these reports, are not accidentally shared publicly or with unauthorized external users, even if the workspace itself has broader sharing settings. Which Fabric feature should the team leverage to enforce this granular control?Govern and administer Fabric (10-15%)
  19. 19.A Microsoft Fabric administrator is responsible for managing multiple capacities across different environments (development, test, production). They need to ensure that the development environment capacity does not consume an excessive amount of resources, potentially impacting the test or production environments. The administrator wants to set a maximum resource usage limit specifically for the development capacity. Which Fabric feature enables this type of resource governance?Govern and administer Fabric (10-15%)
  20. 20.A large enterprise uses Microsoft Fabric across multiple business units. Each business unit has its own dedicated workspace and needs to track the compute consumption of its Fabric items (e.g., Lakehouses, Data Warehouses, Notebooks) to allocate costs accurately. The central IT team needs a consolidated view of resource utilization across all business units. Which tool or method is BEST suited for this requirement?Govern and administer Fabric (10-15%)
  21. 21.A company is implementing a new data analytics solution in Microsoft Fabric. They want to ensure that access to sensitive data stored in a Lakehouse is restricted based on the user's department. Specifically, users from the 'Sales' department should only see sales-related data, and users from 'Marketing' should only see marketing-related data, even when accessing the same Lakehouse table. Which security mechanism in Fabric is MOST appropriate for this scenario?Govern and administer Fabric (10-15%)
  22. 22.A global conglomerate operates multiple subsidiaries, each with its own data governance requirements. They are implementing Microsoft Fabric and need to provision dedicated capacity for each subsidiary to ensure resource isolation and chargeback. Each subsidiary's capacity must be configured with specific region choices and workload settings to optimize for their unique data processing needs. Which capability within Microsoft Fabric allows for this granular management of dedicated resources?Govern and administer Fabric (10-15%)
  23. 23.A data engineering team is developing a critical data pipeline in a Microsoft Fabric workspace. They want to ensure that only authorized users can access the underlying data in the Lakehouse, and that specific columns containing personally identifiable information (PII) are completely hidden from certain roles, even if those roles have access to the table. Which security mechanism should be implemented for the PII columns?Govern and administer Fabric (10-15%)
  24. 24.A data analyst is querying a large Spark Delta table `web_events` in Microsoft Fabric. The table contains `event_id` (string), `user_id` (string), `event_timestamp` (timestamp), and `page_url` (string). The analyst needs to find the immediately preceding `page_url` visited by each user before the current event. Which Spark SQL window function should be used?Explore and analyze data (15-20%)
  25. 25.A data engineer is working with a Delta Lake table named `sensor_readings` in Microsoft Fabric. The table contains columns `device_id` (string), `timestamp` (timestamp), and `temperature` (double). The engineer needs to calculate the average temperature for each device, considering only the readings from the *previous 30 minutes* relative to the current reading. Which PySpark window function clause should be used to define this time-based window?Explore and analyze data (15-20%)
  26. 26.A data analyst needs to retrieve all customer records from a `Customers` table where the `Country` column is 'USA' and the `LastPurchaseDate` is after January 1, 2023. Which SQL clause should be used to filter the rows based on these conditions?Explore and analyze data (15-20%)
  27. 27.A data engineer is working with a large Delta Lake table in a Microsoft Fabric Lakehouse that stores historical sensor readings. Over time, the table has accumulated many small files due to frequent micro-batch ingests, leading to suboptimal query performance. The engineer needs to consolidate these small files into larger, more manageable ones to improve read performance without altering the data. Which Delta Lake command should the engineer execute?Plan and implement data analytics solutions (10-15%)
  28. 28.A data engineer is working with a large Spark Delta table `product_reviews` in Microsoft Fabric. Each review has a `review_text` column. They need to create a new derived column `sentiment_category` based on whether the `review_text` contains specific keywords like 'excellent', 'good', 'bad', or 'poor'. If 'excellent' or 'good' is present, it should be 'Positive'; if 'bad' or 'poor', it should be 'Negative'; otherwise, 'Neutral'. Which PySpark function is most suitable for this conditional logic and string matching?Explore and analyze data (15-20%)
  29. 29.A data engineering team is working with a large dataset in a Spark Delta Lake table. They need to calculate the average `transaction_amount` for each `product_category` and then filter out categories where the average amount is less than $100. Which Spark SQL clause should be used to filter the grouped results?Explore and analyze data (15-20%)
  30. 30.A data analyst needs to query historical sales data stored in a Microsoft Fabric Lakehouse. The data is partitioned by `year` and `month` in a Delta table. The analyst frequently queries data for a specific quarter. To optimize query performance, which file organization strategy within the Lakehouse should be recommended for the Delta table?Plan and implement data analytics solutions (10-15%)
  31. 31.A company is ingesting real-time sensor data into a Microsoft Fabric Lakehouse. The sensor data arrives as small, frequent messages, and needs to be stored in a Delta table. Due to the high volume and velocity, the data engineer is concerned about the 'small file problem' impacting query performance later. Which ingestion pattern should be used to mitigate this issue effectively?Plan and implement data analytics solutions (10-15%)
  32. 32.A data analyst is querying a large Spark Delta table `customer_activity` in Microsoft Fabric. They need to count the number of unique customers who performed an 'add_to_cart' action within the last 7 days. Which Spark SQL query should they use?Explore and analyze data (15-20%)
  33. 33.A data engineering team is designing a Lakehouse solution in Microsoft Fabric for a global e-commerce platform. They need to ingest product catalog data from various sources, including legacy databases and external vendor APIs. The ingested data needs to be validated for schema compliance and data quality rules (e.g., product IDs must be unique, prices must be positive) before being made available for downstream analytics. Where should these validation steps primarily occur within the Medallion Architecture layers?Plan and implement data analytics solutions (10-15%)
  34. 34.A financial analyst is reviewing stock data in a Spark Delta table. They need to identify periods where the `closing_price` of a stock continuously increased for at least 3 consecutive days. Which advanced Spark SQL technique would be most effective for this pattern detection?Explore and analyze data (15-20%)
  35. 35.A data engineer needs to move a large volume of historical data (several terabytes) from an on-premises Hadoop Distributed File System (HDFS) to a Microsoft Fabric Lakehouse. The data consists of various file formats including Parquet, ORC, and CSV. The transfer needs to be secure, reliable, and minimize network egress costs from the on-premises environment. Which Microsoft Azure service, integrated with Fabric, is best suited for the initial bulk transfer of this data?Plan and implement data analytics solutions (10-15%)
  36. 36.A data engineer is working with a PySpark DataFrame named `transactions_df` that contains `transaction_id` (string), `product_category` (string), and `amount` (double). The engineer needs to calculate the total `amount` for each `product_category` and then filter out categories where the total `amount` is less than 1000. Which PySpark operation sequence correctly achieves this?Explore and analyze data (15-20%)
  37. 37.A company is migrating its on-premises data warehouse to Microsoft Fabric. They have a nightly batch process that extracts data from an SAP system as CSV files and stores them on an SFTP server. The data needs to be ingested into a Lakehouse, transformed, and then loaded into a Delta table for analytical purposes. Which Fabric component is best suited for orchestrating this end-to-end data ingestion and transformation workflow?Plan and implement data analytics solutions (10-15%)
  38. 38.A data architect is planning the ingestion of real-time streaming data from IoT devices into a Microsoft Fabric Lakehouse. The data arrives continuously in small, frequent batches. To ensure data consistency and prevent partial reads by downstream consumers, each micro-batch must be committed atomically to the Lakehouse table. Which Delta Lake feature provides this guarantee?Plan and implement data analytics solutions (10-15%)
  39. 39.A data analytics team is planning to implement a Microsoft Fabric Lakehouse for their customer 360 initiative. They have identified various data sources, including CRM systems, website clickstream data, and social media feeds. Before designing the technical solution, the team needs to clearly define the data consumption patterns, required data freshness, and data security requirements for different business units. What is the crucial initial step they must complete?Plan and implement data analytics solutions (10-15%)
  40. 40.A data engineer is optimizing a Spark SQL query that frequently joins a large `FactSales` table with a smaller `DimProduct` table. Both tables are stored in a Delta Lake. To improve join performance, the engineer wants to ensure the smaller table is broadcasted during the join operation. Which Spark SQL hint should be used for this purpose?Explore and analyze data (15-20%)
  41. 41.A data analyst needs to count the number of unique customers from a `customer_orders` table in a Fabric Lakehouse SQL endpoint. The table has a `customer_id` column. Which SQL aggregate function should they use?Explore and analyze data (15-20%)
  42. 42.A data analyst is querying a large Spark Delta table containing customer order data. They need to find the top 5 customers by total order value. The table has `customer_id` and `order_value` columns. Which combination of Spark SQL clauses will achieve this?Explore and analyze data (15-20%)
  43. 43.A data engineer needs to ingest a large volume of CSV files from an external SFTP server into a Microsoft Fabric Lakehouse. The CSV files contain header rows and are comma-delimited. The ingestion process must be robust, handle potential malformed records by quarantining them, and allow for schema inference. Which Data Pipeline activity should be used for this scenario?Plan and implement data analytics solutions (10-15%)
  44. 44.A data engineer is writing a PySpark script to process a large DataFrame `log_data_df` containing web server logs. The DataFrame has a `timestamp` column (string, in 'YYYY-MM-DD HH:MM:SS' format) and `request_path` (string). The engineer needs to extract the hour of the day from the `timestamp` column as an integer for further analysis. Which PySpark function should be used?Explore and analyze data (15-20%)
  45. 45.A retail company is migrating its historical sales data, totaling 50 terabytes, from an on-premises SQL Server database to a Microsoft Fabric Lakehouse. The migration needs to be completed within a strict two-week deadline, and network bandwidth is a significant constraint, making direct online transfer impractical for the entire dataset. What is the most appropriate method for ingesting this data into Microsoft Fabric?Plan and implement data analytics solutions (10-15%)
  46. 46.A data engineer is optimizing a PySpark script that processes a large Delta table named `event_logs`. They frequently need to extract the year and month from a `timestamp` column for filtering and grouping. Which PySpark function is the most efficient and idiomatic way to achieve this for both year and month extraction?Explore and analyze data (15-20%)
  47. 47.A global manufacturing company uses Microsoft Fabric to manage its supply chain data. They have a central Lakehouse and several regional Lakehouses, all within the same Fabric workspace. Data from regional Lakehouses needs to be aggregated into the central Lakehouse for global reporting. The company wants to avoid physically copying data and instead allow the central Lakehouse to directly access the regional data. Which feature should be implemented to achieve this?Plan and implement data analytics solutions (10-15%)
  48. 48.A data analyst is preparing a report on customer demographics. They have a Spark DataFrame named `customer_df` with columns `CustomerID`, `Age`, and `City`. They need to group customers by `City` and then calculate the average `Age` for each city. Which PySpark operation should be used after `groupBy('City')` to compute the average?Explore and analyze data (15-20%)
  49. 49.A data engineer is working with a Microsoft Fabric Lakehouse. They have ingested raw CSV files into the 'Files' section of the Lakehouse. Now, they need to create a managed Delta table from these CSV files in the 'Tables' section. They want to incrementally load new CSV files that arrive daily into this Delta table, ensuring schema evolution is handled gracefully. Which approach is most suitable for this task?Plan and implement data analytics solutions (10-15%)
  50. 50.A data engineering team is planning to ingest data from an on-premises SQL Server database into a Microsoft Fabric Lakehouse. The database contains several tables, and for some tables, only new or modified records need to be ingested daily to minimize data transfer and processing. Which ingestion pattern is best suited for this requirement?Plan and implement data analytics solutions (10-15%)