Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Medium

A manufacturing company uses a legacy system that generates daily production logs as fixed-width text files. Each file contains hundreds of thousands of records. The data engineer needs to ingest these files into a Fabric Lakehouse, parse the fixed-width format, and perform some basic column renaming and data type conversions efficiently. Which approach is most suitable for this task?

  1. AUsing Dataflows Gen2 with the Text/CSV connector and manually splitting columns.
  2. BDeveloping a custom C# application to parse and upload the files to a staging area.
  3. CImplementing a Data Pipeline with a Copy Data activity to ingest as-is, then a Spark Notebook for parsing.
  4. DUsing a Spark Notebook with PySpark to read the fixed-width files and transform them.
Show answer & explanation

Correct answer: D. Using a Spark Notebook with PySpark to read the fixed-width files and transform them.

Spark Notebooks with PySpark provide powerful capabilities for handling complex file formats like fixed-width text. Libraries like `pyspark.sql.functions` or custom parsing logic can be used to efficiently extract data based on column positions, apply transformations, and then load into a Lakehouse, leveraging Spark's distributed processing for large files.

Why the other options are wrong

  • A. While Dataflows Gen2 can read text files, manually splitting fixed-width columns in Power Query can be tedious and less efficient for hundreds of thousands of records across many columns compared to programmatic Spark parsing.
  • B. Developing a custom application outside Fabric adds unnecessary complexity, maintenance overhead, and doesn't leverage Fabric's integrated data engineering capabilities.
  • C. Ingesting as-is with Copy Data and then parsing with Spark is a valid two-step process, but a Spark Notebook can directly handle both reading and parsing of complex formats in one step more efficiently for large files.

PySpark Fixed-Width Parsing

PySpark in a Spark Notebook is well-suited for parsing fixed-width text files. It allows programmatic extraction of data based on character positions, followed by efficient distributed transformations and loading into a Lakehouse.

  • Handles complex file formats programmatically.
  • Leverages Spark's distributed processing for large files.
  • Can use string manipulation functions (`substring`) or regex.
  • Integrates seamlessly with Lakehouse for output.

Memory trick: Fixed-width files parsed by PySpark for the Lakehouse.

More Prepare and transform data (20-25%) questions