AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium

A research institution collects genomic sequencing data from various instruments, generating files up to several terabytes in size. These files are initially stored on a network file system (NFS) in their on-premises data center. The institution needs to regularly move these files to an Amazon S3 Glacier Flexible Retrieval vault for long-term archival and cost-effective storage. The transfer process must be reliable, handle large files efficiently, and be scheduled to run weekly. Which AWS service should be used to automate this ingestion process?

  1. AAWS Storage Gateway (File Gateway)
  2. BAWS Direct Connect
  3. CAWS DataSync
  4. DAmazon S3 Transfer Acceleration
Show answer & explanation

Correct answer: C. AWS DataSync

AWS DataSync is designed for automated, accelerated, and secure data transfer between on-premises storage (including NFS) and AWS storage services like S3, including S3 Glacier. It can handle large files, be scheduled for regular transfers, and ensures transfer reliability, making it perfect for this archival use case.

Why the other options are wrong

  • A. Storage Gateway (File Gateway) provides NFS/SMB access to S3, but it's more for hybrid cloud storage access than for one-time or scheduled large-scale transfers to Glacier.
  • B. Direct Connect provides a dedicated network connection to AWS, improving bandwidth and reducing latency, but it is an underlying network service, not a data transfer automation service itself.
  • D. S3 Transfer Acceleration speeds up transfers to S3 over the internet, but it's not a service for automating or scheduling transfers from on-premises file systems directly.

AWS DataSync

A data transfer service that makes it easier for you to automate moving data between on-premises storage and Amazon S3, Amazon EFS, or Amazon FSx for Windows File Server.

  • Supports various on-premises sources (NFS, SMB, HDFS).
  • Accelerates transfers using a proprietary protocol.
  • Can transfer directly to S3 storage classes like Glacier.

Memory trick: DataSync archives on-prem files reliably to Glacier.

More Data Ingestion and Transformation questions