AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium

A research institution collects genomic sequencing data from various instruments, generating files up to 100 GB each. This data needs to be securely transferred from on-premises storage to Amazon S3 for long-term archival and analysis. The transfer must be highly optimized for network throughput, fault-tolerant, and support scheduled, incremental transfers. Which AWS service is specifically designed for this type of data transfer?

  1. AAmazon Kinesis Data Streams
  2. BAWS DataSync
  3. CAmazon S3 API Gateway
  4. DAWS Transfer Family
Show answer & explanation

Correct answer: B. AWS DataSync

AWS DataSync is a data transfer service that makes it easier for you to automate moving data between on-premises storage and Amazon S3, Amazon EFS, or Amazon FSx. It is optimized for large-scale, high-performance transfers, supports scheduling, filtering for incremental transfers, and provides end-to-end data integrity.

Why the other options are wrong

  • A. Amazon Kinesis Data Streams is for real-time streaming data ingestion, not for transferring large files in batches or incrementally from on-premises storage.
  • C. Amazon S3 API Gateway is a non-existent service; API Gateway is for creating, publishing, maintaining, monitoring, and securing APIs, not for data transfer to S3.
  • D. AWS Transfer Family is for transferring files over SFTP, FTPS, and FTP protocols, primarily for external partners, not typically for optimized, internal, large-scale storage synchronization.

AWS DataSync

A data transfer service that simplifies, automates, and accelerates moving data between on-premises storage and AWS storage services.

  • Optimized for large-scale transfers (files/objects)
  • Supports scheduled and incremental transfers
  • Ensures data integrity and security

Memory trick: DataSync synchronizes data like a pro.

More Data Ingestion and Transformation questions