AWS Certified Data Engineer – Associate flashcards
118 free flashcards. Tap a card to flip it.
Glue Data Catalog Eventual Consistency
Flip cardAWS Glue Data Catalog metadata updates are eventually consistent, meaning changes might take time to propagate across all regions or endpoints. Applications querying immediately after an update might receive stale data.
- Metadata updates are not immediately consistent.
- Can lead to schema mismatches or missing partitions.
- Retry with exponential backoff is the recommended handling.
- Avoids arbitrary delays and ensures eventual success.
Memory trick: Eventually Consistent: Retry, wait, then try again, or your data will be in a state of 'maybe'.
Glue Redshift Write Optimization
Flip cardOptimize AWS Glue writes to Amazon Redshift by using 'use_s3_dist_copy' and 'num_partitions' to stage data in S3 and leverage Redshift's parallel COPY command, reducing cluster load.
- S3 staging + Redshift COPY is highly efficient.
- Reduces direct Redshift cluster load.
- Parameters: use_s3_dist_copy, num_partitions/num_files.
Memory trick: Glue's best Redshift friend is S3, for a fast 'COPY' party.
AWS Database Migration Service (DMS)
Flip cardA cloud service that makes it easy to migrate relational databases, data warehouses, NoSQL databases, and other types of data stores.
- Supports homogeneous and heterogeneous migrations.
- Enables continuous data replication (Change Data Capture).
- Automates schema and data type conversions.
Memory trick: DMS moves and mirrors databases, even different types.
Glue ETL for Data Lake Refinement
Flip cardUtilizing AWS Glue ETL jobs to transform raw data in a data lake into optimized formats (e.g., Parquet, ORC) and structures (e.g., partitioning) for improved query performance and reduced costs.
- Serverless Apache Spark environment.
- Integrates with Glue Data Catalog for schema management.
- Supports various data sources and targets.
- Ideal for batch processing and large-scale transformations.
Memory trick: Glue's ETL is the Gold Standard for transforming your data lake's CSVs to Parquet, saving bucks and boosting speed.
Kinesis-Lambda Throttling
Flip cardResolve Kinesis-Lambda throttling by increasing Lambda's `ConcurrentExecutions` limit to match stream throughput and using `On-demand` concurrency for cost-effective, adaptive scaling.
- Throttling means insufficient Lambda parallelism.
- `ConcurrentExecutions` directly controls parallelism.
- `On-demand` concurrency scales efficiently with load.
Memory trick: More Lambda 'Concurrence' on 'Demand' keeps Kinesis flowing free.
AWS Data-in-Transit Encryption
Flip cardAWS services automatically encrypt data in transit using SSL/TLS when communicating over public endpoints, ensuring secure communication between services by default.
- SSL/TLS is the default encryption protocol.
- Applies to communication over public endpoints.
- Requires no explicit configuration for default behavior.
Memory trick: Default AWS encryption is like a secure tunnel for your data.
Kinesis-Lambda Error Handling
Flip cardFor Lambda functions invoked by Kinesis Data Streams, configure the event source mapping's retry attempts and 'On failure' destination to automatically handle transient errors and send failed records for reprocessing or analysis.
- Kinesis invokes Lambda synchronously.
- Event source mapping controls retries and error destinations.
- Ensures no data loss and prevents stream blocking.
Memory trick: Kinesis event mapping catches the 'fail' and gives it another 'try'.
Proactive Pipeline Automation (CloudWatch + EventBridge)
Flip cardA pattern for automating operational responses in AWS data pipelines, where CloudWatch monitors metrics (including anomaly detection), Alarms trigger on deviations, and EventBridge orchestrates custom remediation (e.g., Lambda) and notification workflows.
- CloudWatch Anomaly Detection learns normal metric patterns.
- CloudWatch Alarms can trigger on anomalies or thresholds.
- EventBridge acts as a central event bus for routing alarm events.
- Lambda functions enable custom, programmatic remediation actions.
Memory trick: Anomaly Detected by CloudWatch, Alarm Triggers, EventBridge Routes to Lambda for Fix.
Lambda Throttling
Flip cardOccurs when an AWS Lambda function attempts to execute more concurrent invocations than its configured concurrency limit or the account's regional concurrency quota allows, leading to invocation failures.
- Default concurrency limit per function is 1000, shared across all functions in a region.
- Can be configured at the function level (reserved concurrency) or account level.
- Throttled invocations are typically retried by event sources like Kinesis (with a delay).
Memory trick: When Lambda's lanes are jammed, more lanes let the cars zoom past without crashing.
Glue Job Optimization (Code-Level)
Flip cardTechniques applied within an AWS Glue ETL script (e.g., using PySpark/Scala) to improve execution efficiency, reduce resource consumption, and decrease runtime, primarily by optimizing data transformations and I/O operations.
- Use `pushdown_predicate` for S3/JDBC sources.
- Choose efficient Spark transformations (e.g., `filter` before `join`).
- Cache frequently used DataFrames.
- Optimize partitioning strategies for writes.
Memory trick: Optimize the Code, then Scale the Load; save your cash, make your job fast.
Glue Security Configuration
Flip cardAn AWS Glue Security Configuration centralizes encryption settings for Glue jobs, allowing consistent application of KMS CMKs for temporary S3 data, CloudWatch Logs, and Job Bookmarks.
- Centralized encryption management for Glue jobs.
- Applies to S3 temporary data, CloudWatch Logs, Job Bookmarks.
- Ensures consistent use of KMS CMKs.
Memory trick: Glue's Security Config is the 'Key Master' for temporary data.
DataSync Troubleshooting with CloudWatch Logs
Flip cardUtilizing AWS CloudWatch Logs to access detailed execution logs generated by AWS DataSync tasks, enabling identification of specific errors and root causes for transfer failures.
- DataSync tasks send logs to CloudWatch Logs.
- Logs contain information on file transfers, errors, network issues, and agent health.
- Essential for diagnosing transfer failures and performance issues.
Memory trick: DataSync fails? CloudWatch Logs will tell you its tales.
Serverless SFTP to S3 ETL Pipeline
Flip cardAn automated pipeline using AWS Transfer Family for SFTP ingestion to S3, S3 Event Notifications or EventBridge for trigger, Lambda for orchestration, and AWS Glue ETL for serverless data transformation.
- Transfer Family provides managed SFTP to S3.
- S3 events trigger automation.
- Lambda orchestrates Glue ETL jobs.
- Fully serverless and automated.
Memory trick: Transfer Family uploads, EventBridge triggers, Lambda calls Glue, data transforms.
Kinesis Data Streams + Kinesis Data Analytics (Flink)
Flip cardA serverless combination for real-time data ingestion and advanced stream processing, enabling complex aggregations, filtering, and analysis on high-volume data streams.
- Kinesis Data Streams: high-throughput, low-latency ingestion.
- Kinesis Data Analytics (Flink): serverless, real-time stream processing.
- Ideal for complex event processing, aggregations, and real-time dashboards.
Memory trick: Kinesis streams, Flink aggregates, Redshift reports.
AWS Glue ETL (Batch)
Flip cardA serverless data integration service that makes it easy to discover, prepare, and combine data for analytics. It's ideal for scheduled, batch-oriented Extract, Transform, and Load (ETL) jobs.
- Serverless: no infrastructure to manage.
- Spark-based: scalable for large datasets.
- Integrated with Data Catalog for metadata management.
Memory trick: Glue's batch jobs make files pristine.
AWS IoT Core Rules Engine
Flip cardA feature of AWS IoT Core that allows users to define rules that process and route messages from IoT devices to other AWS services based on message content, without needing to provision or manage servers.
- Enables filtering, transforming, and routing IoT messages.
- Supports SQL-like syntax for rule definition.
- Integrates with numerous AWS services as destinations (S3, Kinesis, Lambda, etc.).
Memory trick: IoT Core takes device messages, Rules Engine filters and routes them perfectly.
AWS Snowball Edge Storage Optimized
Flip cardA physical device with significant storage and compute capabilities for transferring large amounts of data into and out of AWS, especially when network bandwidth is limited or non-existent.
- Designed for petabyte-scale data transfers.
- Offers local storage and compute for edge processing.
- Bypasses internet for data transfer, shipped physically.
Memory trick: Snowball rolls data fast when the internet is slow or gone.
AWS Glue ETL (Spark)
Flip cardA serverless data integration service that provides a fully managed Apache Spark environment for performing complex ETL (Extract, Transform, Load) operations on large datasets.
- Serverless Apache Spark environment.
- Handles complex transformations, enrichments, aggregations.
- Highly available and fault-tolerant by design.
Memory trick: Glue ETL runs Spark, serverless and smart.
AWS Step Functions + Lambda + Elastic Transcoder
Flip cardA serverless architecture for orchestrating complex, multi-step workflows, combining Lambda for custom logic and Elastic Transcoder for media processing.
- Step Functions: Visual workflow orchestration, state management, error handling.
- Lambda: Serverless compute for event-driven functions.
- Elastic Transcoder: Cloud-based media transcoding service.
Memory trick: Steps orchestrate media, Lambda and Transcoder free.
CloudTrail + Kinesis Data Firehose for S3 Auditing
Flip cardAWS CloudTrail captures S3 data events (object-level API activity) for auditing, and Amazon Kinesis Data Firehose reliably delivers these audit logs to a centralized logging system.
- CloudTrail tracks API calls and S3 data events.
- S3 data events capture object-level reads/writes.
- Firehose delivers high-volume streaming data to destinations.
Memory trick: CloudTrail watches S3, Firehose delivers the evidence fast.
Serverless File Ingestion & Processing
Flip cardA pattern for receiving files from external sources, triggering an event-driven function to process their content, and storing the extracted data in a structured database, all without managing servers.
- Uses services like AWS Transfer Family for secure file reception.
- Leverages S3 events to trigger processing functions (e.g., Lambda).
- Ideal for event-driven, small-to-medium file processing.
Memory trick: Transfer Family brings the XML, Lambda parses the data, DynamoDB stores it fast.
AWS IoT Core + IoT Rules Engine
Flip cardAWS IoT Core enables secure device connection and management, while the IoT Rules Engine processes and routes messages from connected devices to other AWS services based on defined rules.
- IoT Core: Connects billions of IoT devices securely.
- IoT Rules Engine: Filters, transforms, and routes device data.
- Supports various AWS service integrations (S3, DynamoDB, Lambda, Kinesis, etc.).
Memory trick: IoT devices connect, Rules Engine directs.
Kinesis Data Streams + Kinesis Data Analytics
Flip cardA powerful combination for real-time data ingestion and processing. Kinesis Data Streams captures data, and Kinesis Data Analytics (with Flink) performs sophisticated real-time analytics.
- Kinesis Data Streams: highly scalable, durable, real-time data capture.
- Kinesis Data Analytics: executes SQL or Apache Flink applications on streaming data.
- Ideal for real-time dashboards, anomaly detection, and interactive analytics.
Memory trick: Streams flow, Flink thinks, insights grow.
Kinesis Data Streams Encryption
Flip cardAmazon Kinesis Data Streams provides encryption for data both in transit (using SSL/TLS) and at rest (using Server-Side Encryption with AWS KMS).
- In-transit encryption is automatic via SSL/TLS endpoints
- At-rest encryption uses Server-Side Encryption (SSE) with AWS KMS
- Customers can choose AWS managed KMS keys or customer managed KMS keys for SSE
Memory trick: Kinesis is secure: SSL in transit, KMS at rest.
Lake Formation Column-Level Security & Data Filters
Flip cardAWS Lake Formation provides column-level security to restrict access to specific columns, and data filters can dynamically mask or redact sensitive columns at query time for different users.
- Grant access to a subset of columns in a table
- Data filters allow dynamic redaction/masking on sensitive columns
- Enforced at query time without modifying source data
Memory trick: Columns and filters redact PII dynamically.
Redshift Security Best Practices
Flip cardSecuring Redshift involves using temporary credentials, IAM roles for authentication, and fine-grained permissions within Redshift to control data access.
- Use AWS Secrets Manager for database credential rotation
- Utilize IAM roles for Redshift authentication for applications
- Apply Redshift's internal permissions (GRANT/REVOKE) for table-level access control
Memory trick: Secrets for rotation, IAM for roles, Redshift for tables.
S3 Lifecycle Policy Tag Filters
Flip cardS3 Lifecycle policies can use object tags as filters to apply rules to a subset of objects within a bucket, enabling fine-grained control over storage class transitions and object expiration.
- Apply rules to objects matching specific key-value tag pairs
- Allows selective management of objects within a single bucket
- Combines with object age for precise actions
Memory trick: Tags filter S3 lifecycle to delete selected videos.
Lake Formation Row-Level Security
Flip cardAWS Lake Formation row-level security allows data engineers to define policies that restrict access to specific rows in a table based on conditions, applied dynamically at query time.
- Filters data based on column values (predicates)
- Enforced at query time by integrated services like Athena
- Does not modify the underlying data in S3
Memory trick: Lake Formation's rows filter data for team access.
AWS CloudTrail
Flip cardAWS CloudTrail is a service that enables governance, compliance, operational auditing, and risk auditing of your AWS account. It logs API calls and related events.
- Records actions taken by a user, role, or an AWS service
- Logs are delivered to an S3 bucket
- Provides event history of your AWS account activity
Memory trick: CloudTrail logs, S3 stores, for audit success.
S3 Lifecycle Policy with Tag Filters
Flip cardS3 lifecycle policies can leverage object tags to apply distinct storage class transitions and expiration rules to different groups of objects within the same bucket, offering highly flexible data retention management.
- Tags provide flexible metadata for categorization
- Rules apply only to objects matching specified tags
- Enables diverse retention policies within a single bucket
Memory trick: Tags filter lifecycle rules for diverse retention needs.
AWS Lake Formation Data Filters
Flip cardAWS Lake Formation data filters provide fine-grained access control to data lake resources, including column-level and row-level filtering and dynamic data masking.
- Control access to specific columns or rows
- Mask or redact sensitive data dynamically at query time
- Applied to various AWS analytics services like Athena, Redshift Spectrum, EMR
Memory trick: Lake Formation filters mask data for PHI protection.
SSE-S3
Flip cardServer-Side Encryption with S3 managed keys (SSE-S3) encrypts S3 objects using keys managed by AWS. It is the easiest way to encrypt data at rest in S3.
- AWS manages the encryption keys
- Objects are encrypted before saving to disk and decrypted when retrieved
- No additional cost for key management beyond S3 storage
Memory trick: Simple S3 Encryption means AWS handles Keys.
DynamoDB Encryption with CMK
Flip cardDynamoDB encryption at rest with a Customer Managed Key (CMK) uses a key created and managed by the customer within AWS Key Management Service (KMS), providing full control over the key's lifecycle.
- Customer has full control over the key lifecycle (create, rotate, enable, disable)
- Key usage is logged in CloudTrail
- Provides highest level of control for compliance
Memory trick: Customer Managed Key gives full control for DynamoDB.
Amazon Macie
Flip cardAmazon Macie is a data security and data privacy service that uses machine learning to discover, classify, and protect sensitive data (e.g., PII, PCI) stored in Amazon S3.
- Automates discovery of sensitive data in S3.
- Provides visibility into data access patterns and risks.
- Generates findings and alerts for sensitive data exposures.
Memory trick: Macie Makes Monitoring Sensitive Stuff Simple.
S3 Lifecycle Policies with Object Tags
Flip cardS3 Lifecycle Policies automate the management of objects over their lifespan. Object tags enable applying different lifecycle rules to specific subsets of objects within a bucket, providing fine-grained retention and transition control.
- Automates transitions to different storage classes.
- Automates object expiration (deletion).
- Object tags allow multiple, distinct policies within one bucket.
Memory trick: Tags Tailor Timely Transitions & Terminations.
Client-Side Encryption with CloudHSM
Flip cardClient-Side Encryption (CSE) with keys stored in AWS CloudHSM involves encrypting data on the client side before uploading to S3, using master keys managed by the customer within their dedicated FIPS 140-2 Level 3 validated CloudHSM.
- Data encrypted before leaving client control.
- Master keys stored in customer-controlled AWS CloudHSM.
- Provides FIPS 140-2 Level 3 compliance and exclusive key control.
Memory trick: CloudHSM Client Keys Control Cryptography Completely.
AWS CloudHSM
Flip cardAWS CloudHSM is a cloud-based hardware security module (HSM) service that enables you to easily generate and use your own encryption keys on FIPS 140-2 Level 3 validated hardware.
- Dedicated, single-tenant HSMs.
- FIPS 140-2 Level 3 validated.
- Customer controls key generation and storage.
- Integrates with S3 for encryption.
Memory trick: CloudHSM for Highest Key Standards
Redshift Encryption with CMKs
Flip cardEncrypting Amazon Redshift clusters using Customer-managed keys (CMKs) in AWS KMS allows customers to manage and audit their encryption keys, providing greater control over data security.
- Data encrypted at rest in Redshift.
- Keys managed in AWS KMS by the customer.
- Supports manual or automatic key rotation.
- Key usage is auditable via CloudTrail.
Memory trick: KMS Keys Keep Redshift Secure
SSE-KMS
Flip cardServer-Side Encryption with AWS KMS keys (SSE-KMS) is an S3 encryption option that uses customer-managed keys (CMKs) stored in AWS Key Management Service (AWS KMS) to encrypt data at rest.
- Uses CMKs from AWS KMS.
- Provides an audit trail of key usage.
- Gives customers control over encryption key permissions and rotation.
Memory trick: KMS Keys Keep Sensitive Secrets Safe.
Lake Formation Column-Level Security
Flip cardAWS Lake Formation enables column-level security for data lake tables in the Glue Data Catalog, allowing administrators to specify which columns individual users or roles can query, effectively masking sensitive data from unauthorized departments.
- Fine-grained access control for data lakes.
- Integrates with Glue Data Catalog and Athena.
- Restricts access to specific columns.
- Simplifies compliance and data privacy.
Memory trick: Lake Forms Column Limits
CloudTrail for DynamoDB
Flip cardAWS CloudTrail records API activity for Amazon DynamoDB, capturing all control plane and data plane operations as events, providing an audit trail for security and compliance.
- Logs API calls, including item-level data plane operations.
- Logs are immutable and stored in S3.
- Supports long-term retention for compliance.
Memory trick: CloudTrail Captures All DynamoDB Deeds.
S3 Lifecycle Policies with Tags
Flip cardS3 Lifecycle Policies automate the transition of objects to different storage classes or their expiration (deletion) based on rules. Object tags allow applying distinct policies to specific subsets of objects within a bucket.
- Automates storage class transitions and object expiration.
- Supports rules based on object age, creation date, or last access time.
- Object tags enable granular policy application within a single bucket.
Memory trick: Tags Tailor Timely Transitions & Terminations.
AWS Lake Formation Fine-Grained Access
Flip cardAWS Lake Formation provides centralized, fine-grained access control for data lakes, allowing permissions to be defined at the database, table, column, and row level for data stored in S3 and queried by services like Athena.
- Centralized permissions management for data lake resources.
- Supports column-level and row-level security.
- Integrates with query engines like Amazon Athena and Amazon Redshift Spectrum.
Memory trick: Lake Formation: Level-Up Data Access Control.
Lake Formation Fine-Grained Access
Flip cardAWS Lake Formation provides granular access control for data lakes, enabling row-level, column-level, and cell-level security to filter data based on user identity or attributes without duplicating data.
- Centralized access management for S3 data lakes.
- Filters data at query time.
- Supports various data access patterns.
- Simplifies compliance with data privacy regulations.
Memory trick: Lake Forms Fine-Grained Views
CloudTrail with S3 Object Lock
Flip cardAWS CloudTrail logs API activity and events across AWS services. When combined with Amazon S3 Object Lock in WORM mode, it ensures that these audit logs are immutable and cannot be deleted or overwritten for a specified retention period.
- CloudTrail records API calls and events.
- S3 Object Lock provides WORM protection for data.
- Together, they ensure immutable audit trails for compliance.
Memory trick: CloudTrail Charts Logs, S3 Locks 'em Tight.
Redshift Encryption with KMS
Flip cardAmazon Redshift encryption with AWS KMS protects data at rest in a Redshift cluster using customer-managed keys (CMKs) from AWS Key Management Service (KMS), which are backed by FIPS 140-2 Level 2 validated HSMs.
- Encrypts data at rest in Redshift clusters.
- Uses CMKs from AWS KMS.
- Provides auditability of key usage via CloudTrail and customer control over key policies.
Memory trick: KMS Keys Keep Redshift Records Secure.
S3 Access Logs
Flip cardS3 Access Logs provide detailed records for requests made to an S3 bucket, capturing information about who accessed what, when, and how, including successful and failed attempts.
- Deliver log files to a specified S3 bucket.
- Capture object-level operations.
- Essential for auditing and compliance requirements.
Memory trick: Access Logs: Always Capture All S3 Stuff.
Kinesis SSE with KMS
Flip cardKinesis Data Streams Server-Side Encryption (SSE) with AWS KMS keys encrypts data as it is written to the stream and decrypts it as it is read, using customer-managed keys (CMKs) from AWS KMS.
- Encrypts data in transit within the Kinesis stream.
- Uses CMKs from AWS KMS for encryption.
- Enables auditing of key usage and supports key rotation.
Memory trick: KMS Keys Keep Kinesis Streams Secure.
Kinesis SSE with CMK
Flip cardServer-Side Encryption (SSE) for Amazon Kinesis Data Streams using AWS KMS Customer Master Keys (CMKs) encrypts data at rest within the stream, allowing customers to control and audit their encryption keys.
- Encrypts data at rest in Kinesis streams.
- Keys managed by AWS KMS, controlled by customer.
- Supports automatic key rotation.
- Ensures compliance with strict security requirements.
Memory trick: KMS Keys Kinesis Securely
Redshift Data Masking with Views
Flip cardRedshift data masking can be implemented using SQL views that apply masking functions (e.g., `SUBSTRING`, `RPAD`, `MD5`) to sensitive columns. Access to these views can be controlled via user roles.
- Provides dynamic data masking at query time.
- Does not alter the underlying source data.
- Access to masked/unmasked views can be controlled via Redshift user permissions.
Memory trick: Views Veil Vital Information Visually.
CloudTrail with S3 Object Lock & Athena
Flip cardAWS CloudTrail records API calls to AWS services, delivering immutable logs to S3. S3 Object Lock enforces WORM compliance on these logs, and Amazon Athena provides a serverless query engine to analyze them efficiently.
- CloudTrail logs management and data events.
- S3 Object Lock ensures log immutability for compliance.
- Athena allows SQL queries directly on S3-stored logs.
Memory trick: CloudTrail Locks Logs for Long-term Lookup with Athena.
SSE-KMS with CMK
Flip cardServer-Side Encryption with AWS Key Management Service (KMS) and a Customer Managed Key (CMK) allows AWS to encrypt data at rest using keys managed within KMS, giving the customer full control over the encryption key's lifecycle and permissions.
- Uses AWS KMS to manage encryption keys.
- Customer has full control over CMKs (create, rotate, disable, delete).
- KMS uses FIPS 140-2 validated hardware security modules (HSMs).
Memory trick: KMS Keeps Keys Secure for Customers.
CloudTrail for Governance
Flip cardAWS CloudTrail provides a record of actions taken by a user, role, or an AWS service, serving as a critical tool for governance, compliance, and auditing by logging API calls and changes to resources.
- Records API calls and events.
- Tracks changes to policies, roles, keys.
- Logs 'who, what, when, where' of changes.
- Essential for compliance and security audits.
Memory trick: CloudTrail Changes Tracked
CloudTrail S3 Data Events + Kinesis Firehose
Flip cardA robust solution for capturing and delivering all Amazon S3 object-level API activity (read/write) for auditing and compliance to analytics destinations like Amazon Redshift.
- CloudTrail captures S3 Data Events (GetObject, PutObject, etc.)
- CloudTrail logs are delivered to an S3 bucket
- Kinesis Data Firehose can ingest these logs from S3 or direct stream
- Firehose reliably delivers logs to Redshift for analysis
Memory trick: CloudTrail watches S3 like a detective, and Firehose delivers the evidence to Redshift.
AWS Transfer Family
Flip cardA fully managed service that enables the transfer of files directly into and out of Amazon S3 using SFTP, FTPS, and FTP.
- Supports SFTP, FTPS, and FTP protocols
- Directly integrates with Amazon S3 and Amazon EFS
- Fully managed, serverless, and highly available
- Simplifies secure file transfers for external partners
Memory trick: Transfer Family makes file handoffs smooth and secure, like a trusted courier.
AWS IoT Core with Rules Engine
Flip cardA managed cloud platform that lets connected devices easily and securely interact with cloud applications and other devices, using a Rules Engine for data processing and routing.
- Secure device connectivity and management
- MQTT, HTTP, and WebSockets protocol support
- Rules Engine for serverless data transformation and routing
- Integrates with numerous AWS services (S3, SNS, Lambda, Kinesis, etc.)
Memory trick: IoT Core is the central brain for all your smart devices, directing their data wherever it needs to go.
Automated Partition Discovery
Flip cardThe process of automatically updating metadata catalogs (like AWS Glue Data Catalog) with new data partitions created in a data lake, ensuring immediate data discoverability for query engines.
- Crucial for data lakes with frequently updated data.
- Prevents manual metadata management overhead.
- Enables query engines to access the latest data.
Memory trick: Events trigger Glue to get new parts.
Apache ORC
Flip cardAn open-source, columnar data storage format optimized for big data processing engines like Apache Hive, Spark, and Presto. It provides efficient data compression and query performance.
- Columnar storage minimizes I/O for analytical queries.
- Supports various compression codecs.
- Includes strong type support and predicate pushdown for query optimization.
Memory trick: Columnar formats cut query costs.