Microsoft Azure Data Fundamentals flashcards
95 free flashcards. Tap a card to flip it.
Time Series Database
Flip cardA database optimized for storing and retrieving time-stamped data, like sensor readings or stock prices, efficiently.
- Optimized for time-stamped data
- High ingestion rates
- Efficient time-based queries
Memory trick: Time series data is like a 'timeline' of events.
Aggregation
Flip cardAggregation is a data processing operation that combines multiple data points into a single summary value by applying functions like SUM, AVG, COUNT, MIN, or MAX.
- Summarizes data.
- Often used with grouping (e.g., GROUP BY).
- Common in analytical queries and reporting.
Memory trick: Aggregations 'gather' and 'sum' up the data.
Extract, Transform, Load (ETL)
Flip cardETL is a data integration process that involves extracting data from source systems, transforming it into a desired format, and loading it into a target data store, typically a data warehouse.
- Consists of three main stages: Extract, Transform, Load.
- Used for data warehousing and data integration.
- Prepares data for analytical reporting.
Memory trick: ETL is the 'E'asy 'T'rick for 'L'oading data.
Online Transaction Processing (OLTP)
Flip cardA class of software programs capable of supporting transaction-oriented applications, typically for data entry and retrieval transaction processing.
- Handles high volume of short, atomic transactions
- Optimized for reads and writes
- Emphasizes data integrity and concurrency
Memory trick: OLTP is for 'transactions' happening 'online' right now.
Stream Processing
Flip cardA data processing paradigm where data is processed continuously as it arrives, rather than in batches.
- Handles data in motion, often from IoT devices or clickstreams.
- Enables real-time analytics, anomaly detection, and immediate actions.
- Requires low latency and high throughput.
Memory trick: Streams for Sensors, Not Stagnant Data
Online Analytical Processing (OLAP)
Flip cardOLAP is a category of software tools that provide analysis of data for business intelligence. OLAP systems are optimized for complex data analysis rather than transaction processing.
- Focuses on historical data and trend analysis.
- Optimized for complex queries and aggregations.
- Used for business intelligence, reporting, and forecasting.
Memory trick: OLAP helps you 'look' at the 'past'.
Key-Value Database
Flip cardA type of NoSQL database that stores data as a collection of key-value pairs, where a key serves as a unique identifier for its associated value.
- Optimized for fast read/write operations.
- Values can be simple data types or complex objects.
- Does not enforce a schema for values.
- Ideal for session management, user preferences, and caching.
Memory trick: Key-Value pairs are like a fast dictionary for your data.
Document Data Model
Flip cardA NoSQL data model that stores data in flexible, semi-structured documents, often in formats like JSON or XML. Each document is a self-contained unit.
- Offers schema flexibility, allowing documents to have different fields.
- Supports nested data structures.
- Ideal for catalogs, user profiles, and content management.
- Commonly used in Document Databases like MongoDB or Azure Cosmos DB's Document API.
Memory trick: Documents are like flexible folders, holding whatever you put inside.
Relational Data Model
Flip cardOrganizes data into one or more tables (relations) of rows and columns, with predefined schemas and strong consistency.
- Enforces data integrity through schemas, primary keys, and foreign keys.
- Provides ACID (Atomicity, Consistency, Isolation, Durability) properties.
- Uses SQL for querying and data manipulation.
Memory trick: Relational Ensures Reliability for Records
Unstructured Data
Flip cardData that does not have a predefined data model or is not organized in a pre-defined manner. It accounts for a large majority of data generated today.
- Cannot be stored in traditional relational databases without significant effort.
- Examples include text documents, emails, social media posts, audio, video files.
- Requires advanced analytics techniques (NLP, machine learning) to extract insights.
- Often stored in data lakes or NoSQL document databases.
Memory trick: Unstructured is like a wild, free-flowing river of information.
Graph Data Model
Flip cardA NoSQL data model that represents data as nodes (entities) and edges (relationships between entities). It is optimized for storing and querying highly connected data.
- Nodes represent entities (e.g., user, post).
- Edges represent relationships (e.g., 'follows', 'likes').
- Ideal for social networks, recommendation engines, fraud detection.
- Examples include Neo4j, Azure Cosmos DB's Gremlin API.
Memory trick: Graphs show how everything is 'connected'!
Data Lake
Flip cardA centralized repository that stores a vast amount of raw data in its native format, including structured, semi-structured, and unstructured data.
- Stores data without predefined schema (schema-on-read).
- Highly scalable and cost-effective for large volumes.
- Supports various data processing frameworks (e.g., Spark, Hadoop).
- Used for big data analytics, machine learning, and data exploration.
Memory trick: A Data Lake is a big, raw pool where all your data can swim freely.
Azure Time Series Insights
Flip cardAzure Time Series Insights is an end-to-end platform-as-a-service (PaaS) for collecting, processing, storing, querying, and visualizing time series data at IoT scale.
- Optimized for IoT and time-series data.
- Enables near real-time and historical analysis.
- Provides rich visualization and query capabilities over time-stamped data.
Memory trick: Time Series Insights helps you 'see' the 'time' and 'stream' it.
Batch Processing
Flip cardA method of processing data in large groups (batches) at scheduled intervals, rather than in real-time.
- Processes data in chunks
- Scheduled execution
- Suitable for large volumes, non-real-time needs
Memory trick: Batch processing is like baking a 'batch' of cookies once a day.
NoSQL Database
Flip cardNoSQL databases provide a mechanism for storage and retrieval of data that is modeled in means other than the tabular relations used in relational databases.
- Handles semi-structured and unstructured data.
- Offers flexible schemas.
- Provides high scalability and availability.
Memory trick: No-SQL is no problem for flexible data.
Key-Value Data Model
Flip cardA type of NoSQL data model that stores data as a collection of key-value pairs, where each key is unique and maps to a single value.
- Optimized for fast read and write operations.
- Values can be simple data types or complex objects.
- Commonly used for caching, session management, and real-time data.
Memory trick: Keys Unlock Value: Choose the right model for your data's journey.
Azure Event Hubs
Flip cardA highly scalable data streaming platform and event ingestion service that can receive and process millions of events per second.
- Designed for high-throughput, low-latency data ingestion.
- Ideal for IoT telemetry, application logging, clickstreams.
- Acts as a 'front door' for data pipelines, buffering events.
- Integrates with stream processing engines like Azure Stream Analytics.
Memory trick: Event Hubs is the 'express lane' for all your streaming data.
Azure Blob Storage (Archive Tier)
Flip cardA storage tier in Azure Blob Storage optimized for rarely accessed data, offering the lowest storage costs with flexible latency requirements (hours for retrieval).
- Lowest cost storage option for Azure Blob Storage.
- Ideal for long-term backup, archival, and regulatory compliance data.
- Data retrieval can take several hours.
- Supports various data types (structured, semi-structured, unstructured).
Memory trick: Archive is like putting data in a deep, cold freezer for safekeeping.
Azure Cognitive Search
Flip cardA cloud search service that gives developers APIs and tools for adding a rich search experience to their applications.
- Provides full-text search, faceted navigation, and geospatial search.
- Can index various data sources, including structured data, Blob Storage (for PDFs), and databases.
- Integrates AI capabilities for enriching content, such as entity recognition and sentiment analysis.
Memory trick: Cognitive Search for Content and Context
Azure Database for PostgreSQL
Flip cardAzure Database for PostgreSQL is a managed relational database service in the cloud, based on the open-source PostgreSQL database engine.
- Supports structured data and relational models.
- Enforces primary and foreign key constraints.
- Provides ACID transactional consistency.
Memory trick: PostgreSQL is 'post'-perfect for structured data.
Azure Synapse Serverless SQL Pool
Flip cardA serverless query service in Azure Synapse Analytics that enables users to analyze data directly in Azure Data Lake Storage Gen2 using T-SQL.
- Queries data without moving or transforming it.
- Pay-per-query model based on data processed.
- Supports various file formats like Parquet, CSV, and JSON.
Memory trick: Serverless queries explore the data lake.
Power BI
Flip cardA business intelligence service that provides interactive visualizations and business intelligence capabilities with an interface simple enough for end-users to create their own reports and dashboards.
- Connects to a wide range of data sources.
- Enables creation of interactive reports and dashboards.
- Supports data exploration, drill-down, and automatic refreshes.
Memory trick: Power BI gives you the power to see your data.
Time-Series Database
Flip cardA database optimized for storing and analyzing data points that are indexed by time, commonly used for monitoring, sensor data, and IoT applications.
- Designed for data with a timestamp component.
- Efficient for high-volume inserts and time-based queries.
- Supports aggregation and analysis over time windows.
Memory trick: Time-travelers always check the clock.
Azure Synapse Analytics dedicated SQL pool
Flip cardA massively parallel processing (MPP) data warehousing service in Azure Synapse Analytics, optimized for large-scale analytical workloads.
- Uses columnar storage for efficient analytical queries.
- Scales compute and storage independently.
- Ideal for enterprise data warehousing and business intelligence.
Memory trick: Synapse's SQL pool is the warehouse for big insights.
Azure Synapse Dedicated SQL Pool Pause/Resume
Flip cardA feature of Azure Synapse Dedicated SQL pools that allows users to pause compute resources when not in use, significantly reducing operational costs.
- Pausing stops compute billing while retaining data.
- Resuming restarts compute, typically within minutes.
- Ideal for cost optimization during non-business hours or inactive periods.
Memory trick: Dedicated pool, dedicated savings: pause the power!
Azure Data Factory
Flip cardA cloud-based ETL and data integration service that allows you to create data-driven workflows for orchestrating data movement and transforming data at scale.
- Supports connections to hundreds of data sources.
- Provides visual tools for pipeline development.
- Enables scheduling and monitoring of data integration processes.
Memory trick: Data Factory 'builds' your data pipelines.
Azure Stream Analytics (ASA)
Flip cardA real-time analytics service that processes high volumes of streaming data from various sources to detect patterns and relationships.
- Designed for real-time data processing and analysis.
- Supports SQL-like query language for stream processing.
- Ideal for IoT, log analytics, and clickstream analysis.
Memory trick: Streams are analyzed as they flow.
Azure Data Lake Storage Gen2 Archive Tier
Flip cardA highly cost-effective storage tier within Azure Data Lake Storage Gen2, designed for rarely accessed, long-term data retention.
- Lowest cost storage option for ADLS Gen2.
- Optimized for data that can tolerate longer retrieval times.
- Ideal for compliance, audit, and long-term backup data.
Memory trick: Archive your old data to save big bucks over time.
Data Integration
Flip cardThe process of combining data from different sources into a unified view, often involving cleaning, transformation, and standardization.
- Essential for creating a holistic view from disparate datasets.
- Involves ETL (Extract, Transform, Load) or ELT processes.
- Prepares data for analytics, reporting, and machine learning.
Memory trick: Integration brings all the data pieces together into one picture.
Azure Synapse Analytics
Flip cardA unified analytics service that brings together enterprise data warehousing, big data analytics, and data integration into a single platform.
- Combines SQL technologies for data warehousing with Spark technologies for big data.
- Offers a unified experience for data ingestion, exploration, preparation, transformation, management, and serving.
- Supports both serverless and dedicated resource models for different workloads.
Memory trick: Synapse Unites all Analytics, a true Data Hub.
Azure Data Factory (ADF)
Flip cardA cloud-based ETL (Extract, Transform, Load) and ELT (Extract, Load, Transform) service for orchestrating and automating data movement and transformation.
- Used for building complex data pipelines.
- Supports various data sources and destinations.
- Enables data integration and transformation at scale.
Memory trick: The data factory builds the data path.
Azure Stream Analytics
Flip cardA real-time analytics service designed to process high volumes of streaming data with low latency.
- Processes data in motion from sources like Event Hubs and IoT Hub.
- Supports SQL-like queries for complex event processing.
- Enables real-time dashboards, alerts, and data archiving.
Memory trick: Stream Analytics flows like a river, processing data as it comes.
Data Lake Storage Gen2
Flip cardA highly scalable and cost-effective data lake solution built on Azure Blob Storage, optimized for big data analytics workloads.
- Supports petabytes of data and trillions of files.
- Offers hierarchical namespace for folder and file organization.
- Combines benefits of Blob Storage (cost-effectiveness, tiered storage) with Data Lake Store Gen1 (file system semantics).
Memory trick: Think of a lake for all your data, big and small.
Azure Databricks
Flip cardA unified, Apache Spark-based analytics platform optimized for Azure, providing collaborative workspaces for data engineering, data science, and machine learning.
- Managed Apache Spark clusters for big data processing.
- Supports Python, Scala, R, SQL for data science and ML.
- Ideal for large-scale ETL, stream processing, and advanced analytics.
Memory trick: Databricks is where data scientists get to 'spark' their models.
Azure Synapse Analytics Serverless SQL Pool
Flip cardA serverless query service within Azure Synapse Analytics that enables querying data directly in data lakes using T-SQL.
- No infrastructure to set up or manage.
- Pay-per-query model, ideal for ad-hoc analysis.
- Supports various file formats like Parquet, CSV, JSON.
Memory trick: Serverless SQL queries the lake without needing a boat.