CompTIA DataSys+ (DS0-001) flashcards
160 free flashcards. Tap a card to flip it.
Normalization
Flip cardThe process of organizing the columns and tables of a relational database to minimize data redundancy and improve data integrity.
- Involves breaking down large tables into smaller, related tables.
- Defined by Normal Forms (1NF, 2NF, 3NF, BCNF, etc.).
- Increases data integrity and reduces update anomalies.
Memory trick: Normalization cleans up your data house.
Encryption in Transit
Flip cardThe process of encoding data as it moves across a network to protect it from interception and unauthorized access.
- Protects data during network communication.
- Commonly uses protocols like TLS/SSL or IPsec.
- Ensures confidentiality and integrity of transmitted data.
Memory trick: Data traveling needs a safe tunnel, not an open road.
PostgreSQL Network Configuration
Flip cardConfiguring PostgreSQL to accept connections from specific network interfaces and IP addresses for secure and accessible database operations.
- The `listen_addresses` parameter controls which IP addresses the server listens on.
- Setting `listen_addresses` to '0.0.0.0' allows connections from all network interfaces.
- The `pg_hba.conf` file controls client authentication and access permissions.
Memory trick: Listen to the address, authenticate the host, open the port.
Encryption in Transit (TLS)
Flip cardThe process of encrypting data as it travels across a network, protecting it from interception and eavesdropping during transmission.
- Uses protocols like TLS/SSL for secure communication.
- Protects data from client to server and vice-versa.
- Essential for public networks and regulatory compliance (e.g., HIPAA, GDPR).
Memory trick: Data in transit needs a secure tunnel to fly through.
Pseudonymization
Flip cardA data protection technique where personally identifiable information (PII) is replaced with artificial identifiers (pseudonyms) to reduce data's linkability to an individual.
- Differs from anonymization by retaining the possibility of re-identification with additional information.
- Helps comply with privacy regulations like GDPR.
- Balances privacy with data utility for analysis or retention.
Memory trick: Privacy techniques are like different ways to hide or disguise identity.
Data Retention Policy
Flip cardA data retention policy defines how long specific types of data must be kept and how they should be securely disposed of after their retention period, often driven by legal, regulatory, or business requirements.
- Specifies data lifespan.
- Includes secure disposal methods.
- Crucial for compliance and data minimization.
Memory trick: PCI says 'No' to some data after use.
Database Network Security
Flip cardDatabase network security involves controlling access to database ports and encrypting communication channels to protect data in transit and prevent unauthorized access.
- Implement firewalls to restrict inbound/outbound traffic.
- Use the principle of least privilege for network access (only allow necessary IPs).
- Encrypt data in transit using SSL/TLS.
- Avoid exposing database ports directly to the internet.
Memory trick: Restrict 'R'eally 'R'equired 'R'outes.
Database Activity Monitoring (DAM)
Flip cardDatabase Activity Monitoring (DAM) is a security technology that monitors and analyzes database events in real-time, providing visibility into database activity, detecting anomalies, and generating alerts for suspicious behavior.
- Monitors SQL traffic and administrative commands.
- Detects anomalous behavior and policy violations.
- Provides real-time alerts and audit trails.
Memory trick: DAM watches the database dance.
Virtual Machine Resource Reservation
Flip cardResource reservation in virtualization guarantees a specific amount of CPU, memory, and/or I/O resources to a virtual machine, preventing resource contention.
- Ensures consistent performance for critical applications like databases.
- Prevents 'noisy neighbor' syndrome where other VMs consume shared resources.
- Sets a minimum threshold for resource allocation that the hypervisor must enforce.
- Crucial for production databases where predictable performance is paramount.
Memory trick: Reservations 'R'eally 'R'eserve 'R'esources.
HAVING Clause
Flip cardA SQL clause used to filter the results of a `GROUP BY` clause based on aggregate function conditions.
- Filters groups, not individual rows.
- Always comes after `GROUP BY`.
- Can use aggregate functions (e.g., SUM, AVG, COUNT) in its condition.
Memory trick: WHERE filters rows, HAVING filters groups.
CREATE INDEX
Flip cardA SQL DDL command used to create an index on a table, which helps retrieve data from the database more quickly.
- Improves the speed of data retrieval (SELECT statements).
- Can be created on single or multiple columns.
- Consumes disk space and can slow down data modification operations (INSERT, UPDATE, DELETE).
- Primary keys automatically have a unique index.
Memory trick: CREATE INDEX is the speed boost for your queries.
Composite Non-Clustered Index
Flip cardA non-clustered index created on multiple columns, optimized for queries that filter or sort by those columns in a specific order.
- The order of columns in the index is crucial for query optimization.
- Can be used to cover queries, meaning all needed data is in the index itself.
- Useful for queries with `WHERE` clauses on leading columns and `ORDER BY` on subsequent columns.
Memory trick: Composite Index: 'Combine' for 'Complex' queries.
Slowly Changing Dimension (SCD) Type 2
Flip cardA method for handling changes to dimension attributes in a data warehouse by creating a new record for each change, preserving the full history of the attribute.
- Maintains a complete historical record of dimension attribute values.
- Typically uses start_date, end_date, and often a 'current' or 'active' flag.
- Allows for 'as-of' reporting, showing how data appeared at any point in time.
Memory trick: Type 1 forgets, Type 2 remembers all, Type 3 remembers some.
InnoDB Buffer Pool
Flip cardA memory area in MySQL's InnoDB storage engine that caches data and indexes, reducing disk I/O for frequently accessed data.
- Crucial for performance in transactional workloads.
- Larger size generally improves performance until memory limits are hit.
- Holds both data and index pages.
Memory trick: Fast transactions need a big memory pool to swim in.
Database Indexing Strategies
Flip cardTechniques for optimizing database query performance by creating data structures that allow for faster retrieval of records based on column values.
- B-tree indexes are good for range queries and equality checks on high-cardinality data.
- Hash indexes are efficient for equality lookups but not for range queries.
- Bitmap indexes are ideal for data warehousing on low-cardinality columns, enabling efficient combination of multiple conditions.
- Full-text indexes are for searching text data.
Memory trick: Index for access patterns: B-tree for ranges, hash for equality, bitmap for analytics.
Master-Slave Replication with Read Replicas
Flip cardA database deployment strategy where a primary (master) database handles all write operations, and one or more secondary (slave or replica) databases receive replicated data and serve read-only queries.
- Improves read scalability by distributing queries to replicas.
- Enhances high availability; a replica can be promoted to master upon failure.
- Introduces potential for replication lag between master and replicas.
- Common in relational databases like MySQL, PostgreSQL, and cloud services like AWS RDS.
Memory trick: Master 'M'akes 'M'any 'R'eads 'R'eplicate.
Automated VACUUM
Flip cardAn automatic background process in PostgreSQL and similar systems that reclaims storage occupied by 'dead' row versions and updates data statistics.
- Prevents table bloat and improves query performance.
- Runs automatically without manual intervention.
- Essential for long-term database health.
Memory trick: VACUUM cleaner tidies up the database's dusty corners.
UNIQUE Constraint (with NULLs)
Flip cardA database constraint that ensures all values in a specified column or set of columns are distinct, with the exception that most database systems allow for one (or sometimes multiple) NULL values.
- Enforces uniqueness for non-NULL values.
- Typically allows one NULL value (SQL Standard, PostgreSQL, MySQL).
- SQL Server allows multiple NULL values in a UNIQUE constraint.
Memory trick: Unique emails, but some can be 'null' and still be unique.
Anonymization
Flip cardAnonymization is a data privacy technique that transforms data to prevent the re-identification of individuals while preserving the data's analytical utility. It typically involves removing or altering direct and indirect identifiers.
- Prevents re-identification of individuals.
- Retains statistical value of the data.
- Often considered irreversible.
- Techniques include generalization, suppression, perturbation.
Memory trick: Anonymization makes data nameless but useful.
NoSQL Database Types
Flip cardDifferent categories of non-relational databases designed to handle large volumes of diverse data, offering flexible schemas and horizontal scalability.
- Document databases store data in flexible, semi-structured documents (e.g., JSON).
- Key-value stores are simple, fast, and highly scalable for basic data retrieval.
- Column-family databases are optimized for large-scale data with high write throughput.
- Graph databases model and query data based on relationships between entities.
Memory trick: Schema flexibility, data structure, scale, and query needs guide choice.
BASE Properties
Flip cardAn acronym (Basically Available, Soft state, Eventually consistent) that describes the properties of many NoSQL databases, prioritizing availability and partition tolerance over strict consistency.
- Basically Available: The system guarantees availability.
- Soft state: The state of the system can change over time, even without input.
- Eventually consistent: Data will become consistent across all replicas over time, but not immediately.
Memory trick: ACID is for banks, BASE is for web-scale!
Backup Storage Efficiency
Flip cardBackup storage efficiency refers to minimizing the amount of disk space or network bandwidth consumed by backup operations. Incremental backups are typically the most storage-efficient as they only capture changes since the last backup, contrasting with full or differential backups which copy more data.
- Incremental backups save only new or changed data.
- Reduces storage footprint and backup window.
- Recovery can be more complex due to multiple restore steps.
Memory trick: Inc: Smallest save, longest restore. Diff: Balanced, but grows more.
PostgreSQL pg_hba.conf
Flip cardThe PostgreSQL Host-Based Authentication (pg_hba.conf) file controls which hosts are allowed to connect to the database, and how they must authenticate.
- It's a critical security configuration file.
- Defines connection rules based on host, database, user, and authentication method.
- Read in order, first match applies.
Memory trick: Host-Based Authentication is the Gatekeeper for PostgreSQL.
Bitmap Index
Flip cardA database index type that uses bitmaps (binary arrays) to efficiently store and query data, particularly effective for columns with low cardinality in data warehousing environments.
- Excellent for complex analytical queries with multiple predicates (AND/OR).
- Best suited for columns with a small number of distinct values (low cardinality).
- Less efficient for high-cardinality columns or transactional workloads (high update rates).
Memory trick: Bitmap: Bits for Big Analytics.
PostgreSQL High Availability
Flip cardEnsuring continuous operation of a PostgreSQL database system, often involving replication, monitoring, and automatic failover mechanisms.
- Streaming replication is foundational for HA.
- Tools like Patroni or pg_auto_failover automate failover.
- Requires robust monitoring to detect primary failures.
Memory trick: HA for PostgreSQL needs a smart conductor.
Columnar Database
Flip cardA type of database that stores data in columns rather than rows, optimizing for analytical queries, aggregations, and data compression, especially in data warehousing scenarios.
- Excellent for OLAP workloads.
- Efficient for reading specific columns across many rows.
- High data compression rates.
Memory trick: Choose your NoSQL weapon wisely for the data battle.
Cloud Database Scaling
Flip cardThe ability of a cloud database to adjust its resources (CPU, memory, storage) dynamically to meet changing workload demands.
- Vertical scaling: increasing resources of a single instance.
- Horizontal scaling: adding more instances to distribute load.
- On-demand scaling: resources adjust automatically based on real-time usage.
- Cost-effective for variable workloads by paying only for consumed resources.
Memory trick: Scale up and down, save your crown.
PostgreSQL shared_buffers
Flip cardA core PostgreSQL configuration parameter that sets the amount of shared memory used by the database server for caching data blocks and indexes.
- Crucial for database performance; higher values generally mean more data can be cached.
- Should typically be set to 25% of total system RAM, up to a certain limit.
- Too high a value can lead to excessive swapping or OOM errors.
Memory trick: Shared buffers for shared data blocks.
Oracle Listener
Flip cardA background process on an Oracle database server that listens for incoming client connection requests and manages traffic to the database instance.
- Uses the `lsnrctl` utility for management.
- Must be running for remote clients to connect.
- Configured via `listener.ora` file.
Memory trick: Listener Control (lsnrctl) lets you hear the database.
Serverless Database
Flip cardA cloud database deployment model where the database automatically scales its compute and storage resources up or down based on demand, and users are billed only for actual usage.
- Automatic scaling of compute and storage.
- Pay-per-use billing model.
- Ideal for unpredictable and spiky workloads.
- Reduces operational overhead for database administrators.
Memory trick: Scaling in the cloud: how much control, how much auto?
Time-Series Database (TSDB)
Flip cardA Time-Series Database (TSDB) is a database optimized for storing and retrieving time-stamped data, such as sensor readings or financial market data, offering high ingestion rates and efficient time-based queries.
- Optimized for timestamped data.
- Handles very high write (ingestion) rates.
- Efficient for time-range queries and aggregations.
- Commonly used for IoT, monitoring, and financial data.
Memory trick: For 'time' and 'series' data, use the database that has 'time-series' in its name!
Always Encrypted with Secure Enclaves
Flip cardA SQL Server feature that enhances Always Encrypted by enabling in-database computations on encrypted data using secure enclaves, without exposing data or keys to the database engine.
- Maintains client-side encryption.
- Allows rich computations (e.g., pattern matching, range queries) on encrypted data.
- Uses hardware-based secure enclaves for enhanced security.
Memory trick: Encryption: Protect data at rest, in transit, and even during processing.
Master-Replica with Synchronous Read Replicas and Automatic Failover
Flip cardA high-availability and read-scaling database architecture where a primary (master) writes data synchronously to replicas, which can serve reads and be promoted automatically upon primary failure.
- Ensures strong data consistency.
- Provides high availability through automatic failover.
- Scales read operations by distributing them among replicas.
Memory trick: HA is a dance of consistency, copies, and quick recovery.
PostgreSQL Client Authentication
Flip cardThe process by which PostgreSQL verifies the identity of a connecting client and determines if it is allowed to access the database.
- Controlled primarily by the `pg_hba.conf` file.
- Entries specify connection type, database, user, IP address/range, and authentication method.
- Rules are processed in order, with the first matching rule applied.
Memory trick: HBA: Host-Based Access, the bouncer for your database.
Database Storage Sizing
Flip cardThe process of estimating the total disk space required for a database, considering raw data, indexes, overhead, and replication/backup needs.
- Always account for overhead (indexes, logs, temp files).
- Factor in replication for HA/DR.
- Consider future growth and retention policies.
Memory trick: Size it up: Raw, Overhead, then Replicate!
Online Storage Expansion
Flip cardThe process of increasing the storage capacity of a database system while the database and its applications remain fully operational, without requiring downtime.
- Typically achieved with RAID controllers or storage arrays supporting online expansion.
- Allows adding disks, replacing smaller disks with larger ones, or expanding logical volumes.
- Minimizes disruption to critical applications.
Memory trick: Online expansion keeps your data flowing.
MySQL Secure Installation
Flip cardA post-installation script for MySQL servers that guides users through essential security hardening steps to make the database more secure.
- Sets root password.
- Removes anonymous users.
- Disallows remote root login.
- Removes test database and associated privileges.
Memory trick: Secure your MySQL, don't leave it open!
Online Storage Expansion (LVM)
Flip cardOnline storage expansion using LVM involves extending a Logical Volume and then resizing its filesystem while the system and database remain operational, minimizing downtime for critical applications.
- Leverages Logical Volume Management (LVM) on Linux.
- Allows adding physical disks to existing volume groups.
- Enables online extension of Logical Volumes and filesystems (e.g., XFS, Ext4).
- Crucial for high-availability database systems.
Memory trick: LVM is like a 'stretchable' hard drive – you can grow it while it's still running.
Non-Clustered Index
Flip cardA non-clustered index is a special type of index that creates a separate, sorted structure containing key values and pointers to the actual data rows in the table. It does not alter the physical order of the data.
- Stores data in one logical order and index in another physical order.
- Faster for lookups and joins on columns that are not the primary key.
- A table can have multiple non-clustered indexes.
Memory trick: Indexes guide your data search, not sort your books.
Active-Active Cluster
Flip cardAn Active-Active cluster configuration where all nodes in the cluster are actively processing requests and can automatically take over the workload of a failed node.
- All nodes process requests concurrently.
- Provides high availability and load balancing.
- Automatic failover with minimal downtime.
- More complex to set up and manage than active-passive.
Memory trick: Always Active, Always Available.